Endonuclease system
Patent Information
- Application Number
- JP2024529294
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-29
- Filing Date
- 2022-11-23
- Publication Date
- 2025-12-01
AI Technical Summary
The large size of class 2 Cas effectors poses challenges for therapeutic applications, making delivery difficult.
Development of a novel SMART (SMall ARchaeal-associaTed) nuclease system comprising engineered endonucleases with a molecular weight of 96 kDa or less, derived from uncultured microorganisms, featuring RuvC and HNH domains, and guided by engineered guide RNA structures.
The SMART nuclease system enables efficient and targeted DNA manipulation and gene editing with improved delivery and functionality, suitable for therapeutic applications.
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 282,999, filed November 24, 2021, No. 63 / 289,981, filed December 15, 2021, and No. 63 / 356,908, filed June 29, 2022, each of which is incorporated by reference in its entirety herein.
[0002] This application is related to PCT Application No. PCT / US21 / 24945, which is incorporated herein by reference in its entirety.
[0003] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, which is incorporated herein by reference in its entirety. The XML copy created on November 23, 2022 is named 55921-741_601_SL.xml and is 1,897,990 bytes in size. [Background technology]
[0004] Cas enzymes, along with associated clustered regularly interspaced short palindromic repeats (CRISPR)-guided ribonucleic acid (RNA), appear to be widespread components of the immune system of prokaryotes (~45% of bacteria, ~84% of archaea), where they serve to protect these microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, by CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. Although CRISPR DNA elements have been observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR / Cas complexes have only been recognized relatively recently, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA engineering and gene editing applications. The utility of these enzymes has led to their repurposing for a wide variety of biotechnological, gene editing, and therapeutic applications. Due to their single-effector architecture, the majority of systems currently repurposed for genome engineering belong to the CRISPR class 2 category. Summary of the Invention
[0005] The large size (>1200 amino acids) of many class 2 Cas effectors makes them difficult to deliver for therapeutic applications. Thus, described herein are methods, compositions, and systems for novel putative guide dsDNA nucleases, termed SMART (SMall ARchaeal-associated) nuclease systems. These endonuclease effectors are defined by their small size (~400 aa to ~1050 aa), the presence of RuvC and HNH catalytic domains, and other predicted protein features that together suggest a novel biochemical mechanism.
[0006] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. (1) the endonuclease has a molecular weight of about 96 kDa or less, about 80 kDa or less, about 70 kDa or less, or about 60 kDa or less, and (2) the endonuclease has at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83% identity with an arginine-rich region or a domain having PF14239 homology from any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof; (1) the endonuclease comprises an arginine-rich region or a domain having PF14239 homology with at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity; or (2) the endonuclease comprises an arginine-rich region or a domain having PF14239 homology with at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 99%, or 100% sequence identity with at least 84%, at least 85%, at least 86%, at least 87%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 74-675, 975-1002, 1260-1321, or a variant thereof, and at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%,or (3) the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or variants thereof. In some embodiments, (1) the endonuclease is at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least or (2) the endonuclease comprises an arginine-rich region or a domain having PF14239 homology with at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%,or (3) the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 674-675, 975-1002, or 1260-1321, or variants thereof. In some embodiments, the endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the arginine-rich region or domain having homology to PF14239 is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, or the like, selected from any one of SEQ ID NOs: 1 to 198, 221 to 459, 463 to 612, 617 to 668, 674 to 675, 975 to 1002, and 1260 to 1321.at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the REC domain has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the REC domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321. In some embodiments, the endonuclease further comprises a BH (bridge helix) domain, a WED (wedge) domain, or a PI (PAM interaction) domain or a TI (TAM interaction) domain. In some embodiments, the WED domain or the PI domain has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the BH domain, WED domain, or PI domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, and 1260-1321.
[0007] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC-I domain and an HNH domain; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease, wherein the endonuclease is selected from the group consisting of SEQ ID NOs: 674-675, 975-1002, 1260-1321, 1322-1323, 1324-1325, 1326-1327, 1328-1329, 1330-1331, 1332-1333-1334, 1332-1335, 1332-1336, 1332-1337, 1332-1338, 1332-1339 ... The present invention provides an engineered nuclease system comprising a sequence having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of the following: or a variant thereof. In some embodiments, the endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the arginine-rich region or domain having homology to PF14239 has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the arginine-rich region of any one of SEQ ID NOs: 674-675, 975-1002, and 1260-1321. In some embodiments, the endonuclease further comprises a REC (recognition) domain.In some embodiments, the REC domain has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the REC domain of any one of SEQ ID NOs: 674-675, 975-1002, 1260-1321. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain. In some embodiments, the BH domain, WED domain, or PI domain has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the BH domain, WED domain, or PI domain of any one of SEQ ID NOs: 674-675, 975-1002, and 1260-1321. In some embodiments, the endonuclease is derived from an uncultured microorganism.In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or a sequence having at least 80% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259. In some embodiments, the guide nucleic acid structure comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259.
[0008] In some aspects, the disclosure provides an engineered nuclease system, comprising: (a) an engineered guide ribonucleic acid structure, comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to an endonuclease, wherein the ribonucleic acid sequence has at least 80% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or at least 80% sequence identity to any one of SEQ ID NOs: 677-686, 1006-1012, or 1231-1259. The present invention provides an engineered nuclease system comprising: an engineered guide ribonucleic acid structure comprising a sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity; and (b) an RNA-guided endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the RNA-guided endonuclease is an archaeal endonuclease. In some embodiments, the method, wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide comprising a guide ribonucleic acid sequence and a ribonucleic acid sequence configured to bind an endonuclease. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is about 14 to about 28 nucleotides in length, about 18 to about 26 nucleotides in length, about 22 to about 26 nucleotides in length, or about 24 nucleotides in length.In some embodiments, the guide ribonucleic acid sequence comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 462, 676, or 1229-1230. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from any one of SEQ ID NOs: 205-220. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first or second homologous arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the system comprises Mg. 2+In some embodiments, the endonuclease and the ribonucleic acid sequence configured to bind to the endonuclease are derived from distinct species within the same phylum. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 2-24, and the guide RNA structure comprises an RNA sequence predicted to comprise a hairpin comprising a stem and a loop, the stem comprising at least 10 pairs of ribonucleotides and an intervening multiloop. In some embodiments, the guide RNA structure further comprises a second stem and a second loop, the second stem comprising at least 5 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises an RNA structure comprising at least two hairpins.In some embodiments, a) the endonuclease comprises a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof; and b) the guide RNA structure is , any one of SEQ ID NOs: 199 to 200, 460 to 461, or 669 to 673, or any one of SEQ ID NOs: 201 to 203, 613 to 616, 677 to 686, 1006 to 1012, or 1231 to 1259.In some embodiments, a) the endonuclease is at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% , at least 99%, or 100% sequence identity, and b) the guide RNA structure comprises a sequence with at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a class 2, type II sgRNA or tracr sequence. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using Smith-Waterman homology search algorithm parameters. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11, extension of 1, and using a conditional composition score matrix adjustment.In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease.
[0009] In some aspects, the disclosure provides an engineered nuclease system, comprising: (a) an endonuclease configured to be selective for a target adjacent motif (TAM) sequence comprising any one of AGG (SEQ ID NO: 1029), NARAA (SEQ ID NO: 1030), ATGAAA (SEQ ID NO: 1031), ATGA (SEQ ID NO: 1032), or WTGG (SEQ ID NO: 1033), wherein the endonuclease has at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 122%, at least 124%, at least 126%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least 136%, at least 137%, at least 138%, at least 139%, at least 14 In one embodiment, the invention provides an engineered nuclease system comprising: (a) an endonuclease comprising a TAM interaction domain having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a target nucleic acid sequence; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and the engineered guide RNA comprises a spacer sequence configured such that the engineered guide RNA hybridizes to a target nucleic acid sequence.In some embodiments, the TAM interacting domain is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the TAM interacting domain of SEQ ID NO:674 or a variant thereof. or a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the TAM interacting domain of SEQ ID NO: 675 or a variant thereof. In some embodiments, the endonuclease system comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide RNA is 30-280 nucleotides in length. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first or second homologous arm comprises a sequence of at least 40 nucleotides. In some embodiments, the first and second homologous arms are homologous to a eukaryotic genomic sequence. In some embodiments, the single-stranded or double-stranded DNA repair template comprises a transgene donor. In some embodiments, the system further comprises a DNA repair template comprising a double-stranded DNA segment adjacent to one or two single-stranded DNA segments.In some embodiments, the single stranded DNA segment is conjugated to the 5' end of the double stranded DNA segment. In some embodiments, the single stranded DNA segment is conjugated to the 3' end of the double stranded DNA segment. In some embodiments, the single stranded DNA segment has a length of 4-10 nucleotide bases. In some embodiments, the single stranded DNA segment has a nucleotide sequence complementary to a sequence within the spacer sequence. In some embodiments, the double stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein coding sequence, an miRNA coding sequence, an RNA coding sequence, or a transgene. In some embodiments, the double stranded DNA sequence is flanked by nuclease cleavage sites.
[0010] In some embodiments, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease configured to be selective for a protospacer adjacent motif (PAM) sequence comprising an NRR, wherein the endonuclease interacts with the PAM interaction domain of any one of SEQ ID NOs: 1313-1318 by at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least 136%, at least 137%, at least 138, at least 139, at least 140%, at least 141%, at least 142, The present invention provides an engineered nuclease system comprising: (a) an endonuclease comprising a PAM interaction domain having at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a target nucleic acid sequence; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and the engineered guide RNA comprises a spacer sequence configured to hybridize the engineered guide RNA to a target nucleic acid sequence. In some embodiments, the TAM interaction domain comprises a sequence having at least 80% sequence identity to a TAM interaction domain of SEQ ID NO: 674 or a variant thereof, or at least 80% sequence identity to a TAM interaction domain of SEQ ID NO: 675 or a variant thereof. In some embodiments, the endonuclease system comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide RNA is 30-280 nucleotides in length. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first or second homologous arm comprises a sequence of at least 40 nucleotides.In some embodiments, the first and second homologous arms are homologous to a genomic sequence of the eukaryotic organism. In some embodiments, the single-stranded or double-stranded DNA repair template comprises a transgene donor. In some embodiments, the system further comprises a DNA repair template comprising a double-stranded DNA segment adjacent to one or two single-stranded DNA segments. In some embodiments, the single-stranded DNA segment is conjugated to the 5' end of the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment is conjugated to the 3' end of the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment has a length of 4-10 nucleotide bases. In some embodiments, the single-stranded DNA segment has a nucleotide sequence complementary to a sequence within the spacer sequence. In some embodiments, the double-stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein coding sequence, an miRNA coding sequence, an RNA coding sequence, or a transgene. In some embodiments, the double-stranded DNA sequence is adjacent to a nuclease cleavage site.
[0011] In some embodiments, the disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide; and wherein the engineered guide ribonucleic acid polynucleotide is selected from the group consisting of any one of SEQ ID NOs: 674-675, 975-1002, 1260-1321, or a combination thereof. In some embodiments, the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease comprising a variant having at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the variant. In some embodiments, the DNA targeting segment is located 5' to both of the two complementary stretches of nucleotides.In some embodiments, a) the protein-binding segment comprises a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673; or b) the protein The binding segment comprises a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of the non-variable nucleotides of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259.In some embodiments, a) the endonuclease is at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least a) the guide RNA structure comprises a sequence with at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a class 2, type II sgRNA. In some embodiments, the endonuclease further comprises a base editor or a histone editor coupled to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.
[0012] In some aspects, the disclosure provides a deoxyribonucleic acid polynucleotide encoding any of the engineered guide ribonucleic acid polynucleotides described herein.
[0013] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding an endonuclease comprising a RuvC domain and an HNH domain, the endonuclease being derived from an uncultured microorganism, the endonuclease having a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, 60 kDa or less, or 30 kDa or less, and the endonuclease is selected from the group consisting of SEQ ID NOs: 674-675, 975-1002, 1260-1321, or any of those. and variants thereof having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the endonuclease. In some embodiments, the endonuclease further comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 205-220. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human. In some embodiments, the organism is a prokaryote or a bacterium, and the organism is a different organism than the organism from which the endonuclease is derived. In some embodiments, the organism is not an uncultured microorganism.
[0014] In some aspects, the disclosure provides a vector comprising a nucleic acid sequence encoding an RNA-guided endonuclease comprising a RuvC-I domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less, the RNA-guided endonuclease is optionally an archaeal bacteria, and the RNA-guided endonuclease is selected from the group consisting of SEQ ID NOs: 674-675, 975-1002, 1260-132, 1320-132, 1325-1326, 1326-1327, 1325-1328, 1325-1329, 1326-1329, 1327 ... 1, or variants thereof having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity thereto. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a target adjacent motif (TAM) interaction (TI) domain. In some embodiments, the TI domain comprises any one of the TI domains of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, and 1260-1321.
[0015] In some aspects, the disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with an endonuclease, the engineered guide ribonucleic acid structure comprising a) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and b) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
[0016] In some aspects, the disclosure provides a cell comprising any of the vectors described herein. In some embodiments, the cell is a bacterial, archaeal, fungal, eukaryotic, mammalian, or plant cell. In some embodiments, the cell is a bacterial cell.
[0017] In some aspects, the disclosure provides a method of producing an endonuclease comprising culturing any of the cells described herein.
[0018] In some aspects, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease in a complex with the endonuclease and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide; and (b) the double-stranded deoxyribonucleic acid polynucleotide comprises a target adjacent motif (TAM), and the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide, and the TAM comprises any one of SEQ ID NOs: 1023-1044. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide from the TAM by 5-7 nucleotides, 5 nucleotides, 6 nucleotides, or 7 nucleotides. In some embodiments, the endonuclease comprises a variant having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, and 1260-1321.
[0019] In some aspects, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising: (a) contacting a double-stranded deoxyribonucleic acid polynucleotide with an RNA guided archaeal endonuclease in a complex with an endonuclease and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide; (b) the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the endonuclease is selected from the group consisting of SEQ ID NOs: 674-675, 676-677, 678-679, 680-681, 682-683, 684-685, 686-687, 688-689, 690-691, 700-701, 702-703, 704-705, 706-707, 708-710, 712-713, 714-715, 716-717, 718-719, 720-721, 722-723, 724-725, 726-727, 728-729, 730-730, 731-732, 732-733, 733-734, 735-736, 737-738, 739-740, 741-742, 743-744, 745-746, 747-748, 748-749, 750-751, 752-753, 754-75 1002, 1260-1321, or a variant having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of the following: 1002, 1260-1321. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide and the PAM comprises NGG. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide 6 to 9 or 7 nucleotides from the PAM. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the endonuclease is from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, or bacterial double-stranded deoxyribonucleic acid polynucleotide from a species other than the species from which the endonuclease was derived.
[0020] In some aspects, the disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering any of the engineered nuclease systems described herein to the target nucleic acid locus, the endonuclease configured to form a complex with an engineered guide ribonucleic acid structure, the complex configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid locus comprises genomic eukaryotic DNA, archaeal DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid comprises bacterial DNA, the bacterial DNA is from a bacterial or archaeal species different from the species from which the endonuclease was derived. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the endonuclease and the engineered guide nucleic acid structure are encoded by separate nucleic acid molecules. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, an archaeal cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is from a species different from the species from which the endonuclease was derived. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding the endonuclease is operably linked.In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide ribonucleic acid operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus. In some embodiments, the endonuclease induces a double-stranded break proximal to the target locus 5' from the protospacer adjacent motif (PAM). In some embodiments, the endonuclease induces a double-stranded break 6-8 nucleotides or 7 nucleotides 5' from the PAM. In some embodiments, the engineered nuclease system induces a chemical modification of a nucleotide base within or proximal to the target locus. In some embodiments, the chemical modification is deamination of an adenosine or cytosine nucleotide. In some embodiments, the endonuclease further comprises a base editor coupled to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.
[0021] In some aspects, the disclosure provides a method of disrupting a TRAC locus in a cell, comprising contacting the cell with a composition, wherein the composition (a) identifies at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof, or a variant thereof. and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the locus, and wherein the engineered guide RNA is configured to hybridize to any one of SEQ ID NOs: 1079-1082, 1145-1166, and 1169-1170. In some embodiments, the engineered guide RNA comprises a sequence having at least about at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1123-1144 or 1167-1168. In some embodiments, the engineered guide RNA comprises modified nucleotides of any one of SEQ ID NOs: 1123-1144 or 1167-1168.In some embodiments, the engineered guide RNA comprises a sequence having at least about at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence complementary to any one of SEQ ID NOs: 1145-1166 or 1169-1170. In some embodiments, the endonuclease has at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 582, 988, 990, 993, 996, 999, or 1002. In some embodiments, the region is 5' to a protospacer adjacent motif (PAM) comprising any one of SEQ ID NOs: 1023-1044.
[0022] In some aspects, the disclosure provides an isolated RNA molecule comprising a sequence at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1123-1144 or 1167-1168. In some embodiments, the isolated RNA molecule further comprises a pattern of chemical modifications listed in any one of SEQ ID NOs: 1123-1144 or 1167-1168.
[0023] In some aspects, the disclosure provides for the use of any of the isolated RNA molecules described herein to modify the TRAC locus of a cell.
[0024] In some aspects, the disclosure provides a method of disrupting the AAVS1 locus in a cell, comprising contacting the cell with a composition, wherein the composition is: (a) a nucleic acid sequence that is at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, or a variant thereof; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the locus, and wherein the engineered guide RNA is configured to hybridize to any one of SEQ ID NOs: 1105-1122. In some embodiments, the engineered guide RNA comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1087-1104. In some embodiments, the engineered guide RNA comprises a modified nucleotide of any one of SEQ ID NOs: 1087-1104. In some embodiments, the engineered guide RNA comprises a sequence having at least about 80% identity to a sequence complementary to any one of SEQ ID NOs: 1105-1122.In some embodiments, the endonuclease has at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 582, 988, 990, 993, 996, 999, or 1002. In some embodiments, the endonuclease has at least about 75%, 80%, or 90% sequence identity to SEQ ID NO: 582. In some embodiments, the region is 5' to a protospacer adjacent motif (PAM) comprising any one of SEQ ID NOs: 1023-1044.
[0025] In some aspects, the disclosure provides an isolated RNA molecule comprising a sequence at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1087-1104. In some embodiments, the RNA molecule comprises a pattern of chemical modifications listed in any one of SEQ ID NOs: 1087-1104.
[0026] In some aspects, the disclosure provides an engineered nuclease system, comprising: (a) an endonuclease comprising a RuvC domain and an HNH domain, the endonuclease having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 122%, at least 124%, at least 126%, at least 128%, at least 128, at least 126, at least 128 ... %, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to said endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with said endonuclease, and (ii) an engineered guide ribonucleic acid structure comprising a ribonucleic acid sequence configured to bind to an endonuclease, wherein the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 677-681, 686, 1006-1008, 1011-1014, or 1231-1259. In some embodiments, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide comprising a guide ribonucleic acid sequence and a ribonucleic acid sequence configured to bind to an endonuclease, in some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence.In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from any one of SEQ ID NOs: 205-220. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the endonuclease and the ribonucleic acid sequence configured to bind to the endonuclease are derived from distinct species within the same phylum. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the endonuclease does not exhibit concomitant ssDNA cleavage activity.
[0027] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease comprises a sequence having at least 80% sequence identity to any one of the endonuclease effector sequences described herein, or a variant thereof; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease, wherein the endonuclease comprises a sequence having at least 80% sequence identity to any non-degenerate nucleotide of any of the sgRNA sequences described herein, or a variant thereof.
[0028] In some aspects, the disclosure provides isolated RNA molecules comprising a sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the non-degenerate nucleotides of any of the sgRNA sequences described herein.
[0029] In some aspects, the disclosure provides a nucleic acid comprising any of the sequences described herein.
[0030] In some aspects, the disclosure provides vectors comprising any of the nucleic acid sequences described herein.
[0031] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC domain and an HNH domain, the endonuclease being derived from an uncultured microorganism; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease, the endonuclease having a molecular weight of about 96 kDa or less. In some embodiments, the endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease is a class 2, type II Cas endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the arginine-rich region or the domain having PF14239 homology has at least 85%, at least 90%, or at least 95% identity to an arginine-rich region or a domain having PF14239 homology of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some embodiments, the endonuclease further comprises a REC (recognition) domain.In some embodiments, the REC domain has at least 85%, at least 90%, or at least 95% identity to the REC domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some embodiments, the endonuclease further comprises a BH (bridge helix) domain, a WED (wedge) domain, and a PI (PAM interacting) domain. In some embodiments, the BH domain, the WED domain, or the PI domain has at least 85%, at least 90%, or at least 95% identity to the BH domain, the WED domain, or the PI domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof.
[0032] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC-I domain and an HNH domain; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease, wherein the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some embodiments, the endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease is a class 2, type II Cas endonuclease. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the arginine-rich region or the domain having PF14239 homology has at least 85%, at least 90%, or at least 95% identity to an arginine-rich region of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the REC domain has at least 85%, at least 90%, or at least 95% identity to the REC domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain.In some embodiments, the BH domain, the WED domain, or the PI domain has at least 85%, at least 90%, or at least 95% identity to the BH domain, WED domain, or PI domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or variants thereof. In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or a sequence having at least 80% sequence identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259. In some embodiments, the guide nucleic acid structure comprises a sequence having at least 80% identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259.
[0033] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to an endonuclease, the ribonucleic acid sequence comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or a sequence having at least 80% sequence identity to a non-variable nucleotide of any one of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259; and (b) an RNA-guided endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the RNA-guided endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide comprising the guide ribonucleic acid sequence and the tracr ribonucleic acid sequence. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 205-220.In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first or second homologous arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the system comprises Mg. 2+In some embodiments, the endonuclease and the tracr ribonucleic acid sequence are derived from distinct bacterial species within the same phylum. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 2-24, and the guide RNA structure comprises an RNA sequence predicted to comprise a hairpin comprising a stem and a loop, the stem comprising at least 12 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises a second stem and a second loop, the second stem comprising at least 5 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises an RNA structure comprising at least two hairpins. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity to SEQ ID NO: 1, and the guide RNA structure comprises an RNA sequence predicted to comprise at least four hairpins comprising stems and loops. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 1, 2, 10, 17, or 613-616; and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 199-200, or 669-673, or any one of SEQ ID NOs: 201-203, 613-616, non-degenerate nucleotides. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 1-24, 462-488, or 501-612; and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to the non-degenerate nucleotides of any one of SEQ ID NOs: 199-200 or 669-673 or any one of SEQ ID NOs: 201-203, 613-616.In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 2, 10, or 17, and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to the non-degenerate nucleotides of any one of SEQ ID NOs: 202-203 or 613-614. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 25-198, 221-459, or 489-580, and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to a class 2, type II sgRNA, or tracr sequence. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithm parameters using CLUSTALW, hi some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and a BLOSUM62 scoring matrix setting gap costs at 11, presence, and extension of 1, and using a conditional composition score matrix adjustment. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease.
[0034] In some aspects, the disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, and the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease comprising any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof having at least 75% sequence identity. In some embodiments, the DNA-targeting segment is located 5' to both of the two complementary stretches of nucleotides. In some embodiments, a) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to any one of SEQ ID NOs: 199-200 or 669-673, and b) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 201-203 or 613-616. In some embodiments, a) the endonuclease comprises a sequence having at least 70%, at least 80%, or at least 90% identity to any one of SEQ ID NOs: 2, 10, or 17, and b) the guide RNA structure comprises a sequence having at least 70%, at least 80%, or at least 90% identity to at least one of the non-degenerate nucleotides of SEQ ID NO: 200 or SEQ ID NOs: 202-203 or 613-614.In some embodiments, a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 25-198, 221-459, or 489-580, and b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to a class 2, type II sgRNA. In some embodiments, the endonuclease further comprises a base editor or a histone editor coupled to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.
[0035] In some aspects, the disclosure provides a deoxyribonucleic acid polynucleotide encoding any of the engineered guide ribonucleic acid polynucleotides described herein.
[0036] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding a class 2, type II Cas endonuclease comprising a RuvC domain and an HNH domain, the endonuclease being derived from an uncultured microorganism, the endonuclease having a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, 60 kDa or less, or 30 kDa or less. In some embodiments, the endonuclease comprises SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof having at least 70% sequence identity thereto. In some embodiments, the endonuclease further comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 205-220. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human. In some embodiments, the organism is a prokaryote or a bacterium, and the organism is a different organism from the organism from which the endonuclease is derived. In some embodiments, the organism is not the uncultured microorganism.
[0037] In some aspects, the present disclosure provides a vector comprising a nucleic acid sequence encoding an RNA-guided endonuclease comprising a RuvC-I domain and an HNH domain, the endonuclease being derived from an uncultured microorganism, the endonuclease having a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less, and the RNA-guided endonuclease is optionally archaeal. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain.
[0038] In some aspects, the disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure further comprising a) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and b) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
[0039] In some aspects, the disclosure provides a cell comprising any of the vectors described herein. In some embodiments, the cell is a bacterial, archaeal, fungal, eukaryotic, mammalian, or plant cell. In some embodiments, the cell is a bacterial cell.
[0040] In some aspects, the disclosure provides a method of producing an endonuclease comprising culturing any of the cells described herein.
[0041] In some aspects, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with a class 2, type II Cas endonuclease in a complex with the endonuclease and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide; and (b) the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide, and the PAM comprises NGG. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide 6-8 nucleotides or 7 nucleotides from the PAM. In some embodiments, the endonuclease comprises any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to such a variant.
[0042] In some aspects, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with an RNA guided archaeal endonuclease in a complex with the endonuclease and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the endonuclease comprises any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof having at least 70%, at least 75%, at least 80%, at least 90% sequence identity to such a variant. In some embodiments, the endonuclease cleaves the double stranded deoxyribonucleic acid polynucleotide and the PAM comprises NGG. In some embodiments, the endonuclease cleaves the double stranded deoxyribonucleic acid polynucleotide 6 to 8 or 7 nucleotides from the PAM. In some embodiments, the Class 2, Type II Cas endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the Class 2, Type II Cas endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide, hi some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, or bacterial double-stranded deoxyribonucleic acid polynucleotide from a species other than the species from which the endonuclease was derived.
[0043] In some aspects, the disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering any of the engineered nuclease systems described herein to the target nucleic acid locus, the endonuclease configured to form a complex with the engineered guide ribonucleic acid structure, the complex configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid locus comprises genomic eukaryotic DNA, archaeal DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid comprises bacterial DNA, and the bacterial DNA is from a bacterial or archaeal species different from the species from which the endonuclease was derived. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the endonuclease and the engineered guide nucleic acid structure are encoded by separate nucleic acid molecules. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, an archaeal cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is from a species different from the species from which the endonuclease was derived. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked.In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus. In some embodiments, the endonuclease induces a double-stranded break proximal to the target locus 5' from a protospacer adjacent motif (PAM). In some embodiments, the endonuclease induces a double-stranded break 6-8 nucleotides or 7 nucleotides 5' from the PAM. In some embodiments, the engineered nuclease system induces a chemical modification of a nucleotide base at or near the target locus, or a chemical modification of a histone at or near the target locus. In some embodiments, the chemical modification is deamination of an adenosine or cytosine nucleotide. In some embodiments, the endonuclease further comprises a base editor coupled to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.
[0044] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0045] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief description of the drawings]
[0046] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (hereinafter "Figure" and "FIG."). [Figure 1A] Illustrated are dendrograms showing the homology relationships of different classes and types of CRISPR / Cas loci. The SMART I and II Cas enzyme classes described herein are shown in comparison to class 2, type II-A, type II-B, and type II-C Cas systems, demonstrating that these systems are grouped into separate classes from II-A, II-B, and II-C. (Figure 1A) shows the SMART phylogenetic tree in the context of the Cas9 reference sequence, with SMART effectors clustered far away from the Cas9 reference sequence (types II-A, II-B, and II-C), and (Figure 1B) shows the SMART phylogenetic tree illustrating subgroups of SMART enzymes. [Figure 1B]Illustrated are dendrograms showing the homology relationships of different classes and types of CRISPR / Cas loci. The SMART I and II Cas enzyme classes described herein are shown in comparison to class 2, type II-A, type II-B, and type II-C Cas systems, demonstrating that these systems are grouped into separate classes from II-A, II-B, and II-C. (Figure 1A) shows the SMART phylogenetic tree in the context of the Cas9 reference sequence, with SMART effectors clustered far away from the Cas9 reference sequence (types II-A, II-B, and II-C), and (Figure 1B) shows the SMART phylogenetic tree illustrating subgroups of SMART enzymes. [Diagram 2] 1 shows the length distribution of SMART effectors described herein, showing that SMART I and II enzymes are clustered at lower molecular weights than Cas9-like enzymes. SMART nucleases show a bimodal distribution with one peak at about 400 aa (SMART II) and a second peak at about 750 aa (SMART I). Cas9 nucleases also show a bimodal distribution with peaks at about 1,100 aa (e.g., SaCas9) and 1,300 aa (e.g., SpCas9). [Figure 3A]Illustrates the genomic context of the "small" type II nucleases MG33-1, MG35-236. SMART nucleases and CRISPR accessory proteins are indicated with dark grey arrows, other genes are illustrated with light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. The following are shown: (FIG. 3A) Genomic context of the SMART I MG33-1 nuclease encoded upstream of the SMART II nuclease MG35-236, showing the predicted insertion sequence carrying the transposases TnpA and TnpB downstream of SMART II, and the CRISPR locus; (FIG. 3B) Genomic context of the SMART I nuclease MG34-1, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3C) Genomic context of the SMART I nuclease MG34-16, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3D) Genomic fragments were identified as phage-derived based on the virus-specific gene annotation terminase and portal, and MG34-16 in (FIG. 3D). Genomic fragment targeted by spacer 7 from the CRISPR array, inset shows location of MG34-16 spacer 7 targeting the C-terminus of a viral gene of unknown function, the putative NGG PAM of MG34-16 is highlighted by a grey box downstream of the spacer match. [Figure 3B]Illustrates the genomic context of the "small" type II nucleases MG33-1, MG35-236. SMART nucleases and CRISPR accessory proteins are indicated with dark grey arrows, other genes are illustrated with light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. The following are shown: (FIG. 3A) Genomic context of the SMART I MG33-1 nuclease encoded upstream of the SMART II nuclease MG35-236, showing the predicted insertion sequence carrying the transposases TnpA and TnpB downstream of SMART II, and the CRISPR locus; (FIG. 3B) Genomic context of the SMART I nuclease MG34-1, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3C) Genomic context of the SMART I nuclease MG34-16, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3D) Genomic fragments were identified as phage-derived based on the virus-specific gene annotation terminase and portal, and MG34-16 in (FIG. 3D). Genomic fragment targeted by spacer 7 from the CRISPR array, inset shows location of MG34-16 spacer 7 targeting the C-terminus of a viral gene of unknown function, the putative NGG PAM of MG34-16 is highlighted by a grey box downstream of the spacer match. [Figure 3C]Illustrates the genomic context of the "small" type II nucleases MG33-1, MG35-236. SMART nucleases and CRISPR accessory proteins are indicated with dark grey arrows, other genes are illustrated with light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. The following are shown: (FIG. 3A) Genomic context of the SMART I MG33-1 nuclease encoded upstream of the SMART II nuclease MG35-236, showing the predicted insertion sequence carrying the transposases TnpA and TnpB downstream of SMART II, and the CRISPR locus; (FIG. 3B) Genomic context of the SMART I nuclease MG34-1, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3C) Genomic context of the SMART I nuclease MG34-16, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3D) Genomic fragments were identified as phage-derived based on the virus-specific gene annotation terminase and portal, and MG34-16 in (FIG. 3D). Genomic fragment targeted by spacer 7 from the CRISPR array, inset shows location of MG34-16 spacer 7 targeting the C-terminus of a viral gene of unknown function, the putative NGG PAM of MG34-16 is highlighted by a grey box downstream of the spacer match. [Figure 3D]Illustrates the genomic context of the "small" type II nucleases MG33-1, MG35-236. SMART nucleases and CRISPR accessory proteins are indicated with dark grey arrows, other genes are illustrated with light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. The following are shown: (FIG. 3A) Genomic context of the SMART I MG33-1 nuclease encoded upstream of the SMART II nuclease MG35-236, showing the predicted insertion sequence carrying the transposases TnpA and TnpB downstream of SMART II, and the CRISPR locus; (FIG. 3B) Genomic context of the SMART I nuclease MG34-1, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3C) Genomic context of the SMART I nuclease MG34-16, where environmental expression sequencing reads are shown aligned under the CRISPR array and predicted tracrRNA, and the transcriptome coverage of the region is illustrated above the contig sequence; (FIG. 3D) Genomic fragments were identified as phage-derived based on the virus-specific gene annotation terminase and portal, and MG34-16 in (FIG. 3D). Genomic fragment targeted by spacer 7 from the CRISPR array, inset shows location of MG34-16 spacer 7 targeting the C-terminus of a viral gene of unknown function, the putative NGG PAM of MG34-16 is highlighted by a grey box downstream of the spacer match. [Figure 4A]1 shows a multiple sequence alignment of exemplary SMART endonucleases (MG33-1 (SEQ ID NO:1), MG33-2 (SEQ ID NO:463), MG33-3 (SEQ ID NO:464), MG34-1 (SEQ ID NO:2), MG34-9 (SEQ ID NO:10), MG34-16 (SEQ ID NO:17), MG102-1 (SEQ ID NO:581), MG102-2 (SEQ ID NO:582), MG35-1 (SEQ ID NO:25), MG35-2 (SEQ ID NO:26), MG35-3 (SEQ ID NO:27), MG35-102 (SEQ ID NO:126), MG35-236 (SEQ ID NO:284), MG35-419 (SEQ ID NO:222), MG35-420 (SEQ ID NO:223), and MG35-421 (SEQ ID NO:224)). The sequence of SaCas9 used as the reference domain is shown as a rectangle below the reference sequence and catalytic residues are shown as squares above each sequence. Shown is (Figure 4A) an alignment of the endonuclease region containing the RuvC-I and bridge helix domains, (Figure 4B) an alignment of the region containing the RuvC-III domain, and (Figure 4C) an alignment of the region containing the RuvC-II and HNH domains. [Figure 4B]1 shows a multiple sequence alignment of exemplary SMART endonucleases (MG33-1 (SEQ ID NO:1), MG33-2 (SEQ ID NO:463), MG33-3 (SEQ ID NO:464), MG34-1 (SEQ ID NO:2), MG34-9 (SEQ ID NO:10), MG34-16 (SEQ ID NO:17), MG102-1 (SEQ ID NO:581), MG102-2 (SEQ ID NO:582), MG35-1 (SEQ ID NO:25), MG35-2 (SEQ ID NO:26), MG35-3 (SEQ ID NO:27), MG35-102 (SEQ ID NO:126), MG35-236 (SEQ ID NO:284), MG35-419 (SEQ ID NO:222), MG35-420 (SEQ ID NO:223), and MG35-421 (SEQ ID NO:224)). The sequence of SaCas9 used as the reference domain is shown as a rectangle below the reference sequence and catalytic residues are shown as squares above each sequence. Shown is (Figure 4A) an alignment of the endonuclease region containing the RuvC-I and bridge helix domains, (Figure 4B) an alignment of the region containing the RuvC-III domain, and (Figure 4C) an alignment of the region containing the RuvC-II and HNH domains. [Figure 4C]1 shows a multiple sequence alignment of exemplary SMART endonucleases (MG33-1 (SEQ ID NO:1), MG33-2 (SEQ ID NO:463), MG33-3 (SEQ ID NO:464), MG34-1 (SEQ ID NO:2), MG34-9 (SEQ ID NO:10), MG34-16 (SEQ ID NO:17), MG102-1 (SEQ ID NO:581), MG102-2 (SEQ ID NO:582), MG35-1 (SEQ ID NO:25), MG35-2 (SEQ ID NO:26), MG35-3 (SEQ ID NO:27), MG35-102 (SEQ ID NO:126), MG35-236 (SEQ ID NO:284), MG35-419 (SEQ ID NO:222), MG35-420 (SEQ ID NO:223), and MG35-421 (SEQ ID NO:224)). The sequence of SaCas9 used as the reference domain is shown as a rectangle below the reference sequence and catalytic residues are shown as squares above each sequence. Shown is (Figure 4A) an alignment of the endonuclease region containing the RuvC-I and bridge helix domains, (Figure 4B) an alignment of the region containing the RuvC-III domain, and (Figure 4C) an alignment of the region containing the RuvC-II and HNH domains. [Diagram 5] 5 illustrates an exemplary domain organization of a SMART I endonuclease, using MG34-1 as an example. (FIG. 5A) A diagram showing the predicted domain architecture of a SMART I nuclease, including three RuvC domains, a bridge helix ("BH"), a domain with homology to Pfam PF14239 interrupting the recognition domain ("REC"), an HNH endonuclease domain ("HNH"), a wedge domain ("WED"), and a PAM interaction domain (PI), and (FIG. 5B) an outline of a multiple sequence alignment of two SMART I nucleases against a reference Cas9 nuclease sequence, with RuvC and HNH catalytic residues shown as black bars above each sequence, regions that align in 3D space with the crystal structure of SaCas represented by round boxes, and dashed lines representing regions of poor or no alignment in 3D space between the SMART and SaCas9 3D structure predictions. [Figure 6]6 illustrates an exemplary domain organization of SMART II endonucleases, using MG35 family enzymes (MG35-3, MG35-4) as an example. (FIG. 6A) A diagram showing the predicted domain architecture of SMART II nucleases, including three RuvC domains, a domain with homology to Pfam PF14239, an HNH endonuclease domain, an unknown domain, and a recognition domain (REC), and (FIG. 6B) an overview of a multiple sequence alignment of two SMART II nucleases against a reference Cas9 nuclease sequence, where RuvC and HNH catalytic residues are shown as black bars above each sequence, regions that align with the crystal structure of SaCas in 3D space are represented by round boxes, and residues identified from the 3D structure prediction that may be involved in recognition of guide / target / PAM sequences are represented by dark gray boxes above the MG35-419 sequence (within the RRXRR and REC domains). [Figure 7A] Various features of SMART enzymes are illustrated in (FIG. 7A) a dot plot showing the identity of the SMART I domains of various enzymes depicted herein with the SMART I domain of spCas9, showing that they share up to about 35% sequence identity, and (FIG. 7B) a dot plot of the length of the individual SMART I domains of the enzymes described herein. [Figure 7B] Various features of SMART enzymes are illustrated in (FIG. 7A) a dot plot showing the identity of the SMART I domains of various enzymes depicted herein with the SMART I domain of spCas9, showing that they share up to about 35% sequence identity, and (FIG. 7B) a dot plot of the length of the individual SMART I domains of the enzymes described herein. [Figure 8A]Illustrated are the count distributions of predicted motifs in various SMART-specific motifs versus Cas9 nuclease sequences where these motifs occur more commonly in SMART enzymes. Motifs were predicted in 803 reference Cas9 sequences (types II-A, II-B, and II-C), 84 SMART I sequences, and 471 SMART II sequences. (Figure 8A) Box plots of count frequencies of Zn-binding ribbon motifs (CX[2-4]C and CX[2-4]H) in various types of class 2 Cas enzymes, and (B) histograms of count frequencies of RRXRR motifs in various types of class 2 Cas enzymes are shown. In (Figure 8A) and (Figure 8B), the lines follow the average counts, while outliers are represented by dots. [Figure 8B] Illustrated are the count distributions of predicted motifs in various SMART-specific motifs versus Cas9 nuclease sequences where these motifs occur more commonly in SMART enzymes. Motifs were predicted in 803 reference Cas9 sequences (types II-A, II-B, and II-C), 84 SMART I sequences, and 471 SMART II sequences. (Figure 8A) Box plots of count frequencies of Zn-binding ribbon motifs (CX[2-4]C and CX[2-4]H) in various types of class 2 Cas enzymes, and (B) histograms of count frequencies of RRXRR motifs in various types of class 2 Cas enzymes are shown. In (Figure 8A) and (Figure 8B), the lines follow the average counts, while outliers are represented by dots. [Figure 9A] 9A-9D illustrate predicted guide RNA structures of single-stranded RNAs (sgRNAs) designed for cleavage activity with SMART I endonucleases. (FIG. 9A) MG34-1 sgRNA1, (FIG. 9B) MG34-1 sgRNA2, (FIG. 9C) MG34-9 sgRNA1, and (FIG. 9D) MG34-16 sgRNA1 are shown. [Figure 9B]9A-9D illustrate predicted guide RNA structures of single-stranded RNAs (sgRNAs) designed for cleavage activity with SMART I endonucleases. (FIG. 9A) MG34-1 sgRNA1, (FIG. 9B) MG34-1 sgRNA2, (FIG. 9C) MG34-9 sgRNA1, and (FIG. 9D) MG34-16 sgRNA1 are shown. [Figure 9C] 9A-9D illustrate predicted guide RNA structures of single-stranded RNAs (sgRNAs) designed for cleavage activity with SMART I endonucleases. (FIG. 9A) MG34-1 sgRNA1, (FIG. 9B) MG34-1 sgRNA2, (FIG. 9C) MG34-9 sgRNA1, and (FIG. 9D) MG34-16 sgRNA1 are shown. [Figure 9D] 9A-9D illustrate predicted guide RNA structures of single-stranded RNAs (sgRNAs) designed for cleavage activity with SMART I endonucleases. (FIG. 9A) MG34-1 sgRNA1, (FIG. 9B) MG34-1 sgRNA2, (FIG. 9C) MG34-9 sgRNA1, and (FIG. 9D) MG34-16 sgRNA1 are shown. [Figure 10A] Illustrating the cleavage properties of SMART I nuclease as described in Example 1. (FIG. 10A) shows an Agilent TapeStation gel of ligation products of a cleavage assay of MG34-1 with two sgRNA designs versus a negative control. Lane L3: ladder. Lane A4: Apo, no sgRNA. Lanes B4 and C4: MG34-1 sgRNAs tested (sg1: SEQ ID NO: 612, sg2: 613). Cleavage product bands are labeled with arrows. Lanes G3 and H3: shown in grey and not relevant to this experiment. [Figure 10B] Illustrates the cleavage properties of SMART I nuclease as described in Example 1. (FIG. 10B) shows a PCR gel of ligation products showing activity of MG34-1, 34-9 and 34-16. Lane 1: ladder. Lanes 2-7: sgRNA designs with six spacer lengths of MG34-1. Lanes 8 and 9: sgRNA designs of 34-9 and 34-16, respectively. Arrows indicate cleavage confirmation bands. [Figure 11A]Illustrating sequence cleavage preferences for MG34 nuclease. (Figure 11A) shows the SeqLogo representation of the consensus PAM sequence (NGGN) for MG34-1 using sgRNA1 (top, SEQ ID NO: 612) and sgRNA2 (bottom, SEQ ID NO: 613). [Figure 11B] Illustrating sequence cleavage preferences for MG34 nuclease (FIG. 11B) shows a histogram depicting the location of the cleavage site for MG34-1, demonstrating that MG34-1 prefers to cleave at approximately position 7 from the PAM. [Figure 11C] Illustrating sequence cleavage preference for MG34 nuclease. (FIG. 11C) shows a Sanger sequencing chromatogram showing the preferred NGG PAM of MG34-9 (highlighted in a box). The arrow indicates the cleavage site 7 positions from the PAM. [Figure 12A] Illustrates the results of a plasmid targeting experiment in E. coli for MG34-1. (FIG. 12A) shows replica plating of E. coli strains demonstrating plasmid cleavage; E. coli expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid (+sp) containing the sgRNA target. Plate quadrants showing growth defect (+sp) versus negative control (no target and PAM (-sp)) indicate successful targeting and cleavage by the enzyme. Experiments were repeated twice and performed in triplicate. [Figure 12B] 12 illustrates the results of a plasmid targeting experiment in E. coli for MG34-1. (FIG. 12B) A graph of colony forming unit (cfu) measurements from the replica plating experiment in A is shown, showing growth inhibition in the targeting condition (+sp) versus the non-targeting control (-sp), demonstrating that the plasmid was cleaved. (FIG. 12C) shows a bar plot of colony forming unit (cfu) measurements (log scale) showing E. coli growth inhibition in the targeting condition (white bars) versus the non-targeting control (green bars). [Figure 12C]Illustrating the results of plasmid targeting experiments in E. coli against MG34-1. (FIG. 12C) shows a bar plot of colony forming unit (cfu) measurements (logarithmic scale) showing E. coli growth inhibition in the targeting condition (white bars) versus the non-targeting control (green bars). Plasmid interference assays for each nuclease were performed in triplicate with a SpCas9 positive control. [Figure 13A] 13A shows an exemplary genomic context of the SMART system for MG35-419. SMART nucleases are indicated by dark grey arrows, and other genes are illustrated by light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. Environmental expression sequencing reads are shown aligned below the CRISPR array (FIG. 13A) and upstream of the effector (FIG. 13B). Transcriptome coverage of regions showing expression is illustrated above the contig sequence. (FIG. 13A) shows the genomic context of the nearby encoded SMART II MG35-419 effector and CRISPR loci. [Figure 13B] Figure 13 shows an exemplary genomic context of the SMART system for MG35-419. SMART nucleases are indicated with dark grey arrows, and other genes are illustrated with light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. Environmental expression sequencing reads are shown aligned below the CRISPR array (Figure 13A) and upstream of the effector (Figure 13B). Transcriptome coverage of regions showing expression is illustrated above the contig sequence. (Figure 13B) shows the genomic context of the SMART II effector MG35-3 showing the transcribed 5'UTR. [Figure 14]3D structure prediction of SMART II MG35-419 is shown. This 3D model aligns well with regions of the SaCas9 crystal structure, despite being less than half its size. Regions aligned with the SaCas9 template include the catalytic lobe (RuvC-I, HNH, and RuvC-III domains) and a short region of the recognition (REC) lobe. SMART II specific domains include a domain containing an RRXRR motif and homology to Pfam PF14239, and a domain of unknown function. [Figure 15] Illustrates the results of a preliminary cleavage assay for SMART II effectors. MG35-420 (SEQ ID NO: 223) protein preparation was tested for cleavage activity in TXTL extracts where the entire locus was expressed. The experiment incubated the protein preparation with the PAM library (dsDNA target), the predicted repeat region at the locus (cr1) in both forward and reverse orientations (fw and rv), and an intergenic region potentially encoding an associated cofactor. Lanes 2-9 (no cr array): control experiment without repeat region. Apo: only protein preparation with targeted PAM library. Labels 1-2.5 represent the seven different intergenic regions. -IG: no intergenic region was included as a control. PCR gel of ligation products shows putative cleavage bands (arrows) suggestive of dsDNA cleavage. [Figure 16A] Figure 16A illustrates the genomic context of the SMART system. SMART nuclease is indicated by a dark gray arrow, and other genes are illustrated by light gray arrows. Predicted domains for all genes in the genomic fragment are shown as gray boxes under the arrows. Environmental expression sequencing reads are shown aligned upstream of the effector. Figure 16B illustrates the genomic context of the SMART II MG35-419 effector. [Figure 16B]Figure 16B illustrates the genomic context of the SMART system. SMART nucleases are indicated by dark grey arrows, and other genes are illustrated by light grey arrows. Predicted domains for all genes in the genomic fragment are shown as grey boxes below the arrows. Environmental expression sequencing reads are shown aligned upstream of the effector. Figure 16B illustrates the genomic context of the SMART II MG35-102 effector. [Figure 17A] Figure 17A illustrates the data demonstrating that MG35-420 is an active dsDNA nuclease. Figure 17B illustrates the genomic context of the MG34-420 effector. The effector is represented by a dark arrow in the reverse orientation, the predicted PFAM domain is represented by a rectangle below the arrow, and the intergenic region that may code for a guide RNA is annotated as "IG" on the black line. A CRISPR-like repeat region is present in the contig. [Figure 17B] 17A-17C illustrate data demonstrating that MG35-420 is an active dsDNA nuclease. FIG. 17B illustrates the results of purified protein preparations tested for cleavage activity in TXTL. The experiment incubated protein preparations with the PAM library (dsDNA target), the predicted CRISPR-like repeat region at the locus (cr1) in both forward and reverse orientations (fw and rv), and an intergenic region potentially encoding an associated cofactor. Lanes 2-9 (no cr array): control experiment without repeat region. Apo: only protein preparation with targeted PAM library. Labels 1-2.5 represent seven different intergenic regions. -IG: no intergenic region was included as a control. PCR gel of ligation products shows putative cleavage bands (arrows) suggestive of dsDNA cleavage. The band recovered on the lane labeled "4" represents the cleavage band from incubating the enzyme with the CRISPR-like region and the SMART II 5'UTR. [Figure 18]18A illustrates the genomic context of the MG34-420 effector, showing RNASeq reads sequenced from an in vitro transcription reaction of the SMART II effector with its 5'UTR. The effector is represented by a dark arrow in the reverse orientation, the predicted PFAM domain is represented by a rectangle below the arrow, and the predicted guide RNA is annotated on the black line. FIG. 18B illustrates the secondary structure representation of the SMART II MG35-420 putative guide RNA. [Figure 19A] Figure 19 illustrates multiple sequence alignments (MSAs) of conserved UTR regions associated with SMART II effectors. Figure 19A illustrates the full-length MSA of the region immediately upstream of the start codon of SMART II effectors. The percent identity histogram above the alignment indicates the conserved region (annotated as the 5'UTR guide RNA, gray arrow). [Figure 19B] Figure 19B illustrates the multiple sequence alignment (MSA) of conserved UTR regions associated with SMART II effectors. Figure 19B illustrates the highly conserved regions within the putative guide RNA coding sequence. Percent identity histograms and sequence logo representations are shown above the alignment. Identical bases are highlighted with black boxes. [Figure 20A] Figure 20A illustrates data demonstrating that the MG35 effector is an active dsDNA nuclease using sgRNA. Figure 20A illustrates the results of an in vitro cleavage assay. Effectors containing (sg) and without (Apo)sgRNA were assayed in in vitro transcription / translation reactions incubated with PAM library (dsDNA target). Cleavage products were amplified via PCR (successful RNA-guided cleavage with nuclease-generated band of expected size, arrow). [Figure 20B] Figure 20B illustrates data demonstrating that the MG35 effector is an active dsDNA nuclease that uses sgRNA. Figure 20B illustrates the target adjacent motif (TAM). [Figure 21-1]Figure 21 illustrates data demonstrating that SMART enzymes are novel nucleases with diverse targeting capabilities. Figure 21A illustrates the predicted domain architecture of SMART nucleases versus SpCas9. [Figure 21-2] Figure 21 illustrates data demonstrating that SMART enzymes are novel nucleases with diverse targeting capabilities. Figure 21B illustrates the genomic context of the SMART MG102-2 system. The orientation of tracrRNA and CRISPR arrays was confirmed by in vitro cleavage activity with effectors. [Figure 21-3] Data are illustrated demonstrating that SMART enzymes are novel nucleases with diverse targeting capabilities. Figure 21C illustrates the genomic context of the SMART MG34-1 line. Adaptive module genes (Cas1, Cas2, Cas4, and putative Csn2) were identified. Environmental RNASeq reads were forward-oriented mapped to the array and the intergenic region encoding the tracrRNA. Other genes encoded in the locus are represented by yellow arrows. The orientation of the tracrRNA and CRISPR array was confirmed by in vitro cleavage activity with effectors. [Figure 21-4] Illustrates data demonstrating that SMART enzymes are novel nucleases with diverse targeting capabilities. Figure 21D illustrates the HEARO RNA secondary structures of two active SMART HEARO nucleases. Shows SeqLogo representation of consensus target motif sequences. [Figure 21-5]The data demonstrate that SMART enzymes are novel nucleases with diverse targeting capabilities. Figure 21E illustrates the protein phylogenetic tree of SMART nucleases versus Cas9 and IscB reference sequences. SMART effector and archaeal Cas9 sequences (cyan and purple branches) are distantly related to documented Cas9 reference sequences (types II-A, II-B, and II-C, gray branches). Phylogenetic trees were inferred from multiple sequence alignments of the shared RuvC-II / HNH / RuvC-III domains. The SMART MG33 family of nucleases (magenta branch) clusters with CRISPR type II-C variant systems, while other CRISPR-associated SMART nucleases (cyan branch) cluster with sequences recently classified as type II-D. SMART HEARO nucleases (light purple branch) cluster with HEARO ORFs and IscB sequences. Figure 21F illustrates phylogenetic clades of SMART CRISPR type II families. The clades are expanded representations of the phylogenetic tree illustrated in Figure 21E. Local support values for internal family split nodes are shown and range from 0 to 1. A SeqLogo representation of the consensus target motif sequences and sgRNA designs from biochemical cleavage activity assays for active SMART nucleases is shown. [Figure 22A] Figure 22 illustrates data demonstrating that SMART I is a dsDNA nuclease. Figure 22A illustrates a histogram of cleavage site preference showing that MG34-1 preferentially cleaves dsDNA at position 7 from the PAM. The inset shows that MG34-1 generates staggered cuts, with the cut at position 3 occurring on the targeted strand (TS) and the cut at positions 6-7 occurring on the non-targeted strand (NTS). [Figure 22B] Figure 22B illustrates data demonstrating that SMART I is a dsDNA nuclease. Figure 22B illustrates the distribution of percent DNA cleavage with various spacer lengths, showing a preference for the 18 bp spacer for MG34-1. [Figure 22C]Figure 22C illustrates data demonstrating that SMART I is a dsDNA nuclease. Figure 22C illustrates a time-series cleavage assay of MG34-1, suggesting slower kinetics compared to SpCas9. [Figure 22D] 22D illustrates data demonstrating that SMART I is a dsDNA nuclease. FIG. 22D illustrates a plasmid targeting assay. Left: Methods diagram showing engineered E. coli strains expressing effector nucleases (MG34-1 or MG34-9) and sgRNA cofactors. When transformed with a plasmid containing an antibiotic resistance gene with a targeted spacer or a non-targeted spacer (negative control), a growth defect occurs for the targeted plasmid. Center and Right: Bar graphs showing approximately 2-fold growth inhibition for plasmids encoding the MG34-1 (center) or MG34-9 (right) enzyme and sgRNA. [Figure 23] The amino acid content over the entire protein length for the group of SMART HNH endonuclease-associated RNAs and ORFs (HEARO) (35-1, 35-2, 35-3, 35-6, 35-102, and IscB) and SMART (34-1, 102-2, 102-14, 102-35, 102-45) nucleases is illustrated. High content of arginine (R) and lysine (K) is highlighted in green, whereas low content of methionine (M) is highlighted in orange. The amino acid content of most proteins in the Uniref50 database (Carugo, vol. 17, 12 (2008): 2187-91) was used for comparison. [Figure 24A] Illustrates a scatter plot of the average amino acid content of proteins in the Uniref50 database (X-axis) versus the percentage of amino acid content in SMART proteins (Y-axis). Arginine (R) and lysine (K) content deviate from a linear trend. [Figure 24B] Figure 1 illustrates a graph showing the ratio of amino acid percentages in SMART proteins to the percentages in the Uniref50 database. The mean of all ratios is 0.99 with a SD of 0.22. The green line indicates two standard deviations from the mean, assuming normality. [Figure 25A] Figure 25A illustrates data demonstrating that SMART enzymes are dsDNA nucleases. Figure 25A illustrates a histogram of cleavage site preferences for three SMART nucleases on the non-targeted strand (NTS) from next generation sequencing (NGS). The inset shows that SMART nucleases generate staggered cleavages, with cleavage at position 3 occurring on the targeted strand (TS), while cleavage at positions 5-7 from the PAM occurs on the NTS. TS cleavage sites were determined via Sanger run-off sequencing. [Figure 25B] Figure 25B illustrates data demonstrating that SMART enzymes are dsDNA nucleases. Figure 25B illustrates a bar plot of colony forming unit (cfu) measurements (logarithmic scale) showing E. coli growth inhibition in target conditions versus non-target controls. Plasmid interference assays for each nuclease were performed in triplicate with a SpCas9 positive control. [Figure 25C] Figure 25 illustrates data demonstrating that SMART enzymes are dsDNA nucleases. Figure 25C illustrates the measurement of in vitro DNA cleavage efficiency with various spacer lengths, showing a preference for 18-20 bp spacers for SMART nucleases, while SMART HEARO 35-1 prefers a 24 bp spacer. (*) Spacer lengths of 14 bp (34-1) and 30 bp (35-1 and 102-2) were not evaluated. [Figure 25D]Figure 25D illustrates data demonstrating that SMART enzymes are dsDNA nucleases. Figure 25D illustrates mismatch killing assays showing high specificity for targeted spacers at positions -1 to -13 from the PAM. Left: Bar plots of colony forming unit (cfu) measurements (log scale) showing E. coli growth inhibition in targeted conditions, spacers containing mismatches, and non-targeted controls. Top right: Diagram of mismatch killing assay. E. coli containing two plasmids for nuclease expression and guide expression is transformed with a library of targeted plasmids with mismatches in the protospacer. Bottom right: Heat map showing mismatch tolerance at each position of the targeted spacer. For targeted spacers with tolerated mismatches, growth is expected to be inhibited (purple). Positions with the required base pairing are not cleaved efficiently and will be relatively abundant in the output library (yellow). Plasmid interference (killing) assays with libraries for each nuclease were performed in duplicate. [Figure 26] Illustrates data demonstrating that MG102-2 is a highly active nuclease in human cells. Nuclease activity was tested by nucleofecting MG102-2 mRNA and two sgRNA targeting sites (guides A1 and B1) within the TRAC locus with increasing concentrations of sgRNA (150, 300, and 450 pmol / reaction). Mock controls represent background editing levels in the targeted region in the absence of mRNA and guide. [Figure 27] Illustrates mismatch killing assays showing the log fold change cleavage activity of spacers with mismatches at each position in the spacer tested for MG102-2 and MG35-1. [Figure 28] 1 illustrates data demonstrating that SMART nucleases have no activity on ssDNA. [Figure 29]SMART nuclease guide and salt concentration titrations are illustrated. In vitro cleavage assays of MG102-2 (lanes 1-6) and SMART HEARO 35-1 (lanes 7-18) show cleavage of the target plasmid DNA (approximately 3500 bp) into linear DNA products (less than 2500 bp). [Figure 30A] Figure 30A illustrates data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at a site targeted by a specific spacer (guide). Figure 30B illustrates data for MG102-2 targeting the AAVS1 locus. [Figure 30B] Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at the site targeted by a particular spacer (guide). Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data for MG102-39, MG102-42, MG102-48, MG33-34, MG102-26, and MG102-45, respectively, targeting the TRAC locus. [Figure 30C] Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at the site targeted by a particular spacer (guide). Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data for MG102-39, MG102-42, MG102-48, MG33-34, MG102-26, and MG102-45, respectively, targeting the TRAC locus. [Figure 30D]Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at the site targeted by a particular spacer (guide). Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data for MG102-39, MG102-42, MG102-48, MG33-34, MG102-26, and MG102-45, respectively, targeting the TRAC locus. [Figure 30E] Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at the site targeted by a particular spacer (guide). Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data for MG102-39, MG102-42, MG102-48, MG33-34, MG102-26, and MG102-45, respectively, targeting the TRAC locus. [Figure 30F] Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at the site targeted by a particular spacer (guide). Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data for MG102-39, MG102-42, MG102-48, MG33-34, MG102-26, and MG102-45, respectively, targeting the TRAC locus. [Figure 30G]Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data demonstrating SMART I editing efficiency in human cells. Nuclease activity was tested by nucleofecting SMART I mRNA and sgRNA (450 pmol / reaction) targeting multiple sites within the locus. Each bar represents the editing efficiency at the site targeted by a particular spacer (guide). Figure 30B, Figure 30C, Figure 30D, Figure 30E, Figure 30F, and Figure 30G illustrate data for MG102-39, MG102-42, MG102-48, MG33-34, MG102-26, and MG102-45, respectively, targeting the TRAC locus. [Diagram 31] Figure 1 illustrates a multiple sequence alignment of the 5'UTR nucleotide sequences of four SMART HEARO nucleases. The region before the start of the HEARO RNA (box) shows poor similarity, while strong conservation around the first structural hairpin is observed (inset). [Figure 32A] The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. [Figure 32B]The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. [Figure 32C] The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. [Fig. 32D]The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. Figures 32D-G illustrate the HEARO RNA secondary structures of three active nucleases: MG35-104 sg1, MG35-463 sg2 (CRISPR-independent), MG35-463 sg3 (CRISPR-associated), and MG35-556 dual-guide HEARO RNA (CRISPR-associated), respectively. [Figure 32E] The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. Figures 32D-G illustrate the HEARO RNA secondary structures of three active nucleases: MG35-104 sg1, MG35-463 sg2 (CRISPR-independent), MG35-463 sg3 (CRISPR-associated), and MG35-556 dual-guide HEARO RNA (CRISPR-associated), respectively. [Fig. 32F]The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. Figures 32D-G illustrate the HEARO RNA secondary structures of three active nucleases: MG35-104 sg1, MG35-463 sg2 (CRISPR-independent), MG35-463 sg3 (CRISPR-associated), and MG35-556 dual-guide HEARO RNA (CRISPR-associated), respectively. [Fig. 32G] The genomic context of SMART HEARO nucleases is illustrated. The majority of SMART HEARO nucleases are not CRISPR-associated (e.g., MG35-104, FIG. 32A), while a few SMART HEARO nucleases are associated with CRISPR arrays (e.g., MG35-463 and MG35-556 in FIG. 32B and FIG. 32C). SMART HEARO nucleases are represented by dark grey arrows with RRXRR and HNH Pfam domains annotated below the gene. HEARO RNAs predicted from covariance models (CM) are shown upstream of SMART HEARO effector genes (CM HEARO RNAs). RAR: repeat-anti-repeat. Figures 32D-G illustrate the HEARO RNA secondary structures of three active nucleases: MG35-104 sg1, MG35-463 sg2 (CRISPR-independent), MG35-463 sg3 (CRISPR-associated), and MG35-556 dual-guide HEARO RNA (CRISPR-associated), respectively. [Figure 33-1]Illustrates SMART HEARO cleavage activity in vitro. SMART II effectors were assayed in in vitro transcription / translation reactions incubated with their single guide RNA and PAM library (dsDNA target). Cleavage products were amplified via ligation to the cleavage site and subsequent PCR (successful RNA-guided cleavage with nuclease-generated bands of expected size, arrows). For Figure 33A, lane labels are as follows: L: ladder, PC: MG35-1 nuclease as positive control (PC), 1: MG35-94, 2: MG35-104, 3: MG35-346, 4: MG35-350, 5: MG35-423, 6: MG35-422, 7: MG35-461, 8: MG35-465, 9: MG35-515. For Figure 33B, lane labels are as follows: L: ladder, PC: MG35-1 nuclease as positive control (PC), 1: MG35-94, 2: MG35-104, 3: MG35-346, 4: MG35-350, 5: MG35-423, 6: MG35-422, 7: MG35-461, 8: MG35-465, 9: MG35-515. L: ladder, PC: MG35-1 nuclease as positive control (PC), 10: MG35-517, 11: MG35-518 with sgRNA design 1, 12: MG35-518 with sgRNA design 2, 13: MG35-519, 14: MG35-550 with sgRNA design 1, 15: MG35-550 with sgRNA design 2, 16: MG35-553, 17: MG35-554 with sgRNA design 1, 18: MG35-554 with sgRNA design 2, 19: MG35-555, and 20: MG35-556. [Figure 33-2]Illustrating SMART HEARO cleavage activity in vitro. SMART II effectors were assayed in in vitro transcription / translation reactions incubated with their single guide RNA and PAM library (dsDNA target). Cleavage products were amplified via ligation to the cleavage site and subsequent PCR (successful RNA-guided cleavage with a nuclease-generated band of the expected size, arrow). For Figure 33C, SMART II effectors were assayed for cleavage activity via the TAM / PAM enrichment protocol. Effectors were expressed in in vitro transcription / translation (IVTT) reactions in the presence of their single guide RNA and then added to the PAM library (dsDNA target). Cleavage products were amplified via ligation to the cleavage site and subsequent PCR (successful RNA-guided cleavage with a nuclease-generated band of the expected size, arrow). The reaction shown is prior to PCR cleanup, so primer and adapter dimer bands are observed at sizes <100 bp. [Figure 34-1] Illustrated are the TAM recognition motifs of active SMART HEARO nucleases. NGS sequencing of the bands identified in Figure 33A-C was used to generate the TAM and preferred cleavage positions for each nuclease. The structures of the working guides predicted by Geneious (Andronescu 2007) are shown in the inset. Cleavage typically occurs between positions 5-10 on the non-target strand. [Figure 34-2] Illustrated are the TAM recognition motifs of active SMART HEARO nucleases. NGS sequencing of the bands identified in Figure 33A-C was used to generate the TAM and preferred cleavage positions for each nuclease. The structures of the working guides predicted by Geneious (Andronescu 2007) are shown in the inset. Cleavage typically occurs between positions 5-10 on the non-target strand. [Figure 34-3]Illustrated are the TAM recognition motifs of active SMART HEARO nucleases. NGS sequencing of the bands identified in Figure 33A-C was used to generate the TAM and preferred cleavage positions for each nuclease. The structures of the working guides predicted by Geneious (Andronescu 2007) are shown in the inset. Cleavage typically occurs between positions 5-10 on the non-target strand. [Figure 35A] Illustrating the in vitro cleavage efficiency of active SMART HEARO nuclease. For Figure 35A, cleavage was measured by the transition of the reaction products from supercoiled (uncut) to linear (cut) and visualized on an Agilent Tapestation. Arrows indicate the initial dsDNA product (supercoiled) and the dsDNA product after successful targeted cleavage (linearization) by the enzyme. PE: PURExpress, sgRNA, single guide RNA. [Figure 35B] Figure 35B illustrates the in vitro cleavage efficiency of active SMART HEARO nuclease. Figure 35B illustrates a bar plot representation of the quantification from Figure 35A. DNA: DNA only control (negative control) with no RNP reaction, Apo: RNP reaction with no sgRNA added, Holo: RNP reaction with sgRNA. [Figure 36A] 36A illustrates SMART HEARO guide engineering. Five active SMART HEARO sgRNAs had one or more poly-T tracts in their sequences. Three poly-T mutant sgRNAs were designed for each candidate to compare activity versus the original guide. Guides were in vitro transcribed, normalized to the same concentration, and then used in an in vitro cleavage efficiency reaction. Figure 36A illustrates an exemplary guide RNA with a poly-T region and engineered guide sequence of MG35-518. [Figure 36B]36A illustrates SMART HEARO guide engineering. Five active SMART HEARO sgRNAs had one or more poly-T tracts in their sequence. Three poly-T mutant sgRNAs were designed for each candidate to compare activity versus the original guide. Guides were in vitro transcribed, normalized to the same concentration, and then used in an in vitro cleavage efficiency reaction. Figure 36B illustrates the cleavage efficiency of engineered SMART HEARO guide RNA versus native guide. Apo: no guide added (negative control), WT: native guide RNA. [Figure 37A] Phylogenetic analysis of SMART I nucleases is illustrated. Phylogenetic trees were inferred with FastTree or RAxML from global (g-ins-i) or local (l-ins-i) multiple sequence alignments. To account for phylogenetic uncertainty, six reconstructed sequences were taken from multiple trees (nodes highlighted with black circles: MG34-26, MG34-27, MG34-28, MG34-29, MG34-30, and MG34-31). [Figure 37B] Phylogenetic analysis of SMART I nucleases is illustrated. Phylogenetic trees were inferred with FastTree or RAxML from global (g-ins-i) or local (l-ins-i) multiple sequence alignments. To account for phylogenetic uncertainty, six reconstructed sequences were taken from multiple trees (nodes highlighted with black circles: MG34-26, MG34-27, MG34-28, MG34-29, MG34-30, and MG34-31). [Figure 37C] Phylogenetic analysis of SMART I nucleases is illustrated. Phylogenetic trees were inferred with FastTree or RAxML from global (g-ins-i) or local (l-ins-i) multiple sequence alignments. To account for phylogenetic uncertainty, six reconstructed sequences were taken from multiple trees (nodes highlighted with black circles: MG34-26, MG34-27, MG34-28, MG34-29, MG34-30, and MG34-31). [Figure 37D]Phylogenetic analysis of SMART I nucleases is illustrated. Phylogenetic trees were inferred with FastTree or RAxML from global (g-ins-i) or local (l-ins-i) multiple sequence alignments. To account for phylogenetic uncertainty, six reconstructed sequences were taken from multiple trees (nodes highlighted with black circles: MG34-26, MG34-27, MG34-28, MG34-29, MG34-30, and MG34-31). [Figure 38] Illustrates the 3D structure prediction of reconstituted SMART I MG34-30 versus the predicted structure of active MG34-1 nuclease. Good overall structural alignment of the proteins was observed by the overlap between the two structures and low RMSD values. [Figure 39] Illustrated are data demonstrating that reconstituted SMART I effectors are active nucleases. Novel SMART I effectors were assayed for cleavage activity via the PAM enrichment protocol. Effectors were expressed in in vitro transcription / translation (IVTT) reactions in the presence of single guide RNAs from other active MG34 nucleases and added to the PAM library (dsDNA target). Cleavage products were amplified via ligation to the cleavage site and subsequent PCR amplification (successful RNA guide cleavage with the expected 180 bp sized nuclease-generated band, arrow). MG34-27 and MG34-29 showed clear activity with the three tested guide RNAs. [Diagram 40] Illustrating the PAM recognition motifs of active SMART I nucleases from computational reconstructions. NGS sequencing of the bands identified in Figure 39 was used to generate the PAM and preferred cleavage position for each nuclease. Cleavage occurs between positions 6-8 from the PAM on the non-target strand.
[0047] Brief Description of the Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.
[0048] MG33 Nuclease SEQ ID NOs: 1, 463 to 486, 981 to 988, and 1289 to 1312 show full-length peptide sequences of MG33 nuclease.
[0049] SEQ ID NOs: 199 and 669-670 show the nucleotide sequence of tracrRNA predicted to function with MG33 nuclease.
[0050] SEQ ID NOs: 201 and 1003-1005 show the nucleotide sequences of predicted single guide RNA (sgRNA) sequences predicted to function with MG33 nuclease. "N" indicates a variable residue, and non-N residues represent scaffold sequences.
[0051] SEQ ID NOs: 1023 to 1028 show PAM sequences compatible with MG33 nuclease.
[0052] SEQ ID NOs: 1045-1054 show the CRISPR repeats of the MG33 nuclease described herein.
[0053] MG34 Nuclease SEQ ID NOs: 2 to 24, 487 to 488, and 1313 to 1321 show the full-length peptide sequences of MG34 nuclease.
[0054] SEQ ID NO: 200 shows the nucleotide sequence of tracrRNA predicted to function with MG34 nuclease.
[0055] SEQ ID NOs: 202, 203, and 613-616 show the nucleotide sequences of predicted single guide RNA (sgRNA) sequences predicted to function with MG34 nuclease. "N" indicates a variable residue, and non-N residues represent scaffold sequences.
[0056] SEQ ID NOs: 1023 to 1028 show PAM sequences compatible with MG34 nuclease.
[0057] SEQ ID NOs: 1055-1057 show the CRISPR repeats of the MG34 nuclease described herein.
[0058] MG35 Nuclease SEQ ID NOs: 25 to 198, 221 to 459, 489 to 580, 617 to 668, and 674 to 675 show full-length peptide sequences of MG35 nuclease.
[0059] SEQ ID NOs: 460-461 show the nucleotide sequence of MG35 tracrRNA, which is derived from the same locus as MG35 nuclease.
[0060] SEQ ID NOs: 462, 676, and 1229-1230 show the CRISPR repeats of the MG35 nuclease described herein.
[0061] SEQ ID NOs: 677-686, 1006-1012, and 1231-1259 show the nucleotide sequences of MG35 single guide RNAs.
[0062] SEQ ID NOs: 687-974 show the nucleotide sequence of the MG35 single guide RNA coding sequence.
[0063] SEQ ID NOs: 1029 to 1034 show PAM sequences compatible with MG35 nuclease.
[0064] SEQ ID NOs: 1172-1228 show the nucleotide sequence of the locus encoding the MG35 nuclease described herein.
[0065] MG102 Nuclease SEQ ID NOs: 581 to 612, 989 to 1002, and 1260 to 1273 show the full-length peptide sequences of MG102 nuclease.
[0066] SEQ ID NOs: 672-673 show the nucleotide sequence of MG102 tracrRNA from the same locus as MG102 nuclease
[0067] SEQ ID NOs:205-220 show exemplary nuclear localization sequences (NLS) that can be added to nucleases according to the present disclosure.
[0068] SEQ ID NOs: 1013 to 1022 show the nucleotide sequence of the MG102 single guide RNA.
[0069] SEQ ID NOs: 1035 to 1044 show PAM sequences compatible with MG102 nuclease.
[0070] SEQ ID NOs: 1058-1072 show the CRISPR repeats of the MG102 nuclease described herein.
[0071] SEQ ID NO: 1171 shows the nucleotide sequence of the locus encoding the MG102 nuclease described herein.
[0072] MG143 Nuclease SEQ ID NO: 975 shows the full-length peptide sequence of MG143 nuclease.
[0073] SEQ ID NO: 1073 shows the CRISPR repeat of the MG143 nuclease described herein.
[0074] MG144 Nuclease SEQ ID NOs: 976 to 979 and 1274 to 1288 show the full-length peptide sequences of MG144 nuclease.
[0075] SEQ ID NOs: 1074-1077 show the CRISPR repeats of the MG144 nuclease described herein.
[0076] MG145 Nuclease SEQ ID NO: 980 shows the full-length peptide sequence of MG145 nuclease.
[0077] SEQ ID NO: 1078 shows the CRISPR repeat of the MG145 nuclease described herein.
[0078] MG102 TRAC targeting SEQ ID NOs: 1079-1082 and 1145-1166 show the DNA sequences of the TRAC target sites.
[0079] SEQ ID NOs: 1083-1086 and 1123-1144 show the nucleotide sequences of sgRNAs engineered to function with MG102 nuclease to target TRAC.
[0080] MG33 TRAC targeting SEQ ID NOs: 1167-1168 show the nucleotide sequences of sgRNAs engineered to function with MG33 nuclease to target TRAC.
[0081] SEQ ID NOs: 1169-1170 show the DNA sequences of the TRAC target sites.
[0082] AAVS1 targeting SEQ ID NOs: 1087-1104 show the nucleotide sequences of sgRNAs engineered to function with MG102 nuclease to target AAVS1.
[0083] SEQ ID NOs: 1105 to 1122 show the DNA sequences of the AAVS1 target site. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0084] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used.
[0085] The practice of some of the methods disclosed herein may involve immunological, biochemical, chemical, molecular biology, microbiology, cell biology, genomics, and recombinant DNA techniques, unless otherwise indicated. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FMA Usubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)) (incorporated herein in its entirety by reference).
[0086] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, to the extent the terms "comprising," "including," "having," "having," "having," or variations thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0087] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0088] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural, functional, or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-cell eukaryotic cells, protozoan cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, liverworts, mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.), and cells from other organisms (e.g., cereals, vegetables, fruits, and vegetables). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from vertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. In some cases, the cells are not derived from a naturally occurring organism (e.g., the cells may be synthetically produced and sometimes referred to as artificial cells).
[0089] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif.; fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP available from Amersham, Arlington Heights, Ill.; Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Examples of chromosomal labeling nucleotides available from Probes, Eugene, Oreg. include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides may also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs may include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP). Nucleotides may include nucleotide analogs. In some embodiments, nucleotide analogs may include the structure of a natural nucleotide modified at any position to change certain chemical properties of the nucleotide but retain the ability of the nucleotide analog to perform its intended function (e.g., hybridization to other nucleotides in RNA or DNA). Examples of positions of nucleotides that can be derivatized include the 5-position, e.g., 5-(2-amino)propyluridine, 5-bromouridine, 5-propyneuridine, 5-propenyluridine, etc.; the 6-position, e.g., 6-(2-amino)propyluridine; the 8-position for adenosine or guanosine, e.g., 8-bromoguanosine, 8-chloroguanosine, 8-fluoroguanosine, etc. Nucleotide analogs also include deazanucleotides, e.g., 7-deaza-adenosine: O- and N-modified (e.g., alkylated, e.g., N6-methyladenosine) nucleotides, and other heterocyclic modified nucleotide analogs, such as those described in Herdewijn, Antisense Nucleic Acid Drug Dev., 2000 Aug. 10(4):297-310. Nucleotide analogs may also include modifications to the sugar moiety of the nucleotide. For example, the 2'OH group may be replaced by a group selected from H, OR, R, F, Cl, Br, I, SH, SR, NH2, NHR, NR2, COOR, or OR, where R is a substituted or unsubstituted C1-C6 alkyl, alkenyl, alkynyl, aryl, etc. Other possible modifications include those described in U.S. Patent Nos. 5,858,988 and 6,291,438.Examples of nucleotide positions that can be derivatized include the 5-position, e.g., 5-(2-amino)propyluridine, 5-bromouridine, 5-propyneuridine, 5-propenyluridine, etc.; the 6-position, e.g., 6-(2-amino)propyluridine; the 8-position for adenosine or guanosine, e.g., 8-bromoguanosine, 8-chloroguanosine, 8-fluoroguanosine, etc. Nucleotide analogs also include deazanucleotides, e.g., 7-deaza-adenosine: O- and N-modified (e.g., alkylated, e.g., N6-methyladenosine) nucleotides, and other heterocyclic modified nucleotide analogs, such as those described in Herdewijn, Antisense Nucleic Acid Drug Dev., 2000 Aug. 10(4):297-310. Nucleotide analogs may also include modifications to the sugar portion of the nucleotide. For example, the 2'OH group may be replaced by a group selected from H, OR, R, F, Cl, Br, I, SH, SR, NH2, NHR, NR2, COOR, or OR, where R is a substituted or unsubstituted C1-C6 alkyl, alkenyl, alkynyl, aryl, etc. Other possible modifications include those described in U.S. Patent Nos. 5,858,988 and 6,291,438.
[0090] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-stranded, double-stranded, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and may perform any function. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci defined from binding analyses, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers.The sequence of nucleotides may be interrupted by non-nucleotide components.
[0091] The term "transfection" or "transfected" generally refers to the introduction of a nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a genetic sequence encoding a complete protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88, which is incorporated herein by reference in its entirety.
[0092] The terms "peptide", "polypeptide" and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues linked by peptide bonds. The term does not refer to a particular length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers that contain at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by any other manipulation, such as, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include natural amino acids and unnatural amino acids, which are chemically modified to include a non-naturally occurring group or chemical moiety on the amino acid. An amino acid analog may refer to an amino acid derivative. The term "amino acid" includes both D- and L-amino acids.
[0093] As used herein, the term "non-natural" can generally refer to a nucleic acid or polypeptide sequence that is not present in a naturally occurring nucleic acid or protein. Non-natural can refer to an affinity tag. Non-natural can refer to a fusion. Non-natural can refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, or deletion. A non-natural sequence can exhibit or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, an ubiquitination activity, etc.) that can also be exhibited by the nucleic acid or polypeptide sequence to which the non-natural sequence is fused. A non-natural nucleic acid or polypeptide sequence can be linked to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.
[0094] As used herein, the term "promoter" generally refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotides where RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often called transcription factors, which promote the binding of RNA polymerase to DNA, thereby resulting in gene transcription. A "basal promoter", also called a "core promoter", may generally refer to a promoter that contains all the basic elements to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically, but not necessarily, contain a TATA-box or CAAT box.
[0095] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) or by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0096] As used herein, "operably linked," "operably linked," "operably linked," or grammatical equivalents thereof generally refer to the juxtaposition of genetic elements, such as promoters, enhancers, polyadenylation sequences, and the like, where the elements are in a relationship that allows them to operate in an expected manner. For example, a regulatory element, which may include a promoter sequence or an enhancer sequence, is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.
[0097] As used herein, a "vector" generally refers to a polymer or an association of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally includes a genetic element, such as a regulatory element, operably linked to a gene to facilitate expression of the gene in a target.
[0098] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of a gene or genes with regulatory elements that are operably linked for expression.
[0099] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be the ability to affect expression in a manner attributable to the full-length sequence.
[0100] As used herein, an "engineered" object generally indicates that the object has been modified by human intervention. By way of non-limiting examples, a nucleic acid may be modified by altering its sequence to a sequence that does not occur in nature, a nucleic acid may be modified by ligating to a nucleic acid with which it is not naturally associated such that the ligated product has a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro with a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, and an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.
[0101] As used herein, the term "optimally aligned" generally refers to the alignment of two amino acid sequences that gives the highest percent identity score or maximizes the number of matching residues.
[0102] As used herein, "synthetic" and "artificial" are used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) with naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.
[0103] As used herein, the term "tracrRNA" or "tracr sequence" can generally refer to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity or similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S.pyogenes, S.aureus, etc.). A tracrRNA can refer to a nucleic acid having up to about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity or similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S.pyogenes, S.aureus, etc.). A tracrRNA can refer to a modified form of a tracrRNA that may include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeras. tracrRNA may refer to a nucleic acid that may be at least about 60% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from S.pyogenes, S.aureus, etc.) sequence over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may be at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from S.pyogenes, S.aureus, etc.) sequence over a section of at least six consecutive nucleotides. Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions that have complementarity to portions of the repeat sequences in adjacent CRISPR arrays.
[0104] As used herein, a "guide nucleic acid" can generally refer to a nucleic acid that can hybridize to another nucleic acid. A guide nucleic acid can be RNA. A guide nucleic acid can be DNA. A guide nucleic acid can be programmed to bind to a sequence of a nucleic acid in a site-specific manner. The nucleic acid to be targeted, or the target nucleic acid, can include nucleotides. A guide nucleic acid can include nucleotides. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to a guide nucleic acid and hybridizes with the guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore not complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can include a polynucleotide strand and can be referred to as a "single guide nucleic acid". A guide nucleic acid can include two polynucleotide strands and can be referred to as a "double guide nucleic acid". Otherwise, the term "guide nucleic acid" can be inclusive, referring to both single and double guide nucleic acids. A guide nucleic acid may include a segment that may be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence." The nucleic acid targeting segment may include a sub-segment that may be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."
[0105] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a certain percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); or CLUSTALW using parameters of the Smith-Waterman homology search algorithm with parameters of match of 2, mismatch of -1, and gap of -1; MUSCLE using default parameters; MAFFT using parameters retree of 2 and maximum iteration of 1000; Novafold using default parameters; HMMER hmmalign using default parameters.
[0106] As used herein, the term "RuvC_III domain" generally refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three discontinuous segments, RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or a segment thereof (e.g., RuvC_I, RuvC_II, or RuvC_III) can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with annotated domains, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF18541 for RuvC_III).
[0107] As used herein, the term "HNH domain" generally refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).
[0108] As used herein, the term "bridge helix domain" or "BH domain" generally refers to an arginine-rich helical domain present in Cas enzymes that plays a key role in initiating cleavage activity upon binding to target DNA.
[0109] As used herein, the term "recognition domain" or "REC domain" generally refers to a domain that is believed to interact with the repeat:anti-repeat duplex of a gRNA and mediate the formation of a Cas endonuclease / gRNA complex.
[0110] As used herein, the term "wedge domain" or "WED domain" generally refers to a fold comprising a twisted five-stranded beta sheet flanked by four alpha helices, which is generally involved in the recognition of distorted repeat:anti-repeat duplexes for Cas enzymes. WED domains may be involved in the recognition of single guide RNA scaffolds.
[0111] As used herein, the term "PAM interaction domain" or "PI domain" generally refers to a domain found in Cas enzymes that is positioned in the endonuclease DNA complex to recognize the PAM sequence on the non-complementary DNA strand of the guide RNA.
[0112] overview The discovery of new Cas enzymes with unique functionality and structure could confer the potential to further disrupt deoxyribonucleic acid (DNA) editing technologies, improving their speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microbes and the sheer diversity of microbial species, there are relatively few functionally characterized CRISPR / Cas enzymes in the literature. This is in part because the vast number of microbial species are not easily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches representing a large number of microbial species could dramatically increase the number of documented new CRISPR / Cas systems, conferring the potential to expedite the discovery of new oligonucleotide editing functions. A fruitful recent example of such an approach is illustrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0113] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode RNA-based targeting elements; and (ii) an ORF encoding a Cas that encodes a nuclease polypeptide directed by the RNA-based targeting element flanked by accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (PAM is usually a sequence that is not commonly represented within the host genome). Depending on the exact function and composition of the system, CRISPR-Cas systems are commonly organized into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities.
[0114] Class I CRISPR-Cas systems have large, multi-subunit effector complexes and include types I, III, and IV.
[0115] Type I CRISPR-Cas systems are considered to be of intermediate complexity in terms of components. In Type I CRISPR-Cas systems, an array of RNA targeting elements is transcribed as a long precursor crRNA (pre-crRNA) that is processed at the repeat elements to release a short mature crRNA that directs the nuclease complex to the nucleic acid target, followed by an appropriate short consensus sequence called the protospacer adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nuclease functions primarily as a DNA nuclease.
[0116] Type III CRISPR systems can be characterized by the presence of a central nuclease known as Cas10, along with repeat-associated mysterious proteins (RAMPs) that contain Csm or Cmr protein subunits. Similar to type I systems, mature crRNA is processed from pre-crRNA using a Cas6-like enzyme. Unlike type I and II systems, type III systems appear to target and cleave DNA-RNA duplexes (such as the DNA strand used as a template for RNA polymerase).
[0117] Type IV CRISPR-Cas systems possess an effector complex that contains a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) family, and in some cases a predicted small subunit gene; such systems are commonly found on endogenous plasmids.
[0118] Type II CRISPR-Cas systems generally have a single polypeptide multi-domain nuclease effector and include Type II, Type V, and Type VI.
[0119] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, processing of CRISPR arrays into mature crRNA does not require the presence of special endonuclease subunits, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence, which interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure that is cleaved by endogenous RNAse III to generate the mature effector enzyme loaded with both tracrRNA and crRNA. Cas II nuclease is a DNA nuclease. Type II effectors generally exhibit a structure that includes a RuvC-like endonuclease domain that fits into an RNase H fold with an unrelated HNH nuclease domain inserted into the fold of the RuvC-like nuclease domain. The RuvC-like domain is involved in cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is involved in cleavage of the replacement DNA strand.
[0120] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Cas12) structure similar to type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNA, but unlike type II systems that require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems can cleave pre-crRNA using the effector nuclease itself. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.
[0121] Type VI CRISPR-Cas systems have an RNA guide RNA endonuclease. Instead of a RuvC-like domain, the single polypeptide effector of type VI systems (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems also do not appear to require tracrRNA to process pre-crRNA into crRNA. However, like type V systems, some type VI systems (e.g., C2C2) appear to have robust single-stranded non-specific nuclease (ribonuclease) activity that is activated by the first crRNA-directed cleavage of the target RNA.
[0122] Class II CRISPR-Cas are simpler constructs and therefore have been the most widely applied in engineering and development as engineered nuclease / genome editing applications.
[0123] One of the early applications of such a system for in vitro use can be found in Jinek et al. (Science. 2012 Aug 17;337(6096):816-21, incorporated herein by reference in its entirety). The Jinek study first combined (i) recombinantly expressed, purified full-length Cas9 (e.g., class II, type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature ∼42 nt crRNA (total crRNA transcribed in vitro from a synthetic DNA template bearing a T7 promoter sequence) that yields ∼20 nt of 5' sequence complementary to the target DNA sequence to be cleaved, followed by a 3' tracr binding sequence, (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template bearing a T7 promoter sequence, and (iv) Mg 2+ Jinek later described an improved engineered system in which (ii) the crRNA is attached to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can itself guide Cas9 to the target (compare the top and bottom panels of FIG. 2).
[0124] Mali et al. (Science. 2013 Feb 15;339(6121):823-826.), which is incorporated herein by reference in its entirety, later adapted this system for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding a codon-optimized Cas9 (e.g., a class II, type II Cas enzyme) under a suitable mammalian promoter with a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding an sgRNA (having a 5' sequence starting with G followed by 20 nt of complementary targeting nucleic acid sequence attached to the 3' tracr binding sequence, a linker, and the tracrRNA sequence) under a suitable polymerase III promoter (e.g., U6 promoter).
[0125] MG enzyme In one aspect, the present disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) an endonuclease. In some cases, the endonuclease includes a RuvC domain and an HNH domain. The endonuclease may be from an uncultured microorganism. The endonuclease may be a Cas endonuclease. The endonuclease may be a class 2 endonuclease. The endonuclease may be a class 2, type II Cas endonuclease. The engineered nuclease system may include (b) an engineered guide ribonucleic acid structure. The engineered guide ribonucleic acid structure may be configured to form a complex with the endonuclease. In some cases, the engineered guide ribonucleic acid structure configured to form a complex with the endonuclease includes a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize to a target deoxyribonucleic acid sequence. In some cases, the engineered guide ribonucleic acid structure configured to form a complex with an endonuclease comprises a tracr ribonucleic acid sequence. The tracr ribonucleic acid sequence can be configured to bind to the endonuclease. In some cases, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 100 kDa or less, about 90 kDa or less, about 80 kDa or less, about 70 kDa or less, about 60 kDa or less, about 50 kDa or less, about 40 kDa or less, about 30 kDa or less, about 20 kDa or less, or about 10 kDa or less.
[0126] In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof.
[0127] In one aspect, the present disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) an endonuclease. The endonuclease may include a RuvC-1 domain or a RucV domain. The endonuclease may include an HNH domain. The endonuclease may include a RuvC-1 domain and an HNH domain. The endonuclease may be a Cas endonuclease. The endonuclease may be a class 2 endonuclease. The endonuclease may be a class 2, type II Cas endonuclease. The engineered nuclease system may include (b) an engineered guide ribonucleic acid. The engineered guide ribonucleic acid structure may be configured to form a complex with the endonuclease. The engineered guide ribonucleic acid structure configured to form a complex with the endonuclease may include a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize to a target deoxyribonucleic acid sequence. The engineered guide ribonucleic acid structure configured to form a complex with an endonuclease can include a tracr ribonucleic acid sequence. The tracr ribonucleic acid sequence can be configured to bind to the endonuclease. In some embodiments, the endonuclease may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321. The endonuclease may be an archaeal endonuclease. The endonuclease can be a class 2, type II Cas endonuclease.The endonuclease may include an arginine-rich region or a domain having PF14239 homology that includes an RRxRR motif. The arginine-rich region or the domain having PF14239 homology has at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70% identity with the arginine-rich region or the domain having PF14239 homology of any one of SEQ ID NOs: 1 to 198, 221 to 459, 463 to 612, 617 to 668, 674 to 675, 975 to 1002, and 1260 to 1321, or a variant thereof. , at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity. Domain boundaries of the arginine-rich domain or domains with PF14239 homology may be identified by optimal alignment to MG34-1 or MG34-9. The endonuclease may include a REC domain. The REC domain may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the REC domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof.The domain boundaries of the REC domain may be identified by optimal alignment to MG34-1 or MG34-9. The endonuclease may include a BH (bridge helix). The BH domain may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the BH domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. The domain boundaries of the BH domain can be identified by optimal alignment to MG34-1 or MG34-9.
[0128] The endonuclease may include a WED (wedge) domain. The WED domain may include a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the WED domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. The domain boundaries of the WED domain may be identified by optimal alignment to MG34-1 or MG34-9. The endonuclease may contain a PI (PAM interacting) domain. The PI domain may include a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the PI domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. The domain boundaries of the PI domain can be identified by optimal alignment to MG34-1 or MG34-9.
[0129] In some cases, the endonuclease is derived from an uncultured microorganism. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 50, at least 60, at least 70, at least 80 contiguous nucleotides from any one of SEQ ID NOs: 199-200, 460-461, or 669-673; or The present invention relates to a method for producing a nucleic acid sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 50, at least 60, at least 70, or at least 80 consecutive nucleotides of a non-variable nucleotide of any one of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259.
[0130] In some cases, the guide nucleic acid structure comprises SEQ ID NO: 201. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 202. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 203. In some cases, the guide nucleic acid structure comprises SEQ ID NOs: 201-203. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 613. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 614. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 615. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 616.
[0131] In one aspect, the present disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) an engineered guide ribonucleic acid structure. The engineered guide ribonucleic acid structure may include a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize to a target deoxyribonucleic acid sequence. The engineered guide ribonucleic acid structure may include a tracr ribonucleic acid sequence. The tracr ribonucleic acid sequence may be configured to bind to an endonuclease. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 50, at least 60, at least 70, at least 80 contiguous nucleotides from any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or SEQ ID NOs: 201-203, 613-616, 677-686, The present invention relates to a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, or at least 80 contiguous nucleotides of any one of the non-variable nucleotides of 1003 to 1022, or 1231 to 1259.
[0132] In some cases, the engineered nuclease system can include an endonuclease. The endonuclease can be a class 2 endonuclease. The endonuclease can be a Cas endonuclease. The endonuclease can be a class 2, type II Cas endonuclease.
[0133] In some cases, the endonuclease has a specific molecular weight range. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, about 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some cases, the endonuclease comprises a specific number of residues. The endonuclease may comprise about 1,100 or less residues, about 1,000 or less residues, about 950 or less residues, about 900 or less residues, about 850 or less residues, about 800 or less residues, about 750 or less residues, about 700 or less residues, about 650 or less residues, about 600 or less residues, about 550 or less residues, about 500 or less residues, about 450 or less residues, about 400 or less residues, or about 350 or less residues. The endonuclease may comprise about 700 to about 1,100 residues. The endonuclease may comprise about 400 to about 600 residues. In some cases, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide. The single ribonucleic acid polynucleotide may comprise a guide ribonucleic acid sequence and a tracr ribonucleic acid sequence.
[0134] In some cases, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a prokaryotic genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a bacterial genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to an archaeal genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a eukaryotic genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a fungal genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a plant genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a mammalian genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a human genomic sequence.
[0135] In some cases, the guide ribonucleic acid targeting sequence or spacer is 10-30 nucleotides in length, or 12-28 nucleotides in length, or 15-24 nucleotides in length. In some cases, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some cases, the NLS comprises a sequence selected from SEQ ID NOs: 205-220.
[0136] [Table 1]
[0137] The present disclosure includes any variant of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without destroying the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by replacing amino acids with similar hydrophobicity, polarity, and R chain length. Additionally or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that are mutated between species without changing the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of the endonuclease protein sequences described herein. In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions such that the activity of one or more critical active site residues or guide RNA binding residues of the endonuclease is not destroyed. In some embodiments, a functional variant of any of the proteins described herein lacks at least one substitution of the conserved or functional residues called out in Figure 4. In some embodiments, a functional variant of any of the proteins described herein lacks all substitutions of the conserved or functional residues called out in Figure 4.The present disclosure also provides modified active variants of any of the nucleases described herein. Such modified active variants may include inactivating mutations in one or more catalytic residues identified herein (e.g., FIG. 4) or generally described for the RuvC domain. Such modified active variants may include change switch mutations in catalytic residues of the RuvCI, RuvCII, or RuvCIII domains.
[0138] Conservative substitution tables providing functionally similar amino acids are available in a variety of references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). Each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) Serine (S), Threonine (T), and 8) Cysteine (C), Methionine (M)
[0139] Any variant of the endonucleases described herein that have sequence identity with a particular domain are included in the present disclosure. The domain may be an arginine-rich domain (e.g., a domain with PF14239 homology), a REC (recognition) domain, a BH (bridge helix) domain, a WED (wedge) domain, a PI (PAM interaction) domain, a PF14239 homology domain, or any other domain described herein. In some embodiments, the residues that comprise one or more of these domains are identified in the protein by alignment to one of the following proteins (e.g., when one of the following proteins and the protein of interest are optimally aligned), and residue boundaries, e.g., domains, are described.
[0140] [Table 2]
[0141] In some cases, the engineered nuclease system further comprises a single-stranded DNA repair template. In some cases, the engineered nuclease system further comprises a double-stranded DNA repair template. In some cases, the single-stranded or double-stranded DNA repair template comprises a first homologous arm in the 5' to 3' direction, the first homologous arm comprising a sequence of at least 20 nucleotides on the 5' side of the target deoxyribonucleic acid sequence. In some cases, the single-stranded or double-stranded DNA repair template comprises a synthetic DNA sequence of at least 10 nucleotides in the 5' to 3' direction. In some cases, the single-stranded or double-stranded DNA repair template comprises a second homologous arm in the 5' to 3' direction, the second homologous arm comprising a sequence of at least 20 nucleotides on the 3' side of the target sequence. In some cases, the single-stranded or double-stranded DNA repair template comprises, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, or a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence.
[0142] In some cases, the first homology arm comprises a sequence of at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. 2+ In some cases, the endonuclease and the tracr ribonucleic acid sequence are derived from separate bacterial species. In some cases, the endonuclease and the tracr ribonucleic acid sequence are derived from separate bacterial species within the same phylum.
[0143] In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-24 or 462-488. In some cases, the guide RNA structure comprises an RNA sequence predicted to include a hairpin. In some cases, the hairpin comprises a stem and a loop. In some cases, the stem comprises at least 12 pairs, at least 14 pairs, at least 16 pairs, or at least 18 pairs, or ribonucleotides.
[0144] In some cases, the guide RNA structure further comprises a second stem and a second loop. In some cases, the second stem comprises at least 5 pairs, at least 6 pairs, at least 7 pairs, at least 8 pairs, at least 9 pairs, or at least 10 pairs of ribonucleotides. In some cases, the guide RNA structure further comprises an RNA structure, the RNA structure comprising at least two hairpins. In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:1, and the guide RNA structure comprises an RNA sequence predicted to include at least four hairpins. In some cases, each of the four hairpins comprises a stem and a loop.
[0145] In some cases, the engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:1. In some cases, the engineered nuclease system comprises a guide RNA structure that comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least one of the non-variable nucleotides of SEQ ID NO:199 or SEQ ID NO:201.
[0146] In some cases, the engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least one of SEQ ID NOs:1-24 or 462-488. In some cases, the engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of the non-variable nucleotides of any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or any one of SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259.
[0147] In some cases, sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using Smith-Waterman homology search algorithm parameters. In some cases, sequence identity is determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and a BLOSUM62 scoring matrix setting gap costs at presence of 11, extension of 1, and using a conditional composition score matrix adjustment.
[0148] In some cases, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some cases, the endonuclease has less than 80% identity, less than 75% identity, less than 70% identity, less than 65% identity, less than 60% identity, less than 55% identity, or less than 50% identity to a Cas9 endonuclease.
[0149] In one aspect, the present disclosure provides an engineered guide RNA comprising (a) a DNA targeting segment. In some cases, the DNA targeting segment comprises a nucleotide sequence that is complementary to a target sequence in a target DNA molecule. In some cases, the engineered single guide ribonucleic acid polynucleotide comprises a protein binding segment. The protein binding segment comprises two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex. In some cases, the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide. In some cases, the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease comprising a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof.
[0150] In some cases, the DNA targeting segment is located 5' to both of the two complementary stretches of nucleotides. In some cases, the protein binding segment engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of the non-variable nucleotides of any one of SEQ ID NOs: 199-200, 460-461, 669-673, or SEQ ID NOs: 201-203, 613-616, 677-686, 1003-1022, or 1231-1259. In some cases, the deoxyribonucleic acid polynucleotide encodes an engineered guide ribonucleic acid polynucleotide described herein.
[0151] In one aspect, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence. In some cases, the engineered nucleic acid sequence is optimized for expression in an organism. In some cases, the nucleic acid encodes an endonuclease. The endonuclease can be a Cas endonuclease. The endonuclease can be a class 2 endonuclease. The endonuclease can be a class 2, type II Cas endonuclease. In some cases, the endonuclease comprises a RuvC domain and an HNH domain. In some cases, the endonuclease is derived from an uncultured microorganism. In some cases, the endonuclease has a specific molecular weight range. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, about 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some cases, the endonuclease comprises a specific number of residues. The endonuclease may comprise about 1,100 or less residues, about 1,000 or less residues, about 950 or less residues, about 900 or less residues, about 850 or less residues, about 800 or less residues, about 750 or less residues, about 700 or less residues, about 650 or less residues, about 600 or less residues, about 550 or less residues, about 500 or less residues, about 450 or less residues, about 400 or less residues, or about 350 or less residues. The endonuclease may comprise about 700 to about 1,100 residues. The endonuclease may comprise about 400 to about 600 residues.In some cases, the endonuclease comprises SEQ ID NO: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof, having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof. In some cases, the endonuclease further comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some cases, the NLS comprises a sequence selected from SEQ ID NOs: 205-220.
[0152] In some cases, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human. In some cases, the organism is a prokaryote. In some cases, the organism is a bacterium. In some cases, the organism is a eukaryote. In some cases, the organism is a fungus. In some cases, the organism is a plant. In some cases, the organism is a mammal. In some cases, the organism is a rodent. In some cases, the organism is a human. If the organism is a prokaryote or a bacterium, the organism can be a different organism than the organism from which the endonuclease is derived. In some cases, the organism is not an uncultured microorganism.
[0153] In one aspect, the disclosure provides a vector comprising a nucleic acid sequence. In some cases, the nucleic acid encodes an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2, type II Case endonuclease. The endonuclease can include a RuvC-I domain and an HNH domain. In some cases, the endonuclease is derived from an uncultured microorganism. In some cases, the endonuclease has a specific molecular weight range. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, about 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some cases, the endonuclease comprises a specific number of residues. The endonuclease may comprise about 1,100 or less residues, about 1,000 or less residues, about 950 or less residues, about 900 or less residues, about 850 or less residues, about 800 or less residues, about 750 or less residues, about 700 or less residues, about 650 or less residues, about 600 or less residues, about 550 or less residues, about 500 or less residues, about 450 or less residues, about 400 or less residues, or about 350 or less residues. The endonuclease may comprise about 700 to about 1,100 residues. The endonuclease may comprise about 400 to about 600 residues.
[0154] In some aspects, the disclosure provides an endonuclease as described herein configured to induce a double-stranded break proximal to the target locus of interest 5' of a protospacer adjacent motif (PAM). The endonuclease may induce a double-stranded break 6-8 nucleotides from the PAM or 7 nucleotides from the PAM. In some aspects, the disclosure provides an endonuclease as described herein configured to induce a single-stranded break proximal to the target locus of interest 5' of a protospacer adjacent motif (PAM). The endonuclease may induce a single-stranded break 6-8 nucleotides from the PAM or 7 nucleotides from the PAM. In some cases, the endonuclease configured to induce a single-stranded break comprises an inactivating mutation in one or more catalytic residues of an endonuclease as described herein.
[0155] In some aspects, the present disclosure provides an endonuclease system as described herein configured to cause a chemical modification of a nucleotide base in or near a target locus targeted by the endonuclease system. In this case, the chemical modification of a nucleotide base generally refers to a modification of a chemical moiety involved in base pairing, rather than a modification of the sugar or phosphate portion of the nucleotide. The chemical modification is deamination of an adenosine or cytosine nucleotide. In some cases, the endonuclease system configured to cause a chemical modification includes an endonuclease having a base editor attached or fused in frame to the endonuclease. The endonuclease to which the base editor is fused or attached may include an inactivating mutation in at least one catalytic residue of the endonuclease (e.g., the RuvC domain). The base editor may be fused to the N-terminus or C-terminus of the endonuclease, or linked via chemical conjugation. The base editor may be, but is not limited to, RNA-specific adenosine deaminase 1 (ADAR1), RNA-specific adenosine deaminase 2 (ADAR2), apolipoprotein B MRNA editing enzyme catalytic subunit 1 (APOBEC1), apolipoprotein B MRNA editing enzyme catalytic subunit 2 (APOBEC2), apolipoprotein B MRNA editing enzyme catalytic subunit 3A (APOBEC3A), apolipoprotein B MRNA editing enzyme catalytic subunit 3B (APOBEC3B), apolipoprotein B MRNA editing enzyme catalytic subunit 3C (APOBEC3C), apolipoprotein B MRNA editing enzyme catalytic subunit 3D (APOBEC3D), apolipoprotein B MRNA editing enzyme catalytic subunit 3F (APOBEC3F), apolipoprotein B MRNA editing enzyme catalytic subunit 3G (APOBEC3G), apolipoprotein B MRNA editing enzyme catalytic subunit 3H (APOBEC3H), or apolipoprotein B The base editor may include any adenosine or cytosine deaminase, including mRNA editing enzyme catalytic subunit 4 (APOBEC4), or a functional fragment thereof. The base editor may include a yeast, eukaryotic, mammalian, or human base editor.
[0156] In some aspects, the disclosure provides an endonuclease system as described herein configured to cause a chemical modification of a histone in or near a target locus targeted by the endonuclease system. In some cases, the endonuclease system configured to cause a chemical modification of a histone comprises an endonuclease having a histone editor linked or fused in frame to the endonuclease. The histone editor may be N-terminally or C-terminally linked or fused to the endonuclease. In some embodiments, the chemical modification may comprise methylation, acetylation, demethylation, or deacetylation. The endonuclease to which the histone editor is fused or linked may comprise an inactivating mutation in at least one catalytic residue of the endonuclease (e.g., the RuvC domain). Histone editors include histone methyltransferases (e.g., ASH1L, DOT1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, PRDM2, SET, SETBP1, SETD1A, SETD1B, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETD9, SETDB1, SETDB2, SETMAR, SMYD1, SMYD2, SMYD3, SMYD4, SMYD5, SUV39H1, SUV39H2, SUV420H1, or may include SUV420H2), histone demethylases (e.g., KDM1, KDM2, KDM3, KDM4, KDM5, or KDM6 family), histone acetyltransferases (e.g., GNAT or HAT family acetyltransferases), or histone deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC4, HDAC5, HDAC6, HDAC7, HDAC8, HDAC9, HDAC10, HDAC11, SIRT1, SIRT2, SIRT3, SIRT4, SIRT5, SIRT6, or SIRT7). The histone editor may include a yeast, eukaryotic, mammalian, or human histone editor.
[0157] In one aspect, the disclosure provides a vector comprising a nucleic acid described herein. In some cases, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure. The engineered guide ribonucleic acid structure can be configured to form a complex with an endonuclease. In some cases, the engineered guide ribonucleic acid structure comprises a guide ribonucleic acid sequence. In some cases, the guide ribonucleic acid sequence is configured to hybridize to a target deoxyribonucleic acid sequence. In some cases, the engineered guide ribonucleic acid structure comprises a tracr ribonucleic acid sequence. In some cases, the tracr ribonucleic acid sequence is configured to bind to an endonuclease. In some cases, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
[0158] In one aspect, the disclosure provides a cell comprising any of the vectors described herein.
[0159] In one aspect, the disclosure provides a method of producing an endonuclease. The method may include culturing any of the cells described herein.
[0160] In one aspect, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide. The method may include contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2, type II Cas endonuclease. The endonuclease may be complexed with an engineered guide ribonucleic acid structure. In some cases, the engineered guide ribonucleic acid structure is configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM). In some cases, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 100 kDa or less, about 90 kDa or less, about 80 kDa or less, about 70 kDa or less, about 60 kDa or less, about 50 kDa or less, about 40 kDa or less, about 30 kDa or less, about 20 kDa or less, or about 10 kDa or less. In some cases, the endonuclease comprises a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof.
[0161] In one aspect, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide. The method can include contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2, type II Cas endonuclease. The endonuclease can be complexed with an engineered guide ribonucleic acid structure. In some cases, the engineered guide ribonucleic acid structure can be configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide includes a protospacer adjacent motif (PAM). In some cases, the PAM is NGG. In some cases, the endonuclease comprises a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, 617-668, 674-675, 975-1002, 1260-1321, or a variant thereof.
[0162] In some cases, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some cases, the endonuclease is from an uncultured microorganism. In some cases, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, or bacterial double-stranded deoxyribonucleic acid polynucleotide from a species other than the species from which the endonuclease is derived.
[0163] In one aspect, the disclosure provides a method of modifying a target nucleic acid locus. The method may include delivering an engineered nuclease system described herein to a target nucleic acid locus. In some cases, the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure. In some cases, the complex is configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic locus. In some cases, modifying the target nucleic acid locus includes binding, nicking, cleaving, or marking the target nucleic acid locus.
[0164] In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, the target nucleic acid locus comprises genomic eukaryotic DNA, viral DNA, or bacterial DNA. In some cases, the target nucleic acid locus comprises bacterial DNA. The bacterial DNA may be from a bacterial species different from the species from which the endonuclease was derived. In some cases, the target nucleic acid locus is in vitro. In some cases, the target nucleic acid locus is within a cell. In some cases, the endonuclease and the engineered guide nucleic acid structure are provided encoded on separate nucleic acid molecules. In some cases, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some cases, the cell is from a species different from the species from which the endonuclease is derived.
[0165] In some cases, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a nucleic acid described herein or a vector described herein. In some cases, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some cases, the nucleic acid comprises a promoter to which an open reading frame encoding an endonuclease is operably linked. In some cases, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a capped mRNA containing an open reading frame encoding the endonuclease. In some cases, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a translated polypeptide.
[0166] In some cases, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus.
[0167] The disclosed system can be used for a variety of applications, such as, for example, nucleic acid editing (e.g., gene editing), binding to nucleic acid molecules (e.g., sequence-specific binding), etc. Such systems can be used, for example, to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivating enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors. EXAMPLES
[0168] Example 1 - Discovery of novel Cas effectors by metagenomics Metagenome mining Metagenomic samples were collected from sediments, soils, and animals. Deoxyribonucleic acid (DNA) was extracted using Zymobiomics DNA Miniprep Kit and sequenced on an Illumina HiSeq® 2500. Samples were collected with the consent of the owners. DNA was extracted from samples using either the Qiagen DNeasy PowerSoil Kit or the ZymoBIOMICS DNA Miniprep Kit. DNA was sent to the Vincent J. Coates Genomics Sequencing Laboratory at UC Berkeley for sequencing library preparation (Illumina TruSeq) and sequencing on an Illumina HiSeq 4000 or Novaseq (paired 150 base pair (bp) reads with target insert sizes of 400-800 bp). Additionally, publicly available high temperature, and soil and marine metagenomic sequencing data were downloaded from NCBI SRA. Sequencing reads were trimmed using BBMap (Bushnell B., sourceforge.net / projects / bbmap / ) and assembled with Megahit (https: / / paperpile.com / c / QSZG6K / clMrh). Protein sequences were predicted with Prodigal (https: / / paperpile.com / c / QSZG6K / BJ6oW). HMM profiles of documented type II CRISPR nucleases were constructed and searched against all predicted proteins using HMMER3 (hmmer.org). CRISPR arrays were predicted with Minced (https: / / github.com / ctSkennerton / minced or https: / / paperpile.com / c / QSZG6K / OPC44) on assembled contigs. Taxonomy was assigned to proteins with Kaiju https: / / paperpile.com / c / QSZG6K / nMi6k and contig taxonomy was determined by finding the consensus of all encoded proteins.
[0169] Predicted and reference (e.g., SpCas9, SaCas9, and AsCas9) type II effector proteins were aligned with MAFFT (https: / / paperpile.com / c / QSZG6K / sVHNH) and phylogenetic trees were inferred using FastTree2 (https: / / paperpile.com / c / QSZG6K / osZNM). Novel families were identified from clades composed of sequences recovered from this study. From within families, candidates were selected if they contained all components for laboratory analysis (i.e., they were found in fully assembled and annotated contigs using CRISPR arrays and predicted tracrRNA). Selected representative sequences and reference sequences were aligned using MUSCLE (https: / / paperpile.com / c / QSZG6K / ITOla) to identify catalytic and PAM-interacting residues.
[0170] This metagenomic workflow resulted in the delineation of the SMART (SMall ARchaeal-associated) endonuclease system described herein.
[0171] Discovery of SMART endonucleases containing active residue signatures Mining tens of thousands of high-quality CRISPR Cas systems assembled from metagenomic data uncovered novel effectors containing both RuvC and HNH domains, but of unusually small size (<900 aa) (Figure 21A). These effector nucleases showed low sequence similarity (<20% amino acid identity) with the archaeal Cas9 endonuclease as a reference point. Phylogenetic analysis of effector protein sequences showed that SMART systems are clades to well-tested type II systems from subtypes A, B, or C (Figure 1A and Figure 21B).
[0172] These compact "SMART" effectors (~400-1000 amino acids, Figure 2) appeared at loci in the genome adjacent to CRISPR arrays. Some of these adjacent SMART loci also contained sequences predicted to encode tracrRNA and CRISPR-adaptive genes (e.g., genes involved in spacer acquisition) cas1, cas2, or cas4 in the same operon (Figure 3 and Figure 21A). Despite their compact size, SMART effectors contain six putative HNH and RuvC catalytic residues when aligned with the reference SaCas9 sequence (Figure 4). In addition, 3D structure predictions identify residues involved in guide and target binding, as well as recognition of the PAM, suggesting that SMART effectors are active dsDNA endonucleases.
[0173] Multiple families of SMART endonucleases Based on the location of key catalytic and binding residues, SMART nucleases contain three RuvC domains, an arginine-rich region that usually contains an RRxRR motif (e.g., a domain with PF14239 homology), an HNH endonuclease domain, and a putative recognition domain (Figures 5 and 6). These domains share low sequence similarity with reference sequences (Figure 7). In addition, SMART effectors, as well as reference archaeal sequences, contain RRxRR and zinc-binding ribbon motifs (CX[2-4]C or CX[2-4]H) significantly more frequently than Cas9 nucleases (Figure 8). In addition, unlike Cas9 effector sequences, most SMART effectors contain significant hits to the Pfam domain PF14239, which is often associated with diverse endonucleases. Based on differences in SMART effector size, phylogenetic relationships, and both operon and domain architecture, we classified these systems into two major groups, SMART I and SMART II. The salient features of these groups are outlined below in Table 3, which also illustrates the differences compared to class 2, type II A / B / C Cas enzymes.
[0174] [Table 3]
[0175] SMART nucleases, like Cas9, contain RuvC and HNH domains, but the RuvC-I, bridge helix, and recognition domains are poorly aligned. To best understand the evolutionary relationship between SMART nucleases and reference sequences, multiple sequence alignments of documented and cataloged full-length SMARTs, reference type II sequences (see, e.g., Burstein, D. et al. New CRISPR-Cas systems from uncultivated microbes. Nature 2017, 542, 237-241, and Gasiunas, G. et al. A catalogue of biochemically diverse CRISPR-Cas9 orthologs. Nat Commun 2020, 11, 5512, each of which is incorporated by reference in its entirety), and >10,300 recently reported Cas9 homologues and IscB sequences (see, e.g., Altae-Tran, H. et al. The widespread IS200 / 605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 2021,374,57-65). The trimmed well-aligned region encompassing the RuvC-II / HNH / RuvC-III domains was retained. Phylogenetic analysis inferred from this final alignment showed a branched clade of effectors that clustered from the documented Cas9 effectors currently classified as II-A, II-B, and II-C (Figure 21E). Two SMART clades found to be phylogenetically close to the classified type II effectors were likely to be encoded adjacent to CRISPR arrays (Figure 21B, Figure 21C, and Figure 21E). The MG33 family of SMART nuclease clusters with type II-C2 effectors significantly expands this clade (Figure 21E and Figure 21F, Mauve branch). This family contains the largest representatives of the SMART enzymes, between 900 and 1050 aa, and their length distribution overlaps with the smallest classified type II-C enzymes.The more distant SMART clade (Figures 21E and 21F, cyan, green, and yellow branches) contains "early Cas9" sequences recently classified as type II-D (Figures 21E and 21F, light grey branches). These CRISPR systems can generally be collectively referred to as SMART.
[0176] SMART I endonuclease The size of SMART I effectors ranges from approximately 600 amino acids to 1,050 amino acids. Common features in their genomic context are predicted tracrRNAs near adaptive module genes (e.g., genes involved in spacer acquisition) and CRISPR arrays, the organization of which was similar to type II and type V CRISPR systems (Figure 3A, Figure 3B, and Figure 3C). The RRXRR motif-containing region in SMART I effectors is unique but may play a similar functional role as the arginine-rich bridge helix in Cas9 nuclease. When modeled against the SaCas9 crystal structure, the predicted 3D structures of SMART I effectors showed non-aligned regions within the recognition lobe (often containing the Pfam domain PF14239) and the RuvC-II domain (Figure 5). The results indicated that these domains have a distinct origin compared to other type II effectors. Taken together with their branched placement in the type II effector phylogenetic tree and their low sequence similarity to documented type II effectors (FIG. 1A and FIG. 21B), these results indicate that SMART I endonucleases belong to a new group of type II CRISPR systems. Following the accepted classification of CRISPR systems, these SMART I systems were classified as type II-D.
[0177] Putative single guide RNAs (sgRNAs) were engineered using environmental RNA expression data from the SMART I MG34-1 line. In addition, multiple sgRNAs designed from SMART I repeats and tracrRNA predictions were tested in vitro in PAM enrichment assays. For SMART I enzymes, optimal identification of PAM sequences was performed at this stage using end repair and blunt end ligation, suggesting that these enzymes may generate staggered double-stranded DNA cleavage. The assay confirmed dsDNA cleavage of MG34-1 (SEQ ID NO: 2), MG34-9 (SEQ ID NO: 9), and MG34-16 (SEQ ID NO: 17) with multiple sgRNA designs (Figure 7, illustrating the use of SEQ ID NOs: 612-615). MG34-1 demonstrated a preference for the NGGN PAM for target recognition and cleavage (Figure 8A), while members of the MG102 family recognize the 3'NRC PAM for target recognition and cleavage (Figure 21C). Analysis of the cleavage sites showed preferential cleavage at position 7 (Figure 8B and Figure S22A). These results suggest a novel biochemical mechanism compared to cleavage mechanisms from other Type II enzymes, which preferentially cleave at positions 2-3 from the PAM, and support a new classification of SMART I CRISPR systems.
[0178] Environmental expression data for several SMART I systems confirmed the in situ transcription of the CRISPR array and the intergenic region encoding the predicted tracrRNA (Figure 3B and Figure 3C). In addition, cases of active CRISPR targeting were assessed by searching for spacer sequences matching other genome sequences assembled from the same or related metagenomes. Along these lines, we identified a phage genome (Figure 3C and Figure 3D) targeted by one of the spacers encoded in the SMART I CRISPR array. Analysis of the regions flanking the target sequence suggests a 3'PAM sequence containing a GG motif (Figure 3D). These results indicate that SMART I CRISPR systems are active as RNA-guided effectors involved in phage defense in their native environment and likely function as nucleases to cleave or degrade targeted DNA or RNA.
[0179] SMART I effectors are active RNA-guided dsDNA CRISPR endonucleases Putative single guide RNAs (sgRNAs) were engineered using environmental RNA expression data from the SMART I MG34-1 and MG34-16 systems (Figure 3B and Figure 3C, and Figure 9). In addition, multiple sgRNAs designed from SMART I repeats and tracrRNA predictions were tested in vitro in a PAM enrichment assay (Figure 10). The assay confirmed programmable dsDNA cleavage of MG34-1, MG34-9, and MG34-16 with multiple sgRNA designs (Figure 10). MG34-1 and MG34-9 require the NGGN PAM for target recognition and cleavage (Figure 11A and Figure 11C). Analysis of cleavage sites showed preferential cleavage at position 7 (Figure 11B and Figure 11C). These results suggest a novel biochemical cleavage mechanism compared to the Cas9 enzyme, which preferentially cleaves at position 3 from the PAM, providing further support for the new classification of SMART I CRISPR systems.
[0180] PAM enrichment assays without the end-repair procedure showed no activity against SMART I nucleases. The requirement of end-repair to create blunt-ended fragments prior to ligation in the PAM enrichment protocol indicates that these enzymes create staggered double-stranded DNA breaks. Staggered double-stranded breaks were confirmed by sequencing of the cleavage products of MG34-1 nuclease (Figure 22A). These results suggest a novel biochemical mechanism compared to mechanisms from most documented type II enzymes, which preferentially cleave at positions 2-3 from the PAM. In vitro cleavage assays with purified proteins show that MG34-1 is more efficient at targeted DNA cleavage at a target guide 18 bp long, and time-series cleavage assays show that MG34-1 cleaves at a slower rate compared to the reference SpCas9 when tested with the same guide (Figures 22B and 22C).
[0181] Experiments performed in E. coli showed that the system had the necessary activity to function as a nuclease in cells. E. coli strains expressing MG34-1 and MG34-9 sgRNAs were transformed with a kanamycin resistance plasmid containing the sgRNA target. In the presence of antibiotic, successful targeting and cleavage of the antibiotic resistance plasmid would result in a growth defect. The assay showed approximately 2- to 10-fold growth inhibition compared to control experiments performed with a kanamycin resistance plasmid that did not contain the sgRNA target (Figures 12 and 22D).
[0182] SMART II Endonuclease SMART II effectors have a small (~400-600 amino acids) size distribution compared to SMART I effectors. Their genomic context suggested unusual repeat regions or CRISPR arrays. Non-CRISPR repeat regions contain direct repeats ranging in size from ~10-30 bp or more. In some cases, they contain multiple separate repeat units. Sometimes, common CRISPR identification algorithms would flag these regions as CRISPR-based, but closer inspection reveals that regions identified as spacer sequences are repeated within the arrays. Although the arrays are not immediately adjacent to the effectors, they are in the same genomic region (Figure 3A, MG35-236, and Figure 13A, e.g., >20 kb from the effector genes). SMART II-based operons were generally devoid of adaptive module genes (e.g., genes involved in spacer acquisition).
[0183] Structural predictions often identified hallmark residues of Cas enzymes involved in guide RNA binding, target cleavage, and recognition and interaction with the PAM, in addition to all six RuvC and HNH nuclease catalytic residues found in class 2, type II Cas effectors (Figure 6). In addition, SMART II effectors contain multiple RRXRR and zinc-binding ribbon motifs (CXRR) that may be involved in recognition and binding to target nucleic acid motifs. [2-4] C or CX[2-4] Based on the locations of key residues, the predicted domain structure of SMART II nucleases included three RuvC subdomains: an arginine-rich region containing an RRxRR motif (e.g., a domain with PF14239 homology), an HNH endonuclease domain, an unknown domain, and a recognition domain (REC) (Figure 6). The domain architecture of SMART II effectors differed from the documented domain structures of type II Cas9 nucleases (Figures 6 and 14).
[0184] Environmental transcriptome data for several SMART II systems confirmed the in situ expression of CRISPR arrays and other repeat regions in the native environment (Figure 13A). Transcription of the 5' untranslated regions (UTRs) of several SMART II effectors was also observed from the environmental expression data (Figure 13B and Figure 16), suggesting that this region may be important for either nuclease activity or regulation of the SMART system.
[0185] Preliminary in vitro experiments performed with SMART II effector proteins, repeat regions, and associated intergenic regions indicate that these enzymes have the ability to cleave dsDNA, likely in a programmable manner (Figures 15 and 17). The results suggest that SMART II nuclease activity may be RNA or DNA guided, which may be required using repeat regions such as CRISPR arrays or through recognition of features encoded within gene loci such as TIRs or 5'UTRs. The 5'UTRs of SMART II effectors are actively transcribed in in vitro transcription assays and exhibit high secondary structure (Figure 18). Multiple sequence alignment of the region immediately upstream of the start codon of SMART II effectors demonstrates a conserved block (Figure 19), suggesting that the 5'UTRs associated with SMART II effectors encode an RNA guide for the effector to target DNA for cleavage activity.
[0186] Recently, short Cas9 homologs were reported to be programmable dsDNA nucleases using guide RNAs encoded in the 5'UTR region of the effector (Altae-Tran, Kannan, et al. Science 2021). In these systems, targeting "spacers" were identified upstream of the transcribed 5'UTR of the effector, suggesting that the SMART II enzyme could be reprogrammed to target and cleave specific DNA sites by adding a "targeting spacer" to the 5' end of the predicted guide RNA encoded in its 5'UTR. Adding a targeting spacer to the 5' end of the guide RNA encoded in the 5'UTR region of the SMART II effector activated the effector for targeted dsDNA cleavage using various target adjacent motifs (TAMs) (Figure 20).
[0187] Several SMART II effectors were observed next to putative insertion sequences (IS) encoding the transposases TnpA and TnpB (Figure 3A). The ends of the IS were identified as containing terminal inverted repeats (TIRs) with predicted hairpin structures, and target site duplications into which the IS most likely integrated were also identified. In addition, several SMART II loci encoded putative TIRs adjacent to SMART II effectors (e.g., Figure 3).
[0188] The SMART HEARO clade contains virus-associated RNA-guided dsDNA nucleases Phylogenetic analysis shows that SMART nucleases less than 600 aa in length (Figure 21E, light purple branch) cluster together with documented IscB sequences ("insertion sequence Cas9-like" (see, e.g., Kapitonov, VV, Makarova, KS & Koonin, EVISC, a Novel Group of Bacterial and Archaeal DNA Transposons That Encode Cas9 Homologs. J Bacteriol 2016, 198, 797-807, incorporated herein by reference in its entirety)) (Figure 21E, dark grey branch) forming two major clades. Kapitonov et al. reported IscB homology with Cas9 based on the presence of RuvC and HNH domains, and subsequently described the PLMP domain in this same group of enzymes (e.g., Altae-Tran, H. et al. The widespread IS200 / 605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 2021, 374, 57-65, incorporated herein by reference in its entirety). Using 3D structure prediction, we showed that these proteins typically contain an arginine-rich region that contains an RRXRR motif. The arginine-rich region was suggested to be similar to the bridge helix of Cas9, but neither this region nor the RuvC-I domain was found to align well in 3D space with the bridge helix and RuvC-I domain of the reference 3D structure. Such IscB / SMART enzymes lack a PAM interaction domain. Instead, a C-terminal "WED / REC" domain containing a Zn-binding ribbon motif may be involved in target motif recognition. Although the protein domains, catalytic residues, and 3D models suggest an evolutionary relationship with Cas9, most IscB / SMART effectors are not CRISPR-associated (e.g., they are not found proximal to CRISPR repeats in their genomic context). The group that includes the IscB / SMART system is generally compact in size (approximately 400-600 aa) and widely distributed in bacterial and archaeal genomes.More than 16% of the genomic fragments encoding these effectors were classified as likely to be of viral or prophage origin, implicating viruses in the evolution of these systems.
[0189] A search for noncoding RNAs (ncRNAs) associated with the SMART system found that 65% of IscB / SMART 5' untranslated regions (UTRs) contained hits to HNH endonuclease-associated RNAs and ORFs (HEARO) RNAs from the RFam database (RF02033). These ncRNAs were first described as highly structured RNAs from bioinformatics analysis (see, e.g., Weinberg, Z., Perreault, J., Meyer, MM & Breaker, RR Exceptional structured noncoding RNAs revealed by bacterial metagenome analysis. Nature 2009, 462, 656-659, which is incorporated by reference in its entirety), but the functions of their associated HEARO ORFs had not been reported (see, e.g., Harris, KA & Breaker, RR Large Noncoding RNAs in Bacteria. Microbiol Spectr 2018, 6, which is incorporated by reference in its entirety). The putative HEARO HNH endonuclease ORF was also found to contain RuvC and HNH catalytic domains and to cluster with IscB / SMART effectors. Thus, IscB, small SMART, and HEARO ORFs represent a large group of non-Cas endonucleases. Recently, it has been reported that the 5'UTR of IscB encodes a single guide RNA required for dsDNA nuclease activity, which the authors refer to as an omega RNA (see, e.g., Altae-Tran, H. et al. The widespread IS200 / 605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 2021, 374, 57-65, which is incorporated herein by reference in its entirety).To confirm the requirement of guide RNA for function, we observed in situ native expression of the 5'UTR of the IscB / SMART / HEARO system, recapitulated by in vitro transcription assays. The omega RNA structure shares high structural similarity with HEARO RNAs. In recognition of the unifying features of the IscB / SMART / HEARO system (broad taxonomic origins and enrichment of arginine residues) and the chronological discovery of guide RNAs associated with these enzymes, we propose a broad functional classification of the IscB / SMART / HEARO system as SMART HEARO (Figure 21E). We identified the required targeting motifs by evaluating SMART HEARO cleavage activity in vitro and reprogramming the 5' "spacer" region of their HEARO RNAs as described by Altae-Tran and Kannan et al. (Figure 21D) (see, e.g., Altae-Tran, H. et al. The widespread IS200 / 605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 2021, 374, 57-65). Furthermore, plasmid interference assays in E. coli show that SMART HEARO nucleases are highly active compared to SpCas9 (>570-fold suppression against MG35-1 vs. approx. 98-fold suppression shown by SpCas9, Figure 25B), and specificity experiments show low tolerance to mismatches in the protospacer (Figure 25D).
[0190] Example 2 - PAM sequence identification / validation of endonucleases described herein Putative SMART endonucleases were expressed in an E. coli lysate-based expression system (PURExpress, New England Biolabs), in which the endonucleases were codon-optimized for E. coli and cloned into a vector with a T7 promoter and a C-terminal His tag. The genes were PCR amplified with primer binding sites 150 bp upstream and downstream of the T7 promoter and terminator sequences, respectively. The PCR products were added to NEB PURExpress at a concentration of 5 nM and expressed at 37° C. for 2 hours to produce the endonucleases for PAM assays.
[0191] Putative sgRNAs compatible with each SMART Cas enzyme described herein were identified from RNAseq reads assembled into contig CRISPR loci assembled from sequencing data, secondary structures were determined for tracr regions from RNAseq data along with repeat sequences from CRISPR arrays in the Geneious software package (https: / / www.geneious.com), and the resulting helices were trimmed and ligated with GAAA tetraloops. Multiple lengths of repeat-anti-repeat helix trimming were tested, as well as different spacer lengths and different tracr termination points (Figure 12, SEQ ID NOs: 612-615 are demonstrated). Each sgRNA was then assembled via assembly PCR, purified with SPRI beads, and in vitro transcribed (IVT) according to the manufacturer's recommended protocol for short RNA transcripts (HiScribe T7 kit, NEB). RNA transcription reactions were cleaned with Monarch RNA kit and checked for purity via Tapestation (Agilent).
[0192] The PAM sequence was determined by sequencing a plasmid containing randomly generated potential PAM sequences that could be cleaved by the putative nuclease. In this system, an E. coli codon-optimized nucleotide sequence encoding a putative nuclease was transcribed and translated in vitro from a PCR fragment under the control of a T7 promoter. A second PCR fragment carrying a minimal CRISPR array consisting of a T7 promoter followed by a repeat-spacer-repeat sequence was transcribed in the same reaction. Successful expression of the endonuclease and the repeat-spacer-repeat sequence in the TXTL system followed by CRISPR array processing provides an active in vitro CRISPR nuclease complex.
[0193] A library of target plasmids containing spacer sequences (potential PAM sequences) matching the minimal array preceded by 8N mixed degenerate bases was synthesized using the output of a TXTL reaction (5-fold dilution of translated Cas enzyme, 5 nM 8N PAM plasmid library, and 50 nM sgRNA targeting the PAM library in 10 mM Tris pH 7.5, 100 mM NaCl, and 10 mM MgCl). 2) after 1–3 h. The reaction was stopped and DNA was recovered via a DNA clean-up kit. The adapter sequence was blunt-end ligated to DNA with an active PAM sequence that had been cleaved by the endonuclease, whereas uncleaved DNA was inaccessible for ligation. The DNA segment containing the active PAM sequence was then amplified by PCR using primers specific for the library and the adapter sequence. The PCR amplification products were resolved on a gel to identify amplicons corresponding to the cleavage events. The amplified segments of the cleavage reaction were also used as templates for the preparation of NGS libraries or as substrates for Sanger sequencing. Sequencing this resulting library, a subset of the starting 8N library, revealed sequences with PAM activity compatible with the CRISPR complex. For PAM testing with treated RNA constructs, the same procedure was repeated, except that in vitro transcribed RNA was added along with the plasmid library and the minimal CRISPR array / tracr template was omitted. In these assays, the following spacer sequence was used as the target (5'-CGUGAGCCACCACGUCGCAAGCCUCGAC-3').
[0194] After obtaining raw sequencing reads from the PAM assay, the reads were filtered by Phred quality score >20. 24 bp representing the documented DNA sequence from the scaffold adjacent to the PAM were used as a reference to find the PAM-proximal region, and the adjacent 8 bp were identified as putative PAMs. The distance between the PAM and the ligation adapter was also measured for each read. Reads that did not exactly match the reference sequence or the adapter sequence were excluded. PAM sequences were filtered by cleavage site frequency so that PAMs with the most frequent cleavage site ±2 bp were selectively included in the analysis. The filtered list of PAMs was used to generate sequence logos using Logomaker (Tareen A, Kinney JB. Logomaker: beautiful sequence logos in Python. Bioinformatics. 2020;36(7):2272-2274, incorporated herein by reference).
[0195] Example 3 - Predicted RNA folding protocol Calculate the predicted RNA folding of active single RNA sequences at 37° using the method of Andronescu 2007. The color of the base corresponds to the probability of base pairing for that base, with red being high probability and blue being low probability.
[0196] Example 4 - In vitro cleavage efficiency The endonuclease is expressed as a His-tagged fusion protein from an inducible T7 promoter in a protease-deficient E. coli B strain. The endonuclease was fused N-to-C-terminally with two nuclear localization signals (N-terminal NLS nucleoplasmin bipartite and C-terminal Simian Virus 40 T antigen NLS PPKKKRK), a maltose binding protein (MBP) tag, a tobacco etch virus (TEV) protease cleavage site, and a 6XHis tag in the following order: 6XHis-MBP-TEV-NLS-gene-NLS-STOP. The protein was expressed under the pTac promoter in NEB Iq E. coli with autoinduction medium (MagicMedia ThermoFisher), grown at 30°C, and induced at 16°C.
[0197] Cells expressing His-tagged proteins were lysed by sonication and His-tagged proteins were purified by Ni-NTA affinity chromatography on a HisTrap FF column (GE Lifescience) on an AKTA Avant FPLC (GE Lifescience). The eluates were separated by SDS-PAGE on acrylamide gels (Bio-Rad) and stained with InstantBlue Ultrafast Coomassie (Sigma-Aldrich). Purity was determined using densitometry of the protein bands with ImageLab software (Bio-Rad). The purified endonucleases were dialyzed into a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol, pH 7.5, and stored at -80°C.
[0198] A target DNA containing a spacer sequence and a PAM sequence (e.g., as determined in Example 2) was constructed by DNA synthesis. A single representative PAM was selected for testing when the PAM has degenerate bases. The target DNA consisted of 2200 bp linear DNA derived from a plasmid via PCR amplification with the PAM and a spacer located 700 bp from one end. Successful cleavage results in fragments of 700 and 1500 bp. The target DNA, in vitro transcribed single RNA, and purified recombinant protein were mixed in cleavage buffer (10 mM Tris, 100 mM NaCl, 10 mM MgCl) containing excess protein and RNA. 2 ) and incubated for 5 min to 3 h, typically 1 h. The reaction is stopped via the addition of RNAse A and a 60 min incubation. The reactions are then resolved on a 1.2% TAE agarose gel and the fraction of cleaved target DNA is quantified with ImageLab software.
[0199] Example 5 - Activity in E. coli E. coli lacks the ability to efficiently repair double-stranded DNA breaks. Thus, genomic DNA breaks can be a lethal event. Taking advantage of this phenomenon, endonuclease activity is tested in E. coli by recombinantly expressing the endonuclease and guide RNA in a target strain that has the spacer / target sequence and PAM sequence integrated into its genomic DNA.
[0200] For testing of nuclease activity in bacterial cells, BL21(DE3) strain (NEB) was transformed with plasmids containing T7-driven effectors and sgRNAs (10 ng of each plasmid), plated, and grown overnight. The resulting colonies were grown overnight in triplicate, then subcultured in SOB and grown to an OD of 0.4-0.6. 0.5 OD equivalents of the cell culture were made chemically competent according to standard kit protocols (Zymo Mix and Go kits) and transformed with 130 ng of kanamycin plasmids, either with or without spacers and PAMs in the backbone. After heat shock, transformants were allowed to recover for 1 h at 37 °C in SOC, and nuclease efficiency was determined by a 5-fold dilution series grown on induction medium (LB agar plates with antibiotics and 0.05 mM IPTG). Colonies were quantified from the dilution series to measure the overall suppression due to nuclease-driven plasmid cleavage.
[0201] The results of such an assay are shown in FIG. 12. In FIG. 12, panel (A) shows replica plating of E. coli strains demonstrating plasmid cleavage, where E. coli expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid (+sp) containing the target of the sgRNA. Plate quadrants showing growth impairment (+sp) versus negative control (no target and PAM (-sp)) indicate successful targeting and cleavage by the enzyme. Experiments were repeated twice and performed in triplicate. In FIG. 12, panel B shows a graph of colony forming unit (cfu) measurements from the replica plating experiment in A, showing growth inhibition in the targeted condition (+sp) versus the non-targeted control (-sp), demonstrating that the plasmid was cleaved. In FIG. 12, panel C shows a bar plot of colony forming unit (cfu) measurements (log scale) showing E. coli growth inhibition in the targeted condition (white bars) versus the non-targeted control (green bars) for various SMART nucleases. Plasmid interference assays for each nuclease were performed in triplicate along with an SpCas9 positive control.
[0202] The engineered strain with the PAM sequence (e.g., determined as in Example 2) integrated into the genomic DNA is transformed with DNA encoding the endonuclease. The transformants are then made chemically competent and transformed with 50 ng of guide RNA (e.g., crRNA) either specific for the target sequence ("on-target") or non-specific for the target ("non-target"). After heat shock, the transformants are allowed to recover in SOC at 37°C for 2 hours. Nuclease efficiency is then determined by 5-fold dilutions grown on induction medium. Colonies are quantified from the dilution series in triplicate.
[0203] Example 6 - Testing genome cleavage activity of MG CRISPR complexes in mammalian cells To demonstrate targeting and cleavage activity in mammalian cells, the MG Cas effector protein sequence is tested in two mammalian expression vectors, (a) one with a C-terminal SV40 NLS and a 2A-GFP tag, and (b) one without a GFP tag and with two SV40 NLS sequences, one on the N-terminus and one on the C-terminus. The NLS sequences include any of the NLS sequences described herein. In some cases, the nucleotide sequence encoding the endonuclease is codon-optimized for expression in mammalian cells.
[0204] The corresponding crRNA sequence with the targeting sequence added is cloned into a second mammalian expression vector. The two plasmids are co-transfected into HEK293T cells. 72 hours after co-transfection of the expression plasmid and the gRNA targeting plasmid into HEK293T cells, DNA is extracted and used for preparation of NGS libraries. NHEJ percent is measured via indels in the sequencing of the target site to indicate the targeting efficiency of the enzyme in mammalian cells. At least 10 different target sites are selected to test the activity of each protein.
[0205] Example 7 - Predicted activity of the MG family as described herein In situ expression and protein sequence analysis indicate that these enzymes are active nucleases. They contain predicted endonuclease-associated domains (matching the RRXRR and HNH endonuclease Pfam domains, Figures 2, 3A, and 3B) and contain predicted HNH and RuvC catalytic residues (Figures 2, 3A, and 3B, rectangles). Furthermore, the presence of the RRXRR motif found in the ribonuclease H-like protein family indicates potential RNA targeting or nuclease activity (see Figure 2).
[0206] Expression data confirms the in situ native activity of the candidate MG34-1 nuclease, tracrRNA, and CRISPR arrays (Figure 4).
[0207] Example 8 - Activity in mammalian cells following mRNA delivery For genome editing using cell transfection / transformation with mRNA, the coding sequence is mouse or human codon optimized using algorithms from Twist Bioscience or Thermo Fisher Scientific (GeneArt). The cassette is constructed with two nuclear localization signals added to the coding endonuclease sequence: SV40 and nucleoplasmin at the N-terminus and C-terminus, respectively. In addition, untranslated regions from human complement 3 (C3) are added to both the 5' and 3' of the coding sequence in the cassette.
[0208] This cassette is then cloned into an mRNA production vector upstream of a long polyA stretch. The composition of the mRNA construct can be as follows: 5'UTR-SV40 NLS from C3-codon-optimized SMART gene-nucleoplasmin NLS-3'UTR-107 polyA tail from C3. The execution of transcription of the mRNA is then driven by a T7 promoter using an engineered T7 RNA polymerase (Hi-T7: New England Biolabs). 5'-capping of the mRNA occurs co-transcriptionally using CleanCap AG (Trilink Biolabs). The mRNA is then purified using the MEGAclear Transcription Clean-Up Kit (Thermo Fisher Scientific).
[0209] Mammalian cells are co-transfected with the transcribed mRNA and a set of at least 10 guides targeting the genomic region of interest using Lipofectamine Messenger Max (Thermo Fisher Scientific). The cells are incubated for a period (e.g., 48 hours), followed by genomic DNA isolation using Purelink Genomic DNA Extraction Kit (Fisher Scientific). The region of interest is amplified using specific primers. Editing is then assessed by Sanger sequencing using CRISPR editing and NGS inference for a complete analysis of the editing results.
[0210] Example 9 - SMART II guide RNA prediction A region encompassing 400 bp immediately upstream of the start codon of the SMART II effector sequence was extracted as potentially encoding a guide RNA required for activity (UTR). UTR sequences were aligned with MAFFT (mafft-ginsi algorithm), and regions showing conserved blocks were annotated as putative guide RNAs.
[0211] Example 10 - Activity and PAM determination assays Putative guide RNAs predicted from RNASeq or from UTR alignments were folded in Geneious. Single guide RNAs (sgRNAs) were designed by adding a target spacer to either the 5' or 3' end of the guide RNA. sgRNAs were assembled via assembly PCR, purified with SPRI beads, and in vitro transcribed (IVT) according to the manufacturer's recommended protocol for short RNA transcripts (HiScribe T7 kit, NEB). RNA reactions were cleaned with Monarch RNA kit and checked for purity via Tapestation (Agilent).
[0212] Cleavage and PAM determination assays were performed using PURExpress (New England Biolabs). Briefly, proteins were codon-optimized for E. coli and cloned into a vector with a T7 promoter and a C-terminal His tag. Genes were PCR amplified with primer binding sites 150 bp upstream and downstream of the T7 promoter and terminator sequences, respectively. The PCR products were added to NEB PURExpress at a concentration of 5 nM and expressed at 37°C for 2 hours. After this, 5-fold dilutions of PURExpress, 5 nM of the 8N PAM plasmid library, and 50 nM of sgRNA targeting the PAM library were used to incubate the plasmids in 10 mM Tris pH 7.5, 100 mM NaCl, and 10 mM MgCl. 2 The cleavage reaction was assembled in
[0213] Cleavage products from the PURExpress reaction were recovered via cleanup with AMPure SPRI beads (Beckman Coulter). DNA was blunted via the addition of Klenow fragment and dNTPs (New England Biolabs). Blunt-ended products were ligated with 100-fold excess of double-stranded adapter sequences and used as templates for the preparation of NGS libraries, from which PAM requirements were determined by sequence analysis.
[0214] Raw NGS reads were filtered by Phred quality score >20. 24 bp representing the documented DNA sequence from the scaffold adjacent to the PAM were used as a reference to find the PAM-proximal region, and the adjacent 8 bp were identified as putative PAMs. The distance between the PAM and the ligation adapter was also measured for each read. Reads that did not exactly match the reference or adapter sequence were excluded. PAM sequences were filtered by cleavage site frequency such that PAMs with the most frequent cleavage site ±2 bp were selectively included in the analysis. The filtered list of PAMs was used to generate sequence logos using Logomaker.
[0215] Example 11 - SMART Amino Acid Composition To account for the amino acid composition of the SMART protein sequences, the amino acid content for the groups of SMART sequences was calculated as the number of times each residue was observed divided by the total protein length multiplied by 100. The amino acid composition was then compared to the reported content in a large set of protein sequences from the Uniprot50 database (Carugo, Protein Sci. 2008). Both protein groups, SMART HEARO and SMART (type II-D), contain unusually high arginine and lysine amino acid content compared to the content observed in the Uniref50 protein sequences (Figure 23).
[0216] On average, the arginine and lysine composition ratios of SMARTs deviate from the linear trends observed for other residues in SMART sequences, as well as from the residue composition of proteins in the Uniref50 database (Figure 24A). In addition, the methionine content of SMARTs was observed to be statistically lower than that observed in proteins from the Uniref50 database (Figure 24B).
[0217] To describe the physicochemical properties of SMART, the isoelectric point, molecular weight, and charge were determined from the sequences with the "protr" and "peptide" packages in R. The high arginine and lysine content observed in the SMART sequences may contribute to the high isoelectric point and charge at neutral pH (Table 4).
[0218] [Table 4]
[0219] The high arginine and Zn-binding ribbon motif content of SMART nucleases suggests that these enzymes may contain intrinsically disordered regions, which may add flexibility to the protein to interact with large guide RNAs and target DNA. Intrinsically disordered regions are segments of a protein that lack stable tertiary structure in its native unbound state (see, e.g., Bitard-Feildel, T., Lamiable, A., Mornon, J.-P. & Callebaut, I. Order in Disorder as Observed by the “Hydrophobic Cluster Analysis” of Protein Sequences. Proteomics 2018, 18, e1800054, incorporated herein by reference in its entirety), and may be enriched in positively charged arginines that interact with polyanions (such as RNA) (see, e.g., Murthy, A. C. et al. Molecular interactions underlying liquid-liquid phase separation of the FUS low complexity domain. Nat Struct Mol Biol 2018, 18, e1800054, incorporated herein by reference in its entirety). 2019, 26, 637-648), and may be found as linkers between Zn-binding ribbons, aiding in the "search function" (see, e.g., Dyson, H. J. Roles of intrinsic disorder in protein-nucleic acid interactions. Mol Biosyst 2011, 8, 97-104, incorporated herein by reference in its entirety), all of which are features observed in SMART nucleases.
[0220] Example 12 - Mismatch killing assay To determine the specificity of the various SMART enzymes, a mismatch killing assay was developed in which E. coli BL21(DE3) strains (NEB) were transformed with plasmids containing T7-driven effectors (ampicillin resistance) and their T7-driven sgRNAs (chloramphenicol resistance), plated, and grown overnight. The resulting colonies were made competent and transformed with 100 ng of kanamycin plasmid in three conditions: a library of 25 plasmids each containing a single mismatch along the target spacer and PAM in the backbone, a 24 nt spacer and a constant PAM, or a control plasmid without spacer or PAM (Figure 25D). After heat shock, transformants were allowed to recover in SOC medium for 2 hours at 37°C. Cultures were plated and grown overnight at 37°C on induction medium (LB agar plates with antibiotics and 0.05 mM IPTG). Plasmids were extracted from surviving mismatched colonies via a miniprep kit (Qiagen). The target regions were amplified via PCR and analyzed via NGS. The spacers enriched relative to the untreated library cannot be recognized and cleaved by the nuclease, and are therefore considered to be regions where the effector does not tolerate mismatches. If mismatches were tolerated, the enzyme would be expected to cleave the antibiotic resistance plasmid, and growth defects would be observed. It was observed that the MG102-2 nuclease did not tolerate mismatches along the first 13 positions of the target plasmid from the PAM, while variable mismatch tolerance was observed from position 14 onwards (Figure 25D and Figure 27). These results suggest that SMART nucleases may be highly specific and do not exhibit concomitant ssDNA cleavage (Figure 28).
[0221] Example 13 – Human cell editing with SMART nuclease MG102-2 K562 cells (ATCC) were cultured according to the ATCC protocol. Two sgRNAs targeting the TRAC locus were designed based on the MG102-2 PAM and chemically synthesized by IDT. For gene editing experiments, 500 ng of in vitro synthesized MG102-2 mRNA and either 150, 300, or 450 pmol of the indicated sgRNA were transfected into 1.5 × 10 cells using a Lonza 4D Nucleofector (program FF-120). 5 In parallel, cells were nucleofected without mRNA or guide to assess background at sites targeted by TRAC guides. Cells were harvested 72 hours after electroporation for genomic DNA extraction using QuickExtract (Lucigen, #09050) and processed for next generation sequencing on an Illumina Miseq. The resulting data were analyzed using an indel calculation script.
[0222] Delivery of SMART nuclease via mRNA into human cells targeting the T cell receptor alpha constant locus (TRAC) resulted in over 90% editing activity at one of the two TRAC target sites with MG102-2 nuclease (Figure 26). As observed in in vitro experiments (Figure 29), increasing the amount of sgRNA improved editing efficiency at both target loci (Figure 26). Localization of the MG34-1 system (fused to a nuclear localization signal, NLS) to the nucleus of human cells was confirmed, but no nuclease-induced indel formation was detected for this nuclease.
[0223] Example 14 - Cleavage preferences of SMART nucleases Sequencing the cleavage products of MG34-1 and MG102-2 nucleases shows that these enzymes create staggered double-stranded DNA breaks (Figure 25A). Analysis of the cleavage sites shows selective cleavage at positions 5-7 from the PAM (Figure 25A). These results suggest a biochemical cleavage mechanism rarely observed compared to most Cas9 enzymes, which creates blunt ends and staggered cuts with a preference for positions 3-5 from the PAM. In vitro transcription / translation reactions and in vitro cleavage assays using purified proteins show that MG34-1 and MG102-2 are most efficient with 18 and 20 nucleotide spacers (Figure 25C). Furthermore, activity was confirmed in vivo using an E. coli plasmid interference assay, showing growth inhibition of 2-fold (MG34-9) to >500-fold (MG102-2) for five SMART nucleases with the indicated targeted spacers (Figure 25B).
[0224] Example 15 - SMART I enzyme is an active nuclease in human cells K562 cells purchased from ATCC were cultured according to the ATCC protocol. sgRNAs targeting the TRAC or AAVS1 locus were designed based on the PAMs recognized by MG102-2, MG102-36, MG102-39, MG102-42, MG102-45, and MG33-34 and chemically synthesized by IDT. For gene editing experiments, 500 ng of in vitro synthesized nuclease mRNA and 450 pmol of the indicated sgRNA were transfected into 1.5 × 10 cells using a Lonza 4D Nucleofector (program FF-120). 5 Cells were harvested 72 hours after electroporation for genomic DNA extraction using QuickExtract (Lucigen, #09050) and processed for amplicon next-generation sequencing on an Illumina Miseq. The resulting data were analyzed using an in-house indel calculation script.
[0225] As described elsewhere herein, SMART I nuclease MG102-2 is active at two target sites in the TRAC locus of the human genome when delivered via mRNA.MG102-2 (SEQ ID NO:582) is also active at the AAVS1 locus (safe harbor locus) in the human genome, and the cleavage efficiency of the enzyme is as high as 82.6% at eight different target sites, further confirming an editing efficiency of >50% (Figure 30A). In addition, MG102-39 (sequence number 993), MG102-42 (sequence number 996), and MG102-48 (sequence number 1002) showed >40% cleavage activity at the TRAC locus of the human genome when delivered by mRNA (Figures 30B-D), while MG33-34 (sequence number 988), MG102-36 (sequence number 990), and MG102-45 (sequence number 999) showed cleavage efficiency above background (10%) at the TRAC locus (Figures 30E-G).
[0226] [Table 5-1]
[0227] [Table 5-2]
[0228] [Table 5-3]
[0229] [Table 5-4]
[0230] [Table 5-5]
[0231] [Table 5-6]
[0232] [Table 5-7]
[0233] [Table 5-8]
[0234] [Table 5-9]
[0235] Example 16 - SMART HEARO enzymes are active nucleases In silico prediction of SMART HEARO guide RNAs
[0236] To identify guide (HEARO) RNAs associated with novel SMART HEARO nucleases, nucleotide sequences corresponding to the 5'UTR regions of 305 putative effectors were extracted. These 5'UTR nucleotide sequences were aligned with MAFFT (Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013, 30(4), 772-780, incorporated herein by reference in its entirety) using the parameters mafft-xinsi (https: / / mafft.cbrc.jp / alignment / software / ), and conserved regions were used to delineate HEARO RNA boundaries (Figure 31). In addition, the HEARO RNA sequences of active SMART HEARO nucleases were used to generate covariance models to predict additional HEARO RNAs in the genomic fragments encoding novel SMART HEARO nucleases. Covariance models are constructed from multiple sequence alignments (MSA) of active HEARO RNA sequences with mafft-xinsi (https: / / mafft.cbrc.jp / alignment / software / ). Secondary structures of the MSA were determined with RNAalifold (Vienna Package, https: / / www.tbi.univie.ac.at / RNA / ), and covariance models were constructed with the Infernal package (http: / / eddylab.org / infernal / ). Contigs and 305 5'UTR regions containing candidate SMART HEARO nucleases were searched using the covariance models with the Infernal command "cmsearch". HEARO RNAs predicted from the 5'UTR alignment and from the covariance models of novel candidates were tested in vitro.
[0237] In vitro TAM determination assay sgRNAs with targeting spacers at the 5' end (HEARO RNAs) were constructed via assembly PCR, purified with SPRI beads or ordered as gene fragments (IDT), and then in vitro transcribed (IVT, HiScribe T7 kit, New England Biolabs) according to the manufacturer's recommended protocol for short RNA transcripts. RNA reactions were cleaned with Monarch RNA kit and checked for purity via Tapestation (Agilent). Cleavage and TAM determination assays were performed using PURExpress (New England Biolabs). Briefly, proteins were codon-optimized for E. coli and cloned into vectors with a T7 promoter and a C-terminal His tag. Genes were PCR amplified with primer binding sites 150 bp upstream and downstream of the T7 promoter and terminator sequences, respectively. The PCR products were added to PURExpress (New England Biolabs) at a final concentration of 5 nM and expressed for 2 hours at 37 °C. Using a 5-fold dilution of PURExpress, 5 nM 8N PAM plasmid library, and 50 nM sgRNA targeting the PAM library, prepare a 5x 50-well plate in 10 mM Tris pH 7.5, 100 mM NaCl, and 10 mM MgCl. 2Cleavage reactions were assembled in a centrifuge. Cleavage products from PURExpress reactions were recovered via cleanup with SPRI beads (AMPure Beckman Coulter or HighPrep Sigma-Aldritch). DNA was blunted via the addition of Klenow fragment and dNTPs (New England Biolabs). Blunt-end products were ligated with 100-fold excess of double-stranded adapter sequences and used as templates for preparation of NGS libraries, from which PAM requirements were determined from sequence analysis. Raw NGS reads were filtered by Phred quality score >20. Using 14–24 bp representing documented DNA sequence from the backbone adjacent to the PAM as a reference, PAM-proximal regions were found and the adjacent 8 bp were identified as putative target adjacent motifs (TAMs). The distance between the TAM and the ligated adapter was also measured for each read. Reads that did not exactly match the reference or adapter sequences were excluded. TAM sequences were filtered by cleavage site frequency such that only TAMs with the most frequent cleavage site ±2 bp were included in the analysis. The filtered list of TAMs was used to generate sequence logos using Logomaker (Tareen, A. & Kinney, J.B. Logomaker: beautiful sequence logos in Python. Bioinformatics 2020, 36, 2272-2274).
[0238] SMART II (HEARO) effectors are short (approximately 400-600 aa long) nucleases that interact with guide (HEARO) RNAs encoded in their 5'UTR regions for targeted dsDNA cleavage (Figure 32A and Figure 32D). In most cases, SMART HEARO systems are not CRISPR-associated, but a few SMART HEARO nucleases can be associated with CRISPR. For example, SMART HEARO MG35-463 (SEQ ID NO: 530) is encoded downstream of a CRISPR array (Figure 32B). The 5' end of the HEARO guide RNA predicted from the covariance model overlaps with the last CRISPR repeat of the array (Figure 32B and Figure 32F, sg3), suggesting that the fully targeted single guide RNA includes the last spacer and last repeat of the array as well as the HEARO RNA (Figure 32F, sg3). Furthermore, this candidate covariance model predicted a second HEARO RNA upstream and independent of the CRISPR array (Figure 32B and Figure 32E, sg2). Another example of a CRISPR-associated SMART HEARO system is MG35-556 (SEQ ID NO: 659) (Figure 32C), where the HEARO RNA is encoded in the 5'UTR region of the effector, which contains an anti-repeat complementary to one of the CRISPR repeats (Figure 32C). This represents an example of a dual-guide RNA-guided HEARO system, where one CRISPR repeat (likely carrying a targeting spacer at its 5' end) anneals to the 5' end of the HEARO RNA and folds into a structure similar to the other single-guide HEARO RNA (Figure 32G).
[0239] When tested for cleavage activity, many SMART HEARO nucleases were active in the in vitro TAM determination assay, and some of them were active with multiple sgRNA designs (Figure 33A-C). MG35-104 (SEQ ID NO: 128), HEARO MG35-463 (SEQ ID NO: 530), and MG35-518 (SEQ ID NO: 621) were among the most active nucleases, as indicated by strong band intensity readouts (Figure 33). Furthermore, SMART HEARO MG35-463 (SEQ ID NO: 530) is functional with both its CRISPR-associated (SEQ ID NO: 1237) and CRISPR-independent (SEQ ID NO: 1236) HEARO RNAs, even though the guide RNAs share only 65% pairwise nucleotide identity (Figure 32D, Figure 32E, and Figure 33C). The active MG35 candidates recognize a variety of TAMs and display cleavage preference for positions 5 or 7 from the TAM motif (Figure 34).
[0240] Example 17 - SMART HEARO enzymes are efficient nucleases In vitro cleavage assay
[0241] MG35 nuclease was expressed for 2 hours at 37°C using in vitro transcription / translation (IVTT) (New England Biolabs). Transcription was driven by a T7 promoter on a linear DNA template encoding the nuclease. Guide RNA was transcribed separately in vitro and added to the IVTT mixture at a selected concentration, typically between 0.4 at 4 μM. In vitro cleavage reactions were performed in 1× effector buffer (10 mM Tris-HCl pH 7.5, 100 mM NaCl, 10 mM MgCl 2 ) or 1x New England Biolabs 2.1 buffer (10 mM Tris-HCl pH 7.9, 50 mM NaCl, 10 mM MgCl 2Reactions were performed by adding 3 μL of RNP sample to 5 nM supercoiled DNA in 10 μL reaction volume in 10 μL of RNAse A (New England Biolabs, 100 μg / ml BSA). The reaction was incubated at 37° C. for 1 h and then quenched by adding 0.2 μg of RNAse A (New England Biolabs), followed by incubation at 37° C. for 20 min. Then, 4 units of proteinase K (New England Biolabs) were added, followed by incubation at 55° C. for 30 min. Reactions were analyzed by capillary electrophoresis using a D5000 Tapestation kit (Agilent) following the manufacturer's recommended instructions for analysis and visualization. Successful cleavage results in supercoiled 2200 bp DNA being cleaved into linear dsDNA.
[0242] After identifying the active guide RNA and TAM recognition motif, SMART HEARO nucleases were tested for in vitro cleavage efficiency via in vitro transcription / translation co-expression of the nuclease with its guide RNA, followed by incubation with a target plasmid containing a spacer targeted by the guide RNA and a TAM identified in the TAM / PAM enrichment screen. Cleavage is measured by the transfer of uncleaved products (supercoiled) to cleaved linear DNA (Figure 35A and Figure 35B). The results show that MG35-104 (SEQ ID NO: 128) is highly efficient in dsDNA cleavage compared to other active SMART HEARO nucleases (Figure 35A and Figure 35B).
[0243] Example 18 - SMART HEARO Guide Operation Some active SMART II nuclease guide RNAs contain one or more poly-T regions (four or more T bases in a row) (Figure 36A), which may limit transcription efficiency. Three poly-T mutant sgRNAs per candidate were designed and tested for in vitro cleavage activity, and their activity was compared to the activity of the candidates with their native guide RNAs (Figures 36A and 36B). Results show that MG35-94 is active with mutant guides M2 and M3, while MG35-104 is active with all three guide mutations M1-M3, with guide M3 retaining the highest activity compared to the other guides. MG35-518 is active with all three mutants tested, with M1 showing the highest activity (Figure 36B).
[0244] [Table 6-1]
[0245] [Table 6-2]
[0246] Example 19 - Computational reconstruction of novel SMART I nucleases In silico reconstruction of novel sequences
[0247] In an attempt to generate further diversity in SMART I nucleases, divergent nuclease sequences were reconstructed using the ancestral sequence reconstruction algorithm. Ancestral sequence reconstruction is a computational technique that uses existing protein sequences and the inferred relationships between them to reconstruct the sequences of ancient, now extinct proteins (Harms, M. & Thornton JW Analyzing protein structure and function using ancestral gene reconstruction. Current Opinion in Structural Biology 2010, 20, 360-366). Using this technique, novel sequences of the SMART I MG34 family were computationally reconstructed. In this analysis, 190 SMART I protein sequences were aligned using MAFFT with parameters L-INS-i or G-INS-i (Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013, 30(4), 772-780) and phylogenetic trees were constructed using either Fasttree (Price, MN, Dehal, PS, and Arkin, APFastTree 2--Approximately Maximum-Likelihood Trees for Large Alignments. PLoS ONE 2010, 5(3), e9490) or RAxML (Stamatakis, A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 2014, 30(9), 1312-1313) (Figure 37). The phylogenetic tree was rooted using SpCas9 and SaCas9.Sequence reconstructions were performed using the codeml package in PAML4.8 (Yang, Z. PAML 4: a program package for phylogenetic analysis by maximum likelihood. Molecular Biology and Evolution 2007, 24, 1586-1591), applying all four alignment combinations and tree-building methods to account for uncertainties in phylogeny. Insertions and deletions were manually identified for each reconstructed node.
[0248] In vitro PAM determination assay Candidate proteins were codon-optimized for E. coli and cloned into a vector with a T7 promoter and a C-terminal His tag. Genes were PCR amplified with primer binding sites 150 bp upstream and downstream of the T7 promoter and terminator sequences, respectively. The PCR products were added to PURExpress (New England Biolabs) at a final concentration of 5 nM and expressed for 2 hours at 37°C. Five-fold dilutions of PURExpress, 5 nM of the 8N PAM plasmid library, and 50 nM of sgRNA targeting the PAM library were used to express the proteins in 10 mM Tris pH 7.5, 100 mM NaCl, and 10 mM MgCl. 2Cleavage reactions were assembled in a centrifuge. Cleavage products from PURExpress reactions were recovered via cleanup with SPRI beads (AMPure Beckman Coulter or HighPrep Sigma-Aldritch). DNA was blunted via the addition of Klenow fragment and dNTPs (New England Biolabs). Blunt-end products were ligated with 100-fold excess of double-stranded adapter sequences and used as templates for preparation of NGS libraries, from which PAM requirements were determined from sequence analysis. Raw NGS reads were filtered by Phred quality score >20. Using 14–24 bp representing documented DNA sequence from the backbone adjacent to the PAM as a reference, PAM-proximal regions were found and the adjacent 8 bp were identified as putative PAMs. The distance between the PAM and the ligated adapter was also measured for each read. Reads that did not exactly match the reference or adapter sequences were excluded. PAM sequences were filtered by cleavage site frequency such that only PAMs with the most frequent cleavage site ±2 bp were included in the analysis. The filtered list of PAMs was used to generate sequence logos using Logomaker (Tareen, A. & Kinney, J.B. Logomaker: beautiful sequence logos in Python. Bioinformatics 2020, 36, 2272-2274).
[0249] Six sequences of the MG34 family were reconstructed with high confidence (Table 5 and Figure 37), and the catalytic and binding domains were confirmed from multiple sequence alignments and 3D structure prediction (Figure 38).
[0250] [Table 7]
[0251] The main difference between the structures is in the recognition lobe, suggesting that these reconstituted effectors may exhibit nuclease activity similar to MG34-1. Given the strong support for the newly reconstituted candidates, six novel nucleases were tested for in vitro cleavage activity in PAM enrichment assays using guide RNAs from three active MG34 nucleases: MG34-1 sgRNA 1 (SEQ ID NO: 613), MG34-9 sgRNA 1 (SEQ ID NO: 615), and MG34-16 sgRNA 1 (SEQ ID NO: 616). The novel nucleases MG34-27 (SEQ ID NO: 1314) and MG34-29 (SEQ ID NO: 1316) were active with all three tested sgRNAs, as indicated by the expected cleavage band at approximately 180 bp (Figure 39). The PAM targeted by these novel nucleases is likely to be 3'nRR, with nGG being the most commonly recognized PAM (Figure 40). Results indicate that the newly reconstituted nucleases have more relaxed PAM recognition compared to other active MG34 nucleases (e.g., MG34-1 recognizes a 3'nGG PAM) and flexible cleavage preferences at positions 6-9 from the PAM (Figure 40).
[0252] [Table 8-1]
[0253] [Table 8-2]
[0254] [Table 8-3]
[0255] [Table 8-4]
[0256] [Table 8-5]
[0257]
Table 8-6
[0258]
Table 8-7
[0259]
Table 8-8
[0260]
Table 8-9
[0261]
Table 8-10
[0262]
Table 8-11
[0263]
Table 8-12
[0264]
Table 8-13
[0265]
Table 8-14
[0266]
Table 8-15
[0267]
Table 8-16
[0268]
Table 8-17
[0269]
Table 8-18
[0270]
Table 8-19
[0271]
Table 8-20
[0272]
Table 8-21
[0273]
Table 8-22
[0274]
Table 8-23
[0275]
Table 8-24
[0276]
Table 8-25
[0277]
Table 8-26
[0278]
Table 8-27
[0279]
Table 8-28
[0280]
Table 8-29
[0281]
Table 8-30
[0282]
Table 8-31
[0283]
Table 8-32
[0284]
Table 8-33
[0285]
Table 8-34
[0286]
Table 8-35
[0287]
Table 8-36
[0288]
Table 8-37
[0289]
Table 8-38
[0290]
Table 8-39
[0291]
Table 8-40
[0292]
Table 8-41
[0293]
Table 8-42
[0294]
Table 8-43
[0295]
Table 8-44
[0296]
Table 8-45
[0297]
Table 8-46
[0298]
Table 8-47
[0299]
Table 8-48
[0300]
Table 8-49
[0301]
Table 8-50
[0302]
Table 8-51
[0303]
Table 8-52
[0304]
Table 8-53
[0305]
Table 8-54
[0306]
Table 8-55
[0307]
Table 8-56
[0308]
Table 8-57
[0309]
Table 8-58
[0310]
Table 8-59
[0311]
Table 8-60
[0312]
Table 8-61
[0313]
Table 8-62
[0314]
Table 8-63
[0315]
Table 8-64
[0316]
Table 8-65
[0317]
Table 8-66
[0318]
Table 8-67
[0319]
Table 8-68
[0320]
Table 8-69
[0321]
Table 8-70
[0322]
Table 8-71
[0323]
Table 8-72
[0324]
Table 8-73
[0325]
Table 8-74
[0326]
Table 8-75
[0327]
Table 8-76
[0328]
Table 8-77
[0329]
Table 8-78
[0330]
Table 8-79
[0331]
Table 8-80
[0332]
Table 8-81
[0333]
Table 8-82
[0334]
Table 8-83
[0335]
Table 8-84
[0336]
Table 8-85
[0337]
Table 8-86
[0338] [Table 8-87]
[0339] [Table 8-88]
[0340] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided for illustrative purposes only. The present invention is not intended to be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing description, the description and explanation of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in the practice of the present invention. It is therefore contemplated that the present invention will encompass any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. An engineered nuclease system comprising: (a) an endonuclease or a nucleic acid encoding said endonuclease, wherein said endonuclease comprises a sequence having at least 80% sequence identity to SEQ ID NO: 1316; and (b) an engineered guide ribonucleic acid structure or a nucleic acid encoding said engineered guide ribonucleic acid structure, wherein said engineered guide ribonucleic acid structure is configured to form a complex with said endonuclease, and wherein said engineered guide ribonucleic acid structure: (i) a guide ribonucleic acid sequence configured to hybridize to a target nucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease; an engineered guide ribonucleic acid structure or a nucleic acid encoding said engineered guide ribonucleic acid structure, An engineered nuclease system comprising:
2. 2. The engineered nuclease system of claim 1, wherein the endonuclease is an archaeal endonuclease.
3. The engineered nuclease system of claim 1, wherein the endonuclease further comprises one or more of an arginine-rich region containing an RRxRR motif (sequence number 1361), a domain with PF14239 homology, a recognition (REC) domain, a bridge helix (BH) domain, a wedge (WED) domain, or a PAM-interacting (PI) domain.
4. The engineered nuclease system of claim 3, wherein the arginine-rich region containing the RRxRR motif (SEQ ID NO: 1361), the domain with PF14239 homology, the recognition (REC) domain, the bridge helix (BH) domain, the wedge (WED) domain, or the PAM-interacting (PI) domain each comprise a sequence having at least 85% sequence identity to the arginine-rich region containing the RRxRR motif (SEQ ID NO: 1361), the domain with PF14239 homology, the recognition (REC) domain, the bridge helix (BH) domain, the wedge (WED) domain, or the PAM-interacting (PI) domain of SEQ ID NO: 1316.
5. The engineered nuclease system of claim 1, wherein the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease.
6. The engineered nuclease system of claim 1, wherein the endonuclease comprises a sequence having at least 90% sequence identity to SEQ ID NO: 1316.
7. The engineered nuclease system described in claim 6, wherein the endonuclease comprises the sequence of SEQ ID NO: 1316.
8. The engineered nuclease system described in claim 1, wherein the tracr ribonucleic acid sequence comprises a polynucleotide sequence having at least 80% sequence identity to SEQ ID NO:
200.
9. The engineered nuclease system described in claim 8, wherein the tracr ribonucleic acid sequence comprises the polynucleotide sequence of SEQ ID NO:
200.
10. The engineered nuclease system described in claim 6, wherein the tracr ribonucleic acid sequence comprises a polynucleotide sequence having at least 90% sequence identity to SEQ ID NO:
200.
11. The engineered nuclease system of claim 1, wherein the engineered guide ribonucleic acid structure comprises a sequence having at least 80% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 613, 615, or 616.
12. The engineered nuclease system described in claim 11, wherein the engineered guide ribonucleic acid structure comprises any one of the non-degenerate nucleotides of SEQ ID NO: 613, 615, or 616.
13. The engineered nuclease system of claim 6, wherein the engineered guide ribonucleic acid structure comprises a sequence having at least 90% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 613, 615, or 616.
14. The engineered nuclease system of claim 1, wherein the engineered guide ribonucleic acid structure comprises (a) at least two ribonucleic acid polynucleotides, or (b) a single ribonucleic acid polynucleotide comprising the guide ribonucleic acid sequence and the tracr ribonucleic acid sequence.
15. The engineered nuclease system of claim 1, wherein the guide ribonucleic acid sequence is complementary to a eukaryotic, fungal, plant, mammalian, or human genome sequence.
16. The engineered nuclease system of claim 1, further comprising a single-stranded or double-stranded deoxyribonucleic acid repair template.
17. The engineered nuclease system of claim 16, wherein the single-stranded or double-stranded deoxyribonucleic acid repair template comprises, in a 5' to 3' direction, a first homologous arm comprising a sequence of at least 20 nucleotides 5' to the target nucleic acid sequence, a synthetic deoxyribonucleic acid sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target nucleic acid sequence.
18. The engineered nuclease system of claim 16, wherein the single-stranded or double-stranded deoxyribonucleic acid repair template comprises an introduced gene donor.
19. The engineered nuclease system of claim 1, wherein the sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix with parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, and using a conditional composition score matrix adjustment.
20. A nucleic acid encoding an engineered nuclease system according to any one of claims 1 to 19.