Enzymes containing RUVC domains

An engineered nuclease system with a RuvC_III domain and guide RNA structure addresses the limitations of Cas9 by enhancing DNA targeting specificity and versatility, facilitating efficient gene editing across various organisms.

JP7778087B2Active Publication Date: 2025-12-01METAGENOMI THERAPEUTICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022567462
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-27
Filing Date
2021-05-06
Publication Date
2025-12-01
Estimated Expiration
2041-05-06

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems, particularly Cas9 endonucleases, have limitations in binding to diverse protospacer adjacent motifs (PAMs) and require extensive engineering to target specific DNA sequences effectively, limiting their versatility in gene editing applications.

Method used

Development of an engineered nuclease system comprising a Class 2 Type II Cas endonuclease with a RuvC_III domain derived from uncultured microorganisms, combined with a guide RNA structure, to enhance targeting specificity and versatility by binding to various PAM sequences and forming a complex with a tracr RNA.

Benefits of technology

The engineered nuclease system improves DNA targeting accuracy and versatility, enabling efficient gene editing across diverse genomic sequences, including prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, and mammalian genomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778087000050
    Figure 0007778087000050
  • Figure 0007778087000051
    Figure 0007778087000051
  • Figure 0007778087000052
    Figure 0007778087000052
Patent Text Reader

Abstract

The present disclosure provides endonuclease enzymes with distinct domain characteristics, as well as methods of using such enzymes or variants thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims priority to U.S. Provisional Application No. 63 / 022,320, entitled "ENZYMES WITH RUVC DOMAINS," filed May 8, 2020, U.S. Provisional Application No. 63 / 032,464, entitled "ENZYMES WITH RUVC DOMAINS," filed May 29, 2020, U.S. Provisional Application No. 63 / 116,155, entitled "ENZYMES WITH RUVC DOMAINS," filed November 19, 2020, and U.S. Provisional Application No. 63 / 180,570, entitled "ENZYMES WITH RUVC DOMAINS," filed April 27, 2021, all of which are incorporated by reference in their entireties. [Background technology]

[0002] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeat (CRISPR)-guided ribonucleic acid (RNA), appear to be widespread (approximately 45% of bacteria and 84% of archaea) components of prokaryotic immune systems that help protect such microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, through CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. While CRISPR DNA elements were observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR / Cas complexes were only recognized relatively recently, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA manipulation and gene editing applications.

[0003] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format, and is incorporated herein by reference in its entirety. The ASCII copy created on May 29, 2020 is named 55921_712_601_SL.txt and is 24,659,439 bytes in size. Summary of the Invention

[0004] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC_III domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, and the endonuclease is a Class 2 Type II Cas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the RuvC_III domain comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 1827-3637.

[0005] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC_III domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 1827-3637; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease.

[0006] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs:5512-5537, wherein the endonuclease is a Class 2 Type II Cas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease.

[0007] In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has not been engineered to bind to a different PAM sequence. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to about 60 to 90 contiguous nucleotides selected from any one of SEQ ID NOs: 5476-5511 and 5538.

[0008] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to an endonuclease, wherein the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity over about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 5476-5511 and 5538; and (b) a Class 2 Type II Cas endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 5512-5537.

[0009] In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides, hi some embodiments, the engineered guide ribonucleic acid structure comprises one ribonucleic acid polynucleotide comprising a guide ribonucleic acid sequence and a tracr ribonucleic acid sequence.

[0010] In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612.

[0011] In some embodiments, the engineered nuclease system further comprises a single-stranded or double-stranded DNA repair template comprising, in 5' to 3' order, a first homology arm comprising a sequence of at least 20 nucleotides that is 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homology arm comprising a sequence of at least 20 nucleotides that is 3' to the target sequence. In some embodiments, the first homology arm or the second homology arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides.

[0012] In some embodiments, the system further comprises a source of Mg2+.

[0013] In some embodiments, the endonuclease and the tracr ribonucleic acid sequence are derived from different bacterial species within the same phylum. In some embodiments, the endonuclease is derived from a bacterium belonging to the genus Dermabacter. In some embodiments, the endonuclease is derived from a bacterium belonging to the phylum Verrucomicrobia, Candidatus Peregrinibacteria, or Candidatus Melainabacteria. In some embodiments, the endonuclease is derived from a bacterium containing a 16S rRNA gene having at least 90% identity to any one of SEQ ID NOs: 5592-5595.

[0014] In some embodiments, the HNH domain comprises a sequence having at least 70% or at least 80% identity to any one of SEQ ID NOs: 5638-5460. In some embodiments, the endonuclease comprises SEQ ID NOs: 1-1826, or a variant thereof having at least 55% identity thereto. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 1827-1830 or 1827-2140.

[0015] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 3638-3641 or 3638-3954. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5615-5632. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 1-4 or 1-319.

[0016] In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 5461-5464, 5476-5479, or 5476-5489. In some embodiments, the guide RNA structure comprises an RNA sequence that is predicted to comprise a hairpin consisting of a stem and a loop, wherein the stem comprises at least 10, at least 12, or at least 14 base pairs of ribonucleotides and an asymmetric bulge within 4 base pairs of the loop.

[0017] In some embodiments, the endonuclease is configured to bind to a PAM comprising a sequence selected from the group consisting of SEQ ID NOs: 5512-5515 or 5527-5530.

[0018] In some embodiments, (a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1827; (b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5461 or SEQ ID NO: 5476; and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5512 or SEQ ID NO: 5527. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1828; (b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5462 or SEQ ID NO: 5477; and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5513 or SEQ ID NO: 5528. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1829; (b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5463 or SEQ ID NO: 5478; and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5514 or SEQ ID NO: 5529. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1830; (b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5464 or SEQ ID NO: 5479; and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5515 or SEQ ID NO: 5530.

[0019] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2141-2142 or 2141-2241. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 3955-3956 or 3955-4055. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5632-5638. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 320-321 or 320-420. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:5465, SEQ ID NOs:5490-5491, or SEQ ID NOs:5490-5494. In some embodiments, the guide RNA structure comprises a tracr ribonucleic acid sequence comprising a hairpin comprising at least 8, at least 10, or at least 12 base pairs of ribonucleotides. In some embodiments, the endonuclease is configured to bind to a PAM comprising a sequence selected from the group consisting of SEQ ID NO:5516 or SEQ ID NO:5531. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO:2141, (b) the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO:5490, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO:5531.In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2142; (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5465 or SEQ ID NO: 5491; and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5516.

[0020] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2245-2246. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4059-4060. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5639-5648. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 424-425. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 5498-5499 and SEQ ID NO: 5539. In some embodiments, the guide RNA structure comprises a guide ribonucleic acid sequence that is predicted to comprise a hairpin having an uninterrupted base-paired region comprising at least 8 nucleotides of the guide ribonucleic acid sequence and at least 8 nucleotides of the tracr ribonucleic acid sequence, where the tracr ribonucleic acid sequence comprises, from 5' to 3', a first hairpin and a second hairpin, and the first hairpin has a longer stem than the second hairpin.

[0021] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2242-2244 or 2247-2249. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4056-4058 or 4061-4063. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5639-5648. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 421-423 or 426-428. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 5466-5467, 5495-5497, 5500-5502, and 5539. In some embodiments, the guide RNA structure comprises a guide ribonucleic acid sequence that is predicted to comprise a hairpin having an uninterrupted base-paired region comprising at least 8 nucleotides of the guide ribonucleic acid sequence and at least 8 nucleotides of the tracr ribonucleic acid sequence, wherein the tracr ribonucleic acid sequence comprises, from 5' to 3', a first hairpin and a second hairpin, the first hairpin having a longer stem than the second hairpin. In some embodiments, the endonuclease is configured to bind to a PAM comprising a sequence selected from the group consisting of SEQ ID NOs: 5517-5518 or 5532-5534. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2247; (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5500; and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5517 or SEQ ID NO: 5532.In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2248, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5501, and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5518 or SEQ ID NO: 5533. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2249, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5502, and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5534.

[0022] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 2253 or SEQ ID NOs: 2253-2481. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4067 or SEQ ID NOs: 4067-4295. In some embodiments, the endonuclease comprises a peptide motif according to SEQ ID NO: 5649. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 432 or SEQ ID NOs: 432-660. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5468 or SEQ ID NO: 5503. In some embodiments, the endonuclease is configured to bind to a PAM comprising a sequence selected from the group consisting of SEQ ID NO: 5519. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2253; (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5468 or SEQ ID NO: 5503; and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5519.

[0023] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2482-2489. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4296-4303. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 661-668. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2490-2498. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4304-4312. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 669-677. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5504.

[0024] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:2499 or SEQ ID NOs:2499-2750. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:4313 or SEQ ID NOs:4313-4564. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs:5650-5667. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:678 or SEQ ID NOs:678-929. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO:5469 or SEQ ID NO:5505. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5520 or SEQ ID NO: 5535. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2499, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5469 or SEQ ID NO: 5505, and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5520 or SEQ ID NO: 5535.

[0025] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:2751 or SEQ ID NOs:2751-2913. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:4565 or SEQ ID NOs:4565-4727. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs:5668-5678. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:930 or SEQ ID NOs:930-1092. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO:5470 or SEQ ID NO:5506. In some embodiments, the endonuclease is configured to bind to a PAM comprising a sequence selected from the group consisting of SEQ ID NO: 5521 or SEQ ID NO: 5536. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 2751, (b) the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 5470 or SEQ ID NO: 5506, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5521 or SEQ ID NO: 5536.

[0026] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:2914 or SEQ ID NOs:2914-3174. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:4728 or SEQ ID NOs:4728-4988. In some embodiments, the endonuclease comprises at least one, at least two, or at least three peptide motifs selected from the group consisting of SEQ ID NOs:5676-5678. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:1093 or SEQ ID NOs:1093-1353. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:5471, SEQ ID NO:5507, and SEQ ID NOs:5540-5542. In some embodiments, the guide RNA structure comprises a tracr ribonucleic acid sequence that is predicted to include at least two hairpins comprising fewer than five base pairs of ribonucleotides. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO:5522. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO:2914, (b) the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO:5471 or SEQ ID NO:5507, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO:5522.

[0027] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3175 or SEQ ID NOs: 3175-3330. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4989 or SEQ ID NOs: 4989-5146. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5679-5686. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1354 or SEQ ID NOs: 1354-1511. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5472 or SEQ ID NO: 5508. In some embodiments, the endonuclease is configured to bind to a PAM comprising a sequence selected from the group consisting of SEQ ID NO: 5523 or SEQ ID NO: 5537. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 3175, (b) the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 5472 or SEQ ID NO: 5508, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5523 or SEQ ID NO: 5537.

[0028] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:3331 or SEQ ID NOs:3331-3474. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:5147 or SEQ ID NOs:5147-5290. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs:5674-5675 and SEQ ID NOs:5687-5693. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO:1512 or SEQ ID NOs:1512-1655. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5473 or SEQ ID NO: 5509. In some embodiments, the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5524. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 3331, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5473 or SEQ ID NO: 5509, and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5524.

[0029] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3475 or SEQ ID NOs: 3475-3568. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5291 or SEQ ID NOs: 5291-5389. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5694-5699. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1656 or SEQ ID NOs: 1656-1755. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 5474 or SEQ ID NO: 5510. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5525. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 3475, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5474 or SEQ ID NO: 5510, and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5525.

[0030] In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3569 or SEQ ID NOs: 3569-3637. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5390 or SEQ ID NOs: 5390-5460. In some embodiments, the endonuclease comprises at least one, at least two, at least three, at least four, or at least five peptide motifs selected from the group consisting of SEQ ID NOs: 5700-5717. In some embodiments, the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1756 or SEQ ID NOs: 1756-1826. In some embodiments, the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 5475 or SEQ ID NO: 5511. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5526. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 3569, (b) the guide RNA structure comprises a sequence at least 70%, 80%, or 90% identical to SEQ ID NO: 5475 or SEQ ID NO: 5511, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5526. In some embodiments, sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using Smith-Waterman homology search algorithm parameters. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and a BLOSUM62 scoring matrix setting gap cost at 11, presence, and extension of 1, and using a conditional composition score matrix adjustment.

[0031] In some aspects, the present disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA-targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, wherein the engineered guide ribonucleic acid polynucleotide is configured to form a complex with or target a complex to a target sequence in the target DNA molecule, the endonuclease comprising a RuvC_III domain having a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 1827-3637. In some embodiments, the DNA-targeting segment is located 5' of both of the two complementary stretches of nucleotides.

[0032] In some embodiments, (a) the protein-binding segment comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% identity to a sequence selected from the group consisting of SEQ ID NOs: 5476-5479 or SEQ ID NOs: 5476-5489; and (b) the protein-binding segment is selected from the group consisting of (SEQ ID NOs: 5490-5491 or SEQ ID NOs: 5490-5494) and SEQ ID NO: 5538. (c) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5498-5499; and (d) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5495-5497 and SEQ ID NOs: 5500-5502. (e) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5503; (f) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5504; (g) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5505; and (h) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5506. 506, (i) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5507, (j) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5508, (k) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5509,or at least 90% identical to SEQ ID NO: 5510, (l) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5510, or (m) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5511.

[0033] In some embodiments, (a) the guide ribonucleic acid polynucleotide comprises an RNA sequence comprising a hairpin comprising a stem and a loop, wherein the stem comprises at least 10, at least 12, or at least 14 base pairs of ribonucleotides and an asymmetric bulge within 4 base pairs of the loop; (b) the guide ribonucleic acid polynucleotide comprises a tracr ribonucleic acid sequence that is predicted to comprise a hairpin comprising at least 8, at least 10, or at least 12 base pairs of ribonucleotides; (c) the guide ribonucleic acid polynucleotide comprises a guide ribonucleic acid sequence that is predicted to comprise a hairpin having an uninterrupted basepair region comprising at least 8 nucleotides of the guide ribonucleic acid sequence and at least 8 nucleotides of the tracr ribonucleic acid sequence, wherein, from 5' to 3', the tracr ribonucleic acid sequence comprises a first hairpin and a second hairpin, the first hairpin having a longer stem than the second hairpin; or (d) the guide ribonucleic acid polynucleotide comprises a tracr ribonucleic acid sequence that is predicted to comprise at least two hairpins comprising fewer than 5 base pairs of ribonucleotides.

[0034] In some aspects, the present disclosure provides a deoxyribonucleic acid polynucleotide that encodes any of the engineered guide ribonucleic acid polynucleotides described herein.

[0035] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a Class 2 Type II Cas endonuclease comprising a RuvC_III domain and an HNH domain, and the endonuclease is derived from an uncultured microorganism.

[0036] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding an endonuclease comprising a RuvC_III domain having at least 70% sequence identity to any one of SEQ ID NOs: 1827-3637. In some embodiments, the endonuclease comprises an HNH domain having at least 70% or at least 80% sequence identity to any one of SEQ ID NOs: 3638-5460. In some embodiments, the endonuclease comprises SEQ ID NOs: 5572-5591, or a variant thereof having at least 70% sequence identity thereto. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612.

[0037] In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human. In some embodiments, the organism is Escherichia coli (E. coli), and (a) the nucleic acid sequence has at least 70%, 80%, or 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5572-5575, (b) the nucleic acid sequence has at least 70%, 80%, or 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5576-5577, (c) the nucleic acid sequence has at least 70%, 80%, or 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5578-5580, (d) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5581, and (e) the nucleic acid sequence is , SEQ ID NO: 5582, (f) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5583, (g) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5584, (h) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5585, (i) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5586, or (j) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5587. In some embodiments, the organism is a human and (a) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5588 or SEQ ID NO: 5589, or (b) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5590 or SEQ ID NO: 5591.

[0038] In some aspects, the disclosure provides a vector comprising a nucleic acid sequence encoding a Class 2 Type II Cas endonuclease comprising a RuvC_III domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism.

[0039] In some aspects, the disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with an endonuclease, the nucleic acid comprising (a) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (b) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.

[0040] In some aspects, the present disclosure provides a cell comprising any of the vectors described herein.

[0041] In some aspects, the disclosure provides a method of producing an endonuclease, comprising culturing any of the cells described herein.

[0042] In some aspects, the disclosure provides methods for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide in a complex with a Class 2 Type II Cas endonuclease and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide; (b) the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM); and (c) the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 5512-5526 or 5527-5537. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to the sequence of the engineered guide ribonucleic acid structure and a second strand comprising a PAM. In some embodiments, the PAM is immediately adjacent to the 3' end of the sequence complementary to the sequence of the engineered guide ribonucleic acid structure.

[0043] In some embodiments, the Class 2 Type II Cas endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the Class 2 Type II Cas endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.

[0044] In some embodiments, (a) the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 5512-5515 and 5527-5530; (b) the PAM comprises SEQ ID NO: 5516 or SEQ ID NO: 5531; (c) the PAM comprises SEQ ID NO: 5539; (d) the PAM comprises SEQ ID NO: 5517 or SEQ ID NO: 5518; (e) the PAM comprises SEQ ID NO: 5519; (f) the PAM comprises SEQ ID NO: 5520 or SEQ ID NO: 5535; (g) the PAM comprises SEQ ID NO: 5521 or SEQ ID NO: 5536; (h) the PAM comprises SEQ ID NO: 5522; (i) the PAM comprises SEQ ID NO: 5523 or SEQ ID NO: 5537; (j) the PAM comprises SEQ ID NO: 5524; (k) the PAM comprises SEQ ID NO: 5525; or (l) the PAM comprises SEQ ID NO: 5526.

[0045] In some aspects, the present disclosure provides methods of modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus any of the engineered nuclease systems described herein, wherein the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure, and the complex is configured to modify the target nucleic acid locus upon binding to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell.

[0046] In some embodiments, delivering an engineered nuclease system to a target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering an engineered nuclease system to a target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding the endonuclease is operably linked. In some embodiments, the engineered nuclease system to a target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the endonuclease. In some embodiments, the engineered nuclease system to a target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, the engineered nuclease system to a target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or adjacent to the target locus.

[0047] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846, or SEQ ID NO: 6257; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs: 5847-5861 or 6258-6278, wherein the endonuclease is a Class 2 Type II Cas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has not been engineered to bind to a different PAM sequence. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease, hi some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease.In some embodiments, the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of (a) any one of SEQ ID NOs: 5886-5887, 5891, 5893, or 5894, or (b) any one of SEQ ID NOs: 5862-5885, 5888-5890, 5892, 5895-5896, or 6279-6301. In some aspects, the disclosure provides an engineered nuclease system, comprising: (a) an engineered guide ribonucleic acid structure comprising: (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to an endonuclease, wherein the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of (a) any one of SEQ ID NOs: 5886-5887, 5891, 5893, or 5894; or (b) any one of SEQ ID NOs: 5862-5885, 5888-5890, 5892, 5895-5896, or 6279-6301; and a Class 2 Type II Cas endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 5847-5861 or 6258-6278. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length or 19-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in 5' to 3' order, a first homology arm comprising a sequence of at least 20 nucleotides 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homology arm comprising a sequence of at least 20 nucleotides 3' to the target sequence.In some embodiments, the first homology arm or the second homology arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using Smith-Waterman homology search algorithm parameters. In some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and a BLOSUM62 scoring matrix setting gap cost at 11, presence extension of 1, and conditional composition score matrix adjustment.

[0048] In some aspects, the present disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA-targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, wherein the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or SEQ ID NO: 6257, and target the complex to the target sequence in the target DNA molecule. In some embodiments, the DNA-targeting segment is located 5' to both of the two complementary stretches of nucleotides.

[0049] In some aspects, the present disclosure provides a deoxyribonucleic acid polynucleotide that encodes any of the engineered guide ribonucleic acid polynucleotides described herein.

[0050] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or 6257. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.

[0051] In some aspects, the present disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the guide ribonucleic acid structure comprising (a) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (b) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.

[0052] In some aspects, the present disclosure provides a cell comprising any of the vectors described herein.

[0053] In some aspects, the disclosure provides a method of producing an endonuclease, comprising culturing any of the cells described herein.

[0054] In some aspects, the disclosure provides methods for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising contacting the double-stranded deoxyribonucleic acid polynucleotide in a complex with a Class 2 Type II Cas endonuclease and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), wherein the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 5847-5861 or 6258-6278. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the engineered guide ribonucleic acid structure and a second strand comprising the PAM. In some embodiments, the PAM is immediately adjacent to the 3' end of the sequence complementary to the sequence of the engineered guide ribonucleic acid structure. In some embodiments, the Class 2 Type II Cas endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.

[0055] In some aspects, the present disclosure provides methods of modifying a target nucleic acid locus, the method comprising delivering any of the engineered nuclease systems described herein to the target nucleic acid locus, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, the complex configured to modify the target nucleic acid locus upon binding to the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide.In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or near the target locus.

[0056] In some aspects, the disclosure provides a method of editing a TRAC locus in a cell, comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TRAC locus, and wherein the engineered guide RNA comprises at least 19, at least 20, at least 21, or at least 30 of any one of SEQ ID NOs: 5950-5958 or SEQ ID NOs: 5959-5965. In some embodiments, the RNA-guided endonuclease is a Class II Type II Cas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 5950-5958, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 5959-5965, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 5953-5957. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 5960-5961 or 5963-5964.

[0057] In some aspects, the disclosure provides a method of editing a TRBC locus in a cell, comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TRBC locus, and the engineered guide RNA comprises at least 19, at least 20, at least 21, or at least 30 of any one of SEQ ID NOs: 5966-6004 or SEQ ID NOs: 6005-6025. In some embodiments, the RNA-guided endonuclease is a Class II Type II Cas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 5966-6004, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6005-6025, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 5970, 5971, 5983, or 5984. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6006, 6010, 6011, or 6012.

[0058] In some aspects, the disclosure provides a method of editing a GR (NR3C1) locus in a cell, comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the GR (NR3C1) locus, and the engineered guide RNA comprises at least 19, at least 20, at least 21, or at least 30 of any one of SEQ ID NOs: 6026-6090 or SEQ ID NOs: 6091-6121. In some embodiments, the RNA-guided endonuclease is a Class II Type II Cas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6026-6090, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6091-6121, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence that is at least 85% identical to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6027-6028, 6029, 6038, 6043, 6049, 6076, 6080, 6081, or 6086. In some embodiments, the engineered guide RNA comprises a target sequence that is at least 85% identical to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6092, 6115, or 6119.

[0059] In some aspects, the disclosure provides a method of editing the AAVS1 locus in a cell, comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the AAVS1 locus, and wherein the engineered guide RNA is at least 19, at least 20, at least 21, at least 22, at least 30, at least 4 ... In some embodiments, the RNA-guided endonuclease is a Class II Type II Cas endonuclease, comprising a target sequence having at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity over at least 23 or at least 24 contiguous nucleotides. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO:421 or SEQ ID NO:423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6122, 6125-6126, 6128, 6131, 6133, 6136, 6141, 6143, or 6148.

[0060] In some aspects, the disclosure provides a method of editing a TIGIT locus in a cell, comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TIGIT locus, and wherein the engineered guide RNA is at least 19, at least 20, at least 21, at least 22, at least 30, at least 4 ... In some embodiments, the RNA-guided endonuclease is a Class II Type II Cas endonuclease, comprising a target sequence having at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity over at least 23 or at least 24 contiguous nucleotides. In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO:421 or SEQ ID NO:423.In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 66155, 6159, 616, or 6172.

[0061] In some aspects, the disclosure provides a method of editing a CD38 locus in a cell, comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the CD38 locus, and wherein the engineered guide RNA comprises at least 19, at least 20, at least 21, or at least 30 of any one of SEQ ID NOs: 6182-6248 or SEQ ID NOs: 6249-6256. In some embodiments, the RNA-guided endonuclease is a Class II Type II Cas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6182-6248, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6249-6256, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 6182-6183, 6189, 6191, 6208, 6210, 6211, or 6215. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 contiguous nucleotides of SEQ ID NO: 6251.

[0062] In some embodiments of any of the above methods of editing a specific locus in a cell, the cell is a peripheral blood mononuclear cell, a T cell, a NK cell, a hematopoietic stem cell (HSCT), or a B cell, or any combination thereof.

[0063] In some aspects, the disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide; wherein the engineered guide ribonucleic acid polynucleotide is configured to form a complex with a Class 2 Type II Cas endonuclease and target the complex to the target sequence of the target DNA molecule; and wherein the DNA-targeting segment is selected from the group consisting of SEQ ID NOs: 5950-5965, 5966-6025, 6026-612, 6026-613, 6026-614, 6026-615, 6026-616, 6026-617, 6026-618, 6026-619 ...

[0013] In some embodiments, the protein-binding segment comprises a sequence having at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 contiguous nucleotides of any one of SEQ ID NO: 1, 6122-6152, 6153-6181, or 6182-6256. In some embodiments, the protein-binding segment comprises a sequence having at least 85% identity to any one of SEQ ID NO: 5466 or SEQ ID NO: 6304.

[0064] In some aspects, the present disclosure provides a system for generating edited immune cells, the system comprising: (a) an RNA-guided endonuclease; (b) an engineered guide ribonucleic acid polynucleotide of claim 97 configured to bind to the RNA-guided endonuclease; and (c) a single-stranded or double-stranded DNA repair template comprising a first homology arm and a second homology arm flanking a sequence encoding a chimeric antigen receptor (CAR). In some embodiments, the cell is a peripheral blood mononuclear cell, a T cell, a NK cell, a hematopoietic stem cell (HSCT), or a B cell, or any combination thereof. In some aspects, the RNA-guided endonuclease is a Class II Type II Cas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain. In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO:421 or SEQ ID NO:423.

[0065]

[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. Incorporation by Reference

[0066] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]

[0067] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG.").

[0068] [Figure 1] Figure 1 shows the typical organization of different classes and types of CRISPR / Cas loci. [Figure 2] Figure 2 shows the structure of the natural Class 2 / Type II crRNA / tracrRNA pair compared to the hybrid sgRNA to which both are bound. [Figure 3] FIG. 3 shows a schematic diagram illustrating the organization of CRISPR loci, which encode enzymes from the MG1 family. [Figure 4] FIG. 4 shows a schematic diagram illustrating the organization of CRISPR loci, which encode enzymes from the MG2 family. [Figure 5] FIG. 5 shows a schematic diagram illustrating the organization of CRISPR loci, which encode enzymes from the MG3 family. [Figure 6] Figure 6 shows a structure-based alignment of the enzyme of the present disclosure (MG1-1) to Cas9 from Staphylococcus aureus (SEQ ID NO: 5613). Predicted essential residues for function are called out below the sequence. Conserved residues are highlighted in black. [Figure 7] Figure 7 shows a structure-based alignment of the enzyme of the present disclosure (MG2-1) to Cas9 from Staphylococcus aureus (SEQ ID NO: 5613). Predicted essential residues for function are called out below the sequence. Conserved residues are highlighted in black. [Figure 8] Figure 8 shows a structure-based alignment of the enzyme of the present disclosure (MG3-1) to Cas9 from Actinomyces naeslundii (SEQ ID NO: 5614). Predicted essential residues for function are called out below the sequence. Conserved residues are highlighted in black. [Figure 9A] FIG. 9A shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9B] FIG. 9B shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9C] FIG. 9C shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9D] FIG. 9D shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9E] FIG. 9E shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9F]FIG. 9F shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9G] FIG. 9G shows a structure-based alignment of the MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). [Figure 9H] Figure 9H shows a structure-based alignment of MG1 family enzymes MG1-1 to MG1-6 (SEQ ID NOs: 5, 6, 9, 1, 2, and 3). Predicted essential residues for function are called out below the sequences. Conserved residues are highlighted in black. [Figure 10] Figure 10 shows in vitro cleavage of DNA by MG1-4 in complex with its corresponding sgRNA containing targeting sequences of various lengths. [Figure 11] Figure 11 shows cellular cleavage of E. coli genomic DNA using MG1-4 with its corresponding sgRNA. A dilution series of cells transformed with MG1-4 is shown with a target or non-target spacer (top). The top panel shows quantified data, where the left bar represents a non-target sgRNA and the right bar represents a target sgRNA. [Figure 12] Figure 12 shows cellular indel formation produced by transfection of HEK cells with the MG1-4 or MG1-6 constructs described in Example 11, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting various locations in the human genome. [Figure 13] Figure 13 shows the in vitro cleavage of DNA by MG3-6 in complex with its corresponding sgRNA containing targeting sequences of various lengths. [Figure 14] Figure 14 shows cellular cleavage of E. coli genomic DNA using MG3-7 with its corresponding sgRNA. A dilution series of cells transformed with MG3-7 is shown with a target or non-target spacer (top). The bottom panel shows quantified data, where the left bar represents a non-target sgRNA and the right bar represents a target sgRNA. [Figure 15]Figure 15 shows cellular indel formation produced by transfection of HEK cells with the MG3-7 constructs described in Example 13, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting various locations in the human genome. [Figure 16] Figure 16 shows in vitro cleavage of DNA by MG15-1 in complex with its corresponding sgRNA containing targeting sequences of various lengths. [Figure 17] Figure 17 shows an agarose gel showing the results of PAM vector library cleavage in the presence of TXTL extracts containing various MG family nucleases and their corresponding tracrRNAs or sgRNAs. [Figure 18] Figure 18 shows an agarose gel showing the results of PAM vector library cleavage in the presence of TXTL extracts containing various MG family nucleases and their corresponding tracrRNAs or sgRNAs. [Figure 19] Figure 19 shows an agarose gel showing the results of PAM vector library cleavage in the presence of TXTL extracts containing various MG family nucleases and their corresponding tracrRNAs or sgRNAs. [Figure 20] Figure 20 shows an agarose gel showing the results of PAM vector library cleavage in the presence of TXTL extracts containing various MG family nucleases and their corresponding tracrRNAs or sgRNAs. [Figure 21] Figure 21 shows the predicted structures (e.g., predicted as in Example 7) of the corresponding sgRNAs of the MG enzymes described herein. [Figure 22] Figure 22 shows the predicted structures (e.g., predicted as in Example 7) of the corresponding sgRNAs of the MG enzymes described herein. [Figure 23] Figure 23 shows the predicted structures (e.g., predicted as in Example 7) of the corresponding sgRNAs of the MG enzymes described herein. [Figure 24]Figure 24 shows the predicted structures (e.g., predicted as in Example 7) of the corresponding sgRNAs of the MG enzymes described herein. [Figure 25] Figure 25 shows the predicted structures (e.g., predicted as in Example 7) of the corresponding sgRNAs of the MG enzymes described herein. [Figure 26] Figure 26 shows the predicted structures (e.g., predicted as in Example 7) of the corresponding sgRNAs of the MG enzymes described herein. [Figure 27] FIG. 27 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 28] FIG. 28 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 29] FIG. 29 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 30] Figure 30 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 31] FIG. 31 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 32] FIG. 32 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 33] Figure 33 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 34]Figure 34 shows cellular cleavage of E. coli genomic DNA using MG2-7 with its corresponding sgRNA. A dilution series of cells transformed with MG2-7 is shown with a target or non-target spacer (top). The bottom panel shows quantified data, where the right bar represents a non-target sgRNA and the left bar represents a target sgRNA. [Figure 35] Figure 35 shows cellular cleavage of E. coli genomic DNA using MG14-1 with its corresponding sgRNA. A dilution series of cells transformed with MG14-1 is shown with a target or non-target spacer (top). The bottom panel shows quantified data, where the right bar represents a non-target sgRNA and the left bar represents a target sgRNA. [Figure 36] Figure 36 shows cellular cleavage of E. coli genomic DNA using MG15-1 with its corresponding sgRNA. A dilution series of cells transformed with MG15-1 is shown with a target or non-target spacer (top). The bottom panel shows quantified data, where the right bar represents a non-target sgRNA and the left bar represents a target sgRNA. [Figure 37] Figure 37 shows cellular indel formation produced by transfection of HEK cells with the MG1-4, MG1-6, and MG1-7 constructs described in Example 11, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting different locations in the human genome. [Figure 38] Figure 38 shows cellular indel formation produced by transfection of HEK cells with the MG1-4, MG1-6, and MG1-7 constructs described in Example 11, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting different locations in the human genome. [Figure 39] Figure 39 shows cellular indel formation produced by transfection of HEK cells with the MG1-4, MG1-6, and MG1-7 constructs described in Example 11, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting different locations in the human genome. [Figure 40] Figure 40 shows cellular indel formation produced by transfection of HEK cells with the MG3-6, MG3-7, and MG3-8 constructs described in Example 13, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting different locations in the human genome. [Figure 41] Figure 41 shows cellular indel formation produced by transfection of HEK cells with the MG3-6, MG3-7, and MG3-8 constructs described in Example 13, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting different locations in the human genome. [Figure 42] Figure 42 shows cellular indel formation produced by transfection of HEK cells with the MG3-6, MG3-7, and MG3-8 constructs described in Example 13, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting different locations in the human genome. [Figure 43] Figure 43 shows cellular indel formation produced by transfection of HEK cells with the MG14-1 constructs described in Example 14, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting various locations in the human genome. [Figure 44] Figure 44 shows cellular indel formation produced by transfection of HEK cells with the MG18-1 constructs described in Example 17, along with their corresponding sgRNAs containing a variety of different targeting sequences targeting various locations in the human genome. [Figure 45] Figure 45 shows the environmental distribution of nucleases described herein. Protein lengths are shown for representatives of selected protein families. Color indicates the environment or environment type in which each protein was identified. [Figure 46]Figure 46 shows the predicted catalytic residues of the nucleases described herein. Protein length is shown for representatives of selected protein families. Color indicates the number of predicted catalytic residues for each protein. For the effector enzymes described herein, six catalytic residues corresponding to the HNH and RuvC domains were searched. [Figure 47] Figure 47 shows the candidate activity versus protein length of the nucleases described herein. [Figure 48] Figure 48 shows the number of predicted catalytic residues for the nucleases described herein. [Figure 49] Figure 49 shows a table of various characteristic information for the selected nucleases described herein. [Figure 50] Figure 50 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 51] Figure 51 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 52] Figure 52 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 53] Figure 53 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 54] Figure 54 shows a seqLogo representation of a PAM sequence derived by NGS as described herein (e.g., as described in Example 6). [Figure 55] Figure 55 shows guide RNA screening in TRAC using MG3-6 and MG3-8. For the top panel (MG3-6), the x-axis numbers refer to spacers corresponding to SEQ ID NOs: 5950 to 5958. For the bottom panel (MG3-8), the x-axis numbers refer to spacers corresponding to SEQ ID NOs: 5959 to 5965. [Figure 56] Figure 56 shows the activity (indels%) of MG3-6 with guide RNAs of various core sequences, lengths, and doses. [Figure 57] Figure 57 shows the activity (indel %) of MG3-8 using guide RNAs of various sequences and lengths. [Figure 58] Figure 58 shows the activity (indel %) of MG3-6 with TRAC guide 6 and MG3-8 with TRAC guide 8. [Figure 59] Figure 59 shows the effect of MG3-6 with TRAC6 guide RNA on T cell receptor expression by flow cytometry. There was no change in survival rate after editing. [Figure 60] Figure 60 shows increased TRAC editing efficiency with higher amounts of gRNA. [Figure 61] Figure 61 shows how TCR expression can be eliminated and replaced with CAR expression. [Figure 62] Figure 62 shows targeted CAR integration with MG3-6. [Figure 63] Figure 63 shows GR(NR3Cl) editing by MG3-6 with different guide RNAs targeting different exons of the NR3Cl gene. [Figure 64] Figure 64 shows GR(NR3Cl) editing by MG3-8 with different guide RNAs targeting different exons of the NR3Cl gene. [Figure 65] Figure 65 compares GR editing with two MG3-6 batches and various guide RNAs. [Figure 66] Figure 66 shows the process of how gene editing can be used to create allogeneic CAR-NK cells. [Figure 67] Figure 67 shows TRAC editing using MG3-6 with TRAC6 guide RNA. [Figure 68] Figure 68 shows MG3-6 CAR expression (Y axis) in CD56+ NK cells by flow cytometry. [Figure 69]Figure 69 shows CD38 editing in primary NK cells using MG3-6 and MG3-8 with various guide RNAs. [Figure 70] Figure 70 shows TRAC editing in hematopoietic stem cells by MG3-6 and MG3-8 with various guide RNAs. [Figure 71] Figure 71 shows TRAC editing in B cells by MG3-6 with TRAC guide 6 using two different buffers. [Figure 72] Figure 72 shows the consensus PAM sequences of MG48-1 (A) and MG48-3 (B), determined by the method of Example 25. [Figure 73] Figure 73 shows RNAseq mapping of MG48-1 (A) and MG48-3 (B) as performed by the method of Example 25, highlighting the sequenced tracr region.

[0069] Brief Description of Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems of the present disclosure. Below are exemplary descriptions of the sequences therein.

[0070] MG1

[0071] SEQ ID NOs: 1 to 319 show the full-length peptide sequences of MG1 nuclease.

[0072] SEQ ID NOs: 1827 to 2140 show the peptide sequences of the RuvC_III domain of the above-mentioned MG1 nuclease.

[0073] SEQ ID NOs: 3638 to 3955 show peptides of the HNH domain of the above-mentioned MG1 nuclease.

[0074] SEQ ID NOs: 5476 to 5479 show the nucleotide sequences of MG1 tracrRNAs derived from the same loci as the above-mentioned MG1 nucleases (e.g., the same loci as SEQ ID NOs: 1 to 4, respectively).

[0075] SEQ ID NOs: 5461-5464 set forth the nucleotide sequences of sgRNAs engineered to function with MG1 nuclease (e.g., SEQ ID NOs: 1-4, respectively), where Ns represents a nucleotide in the targeting sequence.

[0076] SEQ ID NOs: 5572 to 5575 show the nucleotide sequences of E. coli codon-optimized coding sequences of MG1 family enzymes (SEQ ID NOs: 1 to 4).

[0077] SEQ ID NOs: 5588-5589 show the nucleotide sequences of human codon-optimized coding sequences for MG1 family enzymes (SEQ ID NOs: 1 and 3).

[0078] SEQ ID NOs: 5616 to 5632 show peptide motifs characteristic of MG1 family enzymes.

[0079] MG2

[0080] SEQ ID NOs: 320 to 420 show the full-length peptide sequences of MG2 nuclease.

[0081] SEQ ID NOs: 2141 to 2241 show the peptide sequences of the RuvC_III domain of the above-mentioned MG2 nuclease.

[0082] SEQ ID NOs: 3955 to 4055 show peptides of the HNH domain of the above-mentioned MG2 nuclease.

[0083] SEQ ID NOs: 5490-5494 show the nucleotide sequences of MG2 tracrRNAs derived from the same loci as the MG2 nucleases described above (e.g., the same loci as SEQ ID NOs: 320, 321, 323, 325, and 326, respectively).

[0084] SEQ ID NO: 5465 shows the nucleotide sequence of an sgRNA engineered to function with MG2 nuclease (e.g., SEQ ID NO: 321 above).

[0085] SEQ ID NOs: 5572 to 5575 show the nucleotide sequences of E. coli codon-optimized coding sequences for MG2 family enzymes.

[0086] SEQ ID NOs: 5631 to 5638 show peptide sequences characteristic of MG2 family enzymes.

[0087] MG3

[0088] SEQ ID NOs: 421 to 431 show the full-length peptide sequences of MG3 nuclease.

[0089] SEQ ID NOs: 2242 to 2252 show the peptide sequences of the RuvC_III domain of the above-mentioned MG3 nuclease.

[0090] SEQ ID NOs: 4056 to 4066 show peptides of the HNH domain of the above-mentioned MG3 nuclease.

[0091] SEQ ID NOs: 5495-5502 show the nucleotide sequences of MG3 tracrRNAs derived from the same loci as the MG3 nucleases described above (e.g., the same loci as SEQ ID NOs: 421-428, respectively).

[0092] SEQ ID NOs: 5466-5467 show the nucleotide sequences of sgRNAs engineered to function with MG3 nuclease (e.g., SEQ ID NOs: 421-423).

[0093] SEQ ID NOs: 5578-5580 show the nucleotide sequences of E. coli codon-optimized coding sequences for MG3 family enzymes.

[0094] SEQ ID NOs: 5639 to 5648 show peptide sequences characteristic of MG3 family enzymes.

[0095] MG4

[0096] SEQ ID NOs: 432 to 660 show the full-length peptide sequences of MG4 nuclease.

[0097] SEQ ID NOs: 2253 to 2481 show the peptide sequences of the RuvC_III domain of the above-mentioned MG4 nuclease.

[0098] SEQ ID NOs: 4067 to 4295 show peptides of the HNH domain of the above-mentioned MG4 nuclease.

[0099] SEQ ID NO: 5503 shows the nucleotide sequence of the MG4 tracrRNA, which is derived from the same locus as the MG4 nuclease described above.

[0100] SEQ ID NO: 5468 shows the nucleotide sequence of an sgRNA engineered to function with MG4 nuclease.

[0101] SEQ ID NO: 5649 shows a peptide sequence characteristic of MG4 family enzymes.

[0102] MG6

[0103] SEQ ID NOs: 661 to 668 show the full-length peptide sequences of MG6 nuclease.

[0104] SEQ ID NOs: 2482 to 2489 show the peptide sequences of the RuvC_III domain of the above-mentioned MG6 nuclease.

[0105] SEQ ID NOs: 4296 to 4303 show peptides of the HNH domain of the above-mentioned MG3 nuclease.

[0106] MG7

[0107] SEQ ID NOs: 669 to 677 show the full-length peptide sequences of MG7 nuclease.

[0108] SEQ ID NOs: 2490 to 2498 show the peptide sequences of the RuvC_III domain of the above-mentioned MG7 nuclease.

[0109] SEQ ID NOs: 4304 to 4312 show peptides of the HNH domain of the above-mentioned MG3 nuclease.

[0110] SEQ ID NO: 5504 shows the nucleotide sequence of the MG7 tracrRNA, which is derived from the same locus as the MG7 nuclease described above.

[0111] MG14

[0112] SEQ ID NOs: 678 to 929 show the full-length peptide sequences of MG14 nuclease.

[0113] SEQ ID NOs: 2499 to 2750 show the peptide sequences of the RuvC_III domain of the above-mentioned MG14 nuclease.

[0114] SEQ ID NOs: 4313 to 4564 show peptides of the HNH domain of the above-mentioned MG14 nuclease.

[0115] SEQ ID NO: 5505 shows the nucleotide sequence of the MG14 tracrRNA, which is derived from the same locus as the MG14 nuclease described above.

[0116] SEQ ID NO: 5581 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG14 family enzyme.

[0117] SEQ ID NOs: 5650 to 5667 show peptide sequences characteristic of MG14 family enzymes.

[0118] MG15

[0119] SEQ ID NOs: 930 to 1092 show the full-length peptide sequences of MG15 nuclease.

[0120] SEQ ID NOs: 2751 to 2913 show the peptide sequences of the RuvC_III domain of the above-mentioned MG15 nuclease.

[0121] SEQ ID NOs: 4565 to 4727 show peptides of the HNH domain of the above-mentioned MG15 nuclease.

[0122] SEQ ID NO: 5506 shows the nucleotide sequence of the MG15 tracrRNA, which is derived from the same locus as the MG15 nuclease described above.

[0123] SEQ ID NO: 5470 shows the nucleotide sequence of an sgRNA engineered to function with MG15 nuclease.

[0124] SEQ ID NO: 5582 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG15 family enzyme.

[0125] SEQ ID NOs: 5668 to 5675 show peptide sequences characteristic of MG15 family enzymes.

[0126] MG16

[0127] SEQ ID NOs: 1093 to 1353 show the full-length peptide sequences of MG16 nuclease.

[0128] SEQ ID NOs: 2914 to 3174 show the peptide sequences of the RuvC_III domain of the above-mentioned MG16 nuclease.

[0129] SEQ ID NOs: 4728 to 4988 show peptides of the HNH domain of the above-mentioned MG16 nuclease.

[0130] SEQ ID NO: 5507 shows the nucleotide sequence of the MG16 tracrRNA, which is derived from the same locus as the MG3 nuclease described above.

[0131] SEQ ID NO: 5471 shows the nucleotide sequence of an sgRNA engineered to function with MG16 nuclease.

[0132] SEQ ID NO: 5583 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG16 family enzyme.

[0133] SEQ ID NOs: 5676 to 5678 show peptide sequences characteristic of MG16 family enzymes.

[0134] MG18

[0135] SEQ ID NOs: 1354 to 1511 show the full-length peptide sequences of MG18 nuclease.

[0136] SEQ ID NOs: 3175 to 3330 show the peptide sequences of the RuvC_III domain of the above-mentioned MG18 nuclease.

[0137] SEQ ID NOs: 4989 to 5146 show peptides of the HNH domain of the above-mentioned MG18 nuclease.

[0138] SEQ ID NO: 5508 shows the nucleotide sequence of the MG18 tracrRNA, which is derived from the same locus as the MG18 nuclease described above.

[0139] SEQ ID NO: 5472 shows the nucleotide sequence of an sgRNA engineered to function with MG18 nuclease.

[0140] SEQ ID NO: 5584 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG18 family enzyme.

[0141] SEQ ID NOs: 5679 to 5686 show peptide sequences characteristic of MG18 family enzymes.

[0142] MG21

[0143] SEQ ID NOs: 1512 to 1655 show the full-length peptide sequences of MG21 nuclease.

[0144] SEQ ID NOs: 3331 to 3474 show the peptide sequences of the RuvC_III domain of the above-mentioned MG21 nuclease.

[0145] SEQ ID NOs: 5147 to 5290 show peptides of the HNH domain of the above-mentioned MG21 nuclease.

[0146] SEQ ID NO: 5509 shows the nucleotide sequence of the MG21 tracrRNA, which is derived from the same locus as the MG21 nuclease described above.

[0147] SEQ ID NO: 5473 shows the nucleotide sequence of an sgRNA engineered to function with MG21 nuclease.

[0148] SEQ ID NO: 5585 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG21 family enzyme.

[0149] SEQ ID NOs: 5687 to 5692 and 5674 to 5675 show peptide sequences characteristic of MG21 family enzymes.

[0150] MG22

[0151] SEQ ID NOs: 1656 to 1755 show the full-length peptide sequences of MG22 nuclease.

[0152] SEQ ID NOs: 3475 to 3568 show the peptide sequences of the RuvC_III domain of the above-mentioned MG22 nuclease.

[0153] SEQ ID NOs: 5291 to 5389 show peptides of the HNH domain of the above-mentioned MG22 nuclease.

[0154] SEQ ID NO: 5510 shows the nucleotide sequence of the MG22 tracrRNA, which is derived from the same locus as the MG22 nuclease described above.

[0155] SEQ ID NO: 5474 shows the nucleotide sequence of an sgRNA engineered to function with MG22 nuclease.

[0156] SEQ ID NO: 5586 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG22 family enzyme.

[0157] SEQ ID NOs: 5694 to 5699 show peptide sequences characteristic of MG22 family enzymes.

[0158] MG23

[0159] SEQ ID NOs: 1756 to 1826 show the full-length peptide sequences of MG23 nuclease.

[0160] SEQ ID NOs: 3569 to 3637 show the peptide sequences of the RuvC_III domain of the above-mentioned MG23 nuclease.

[0161] SEQ ID NOs: 5390 to 5460 show peptides of the HNH domain of the above-mentioned MG23 nuclease.

[0162] SEQ ID NO: 5511 shows the nucleotide sequence of the MG23 tracrRNA, which is derived from the same locus as the MG23 nuclease described above.

[0163] SEQ ID NO: 5475 shows the nucleotide sequence of an sgRNA engineered to function with MG23 nuclease.

[0164] SEQ ID NO: 5587 shows the nucleotide sequence of an E. coli codon-optimized coding sequence for an MG23 family enzyme.

[0165] SEQ ID NOs: 5700 to 5717 show peptide sequences characteristic of MG23 family enzymes.

[0166] MG40

[0167] SEQ ID NOs: 5718 to 5750 show the full-length peptide sequences of MG40 nuclease.

[0168] SEQ ID NOs: 5847-5852 show protospacer adjacent motifs associated with MG40 nuclease.

[0169] SEQ ID NOs: 5862-5873 show the nucleotide sequences of sgRNAs engineered to function with MG40 nuclease.

[0170] MG47

[0171] SEQ ID NOs: 5751 to 5768 show the full-length peptide sequences of MG47 nuclease.

[0172] SEQ ID NOs: 5853-5854 show protospacer adjacent motifs associated with MG47 nuclease.

[0173] SEQ ID NOs: 5878-5881 show the nucleotide sequences of sgRNAs engineered to function with MG47 nuclease.

[0174] MG48

[0175] SEQ ID NOs: 5769 to 5804 show the full-length peptide sequences of MG48 nuclease.

[0176] SEQ ID NOs: 5855-5856 show protospacer adjacent motifs associated with MG48 nuclease.

[0177] SEQ ID NOs: 5886, 5890, and 5893 show the nucleotide sequence of the MG48 tracrRNA, which is derived from the same locus as the MG48 nuclease described above.

[0178] SEQ ID NOs: 5887, 5891, and 5894 show CRISPR repeats associated with the MG48 nuclease described herein.

[0179] SEQ ID NOs: 5888-5889, 5892, and 5895-5896 show putative sgRNAs designed to function with MG48 nuclease.

[0180] MG49

[0181] SEQ ID NOs: 5805 to 5823 show the full-length peptide sequences of MG49 nuclease.

[0182] SEQ ID NOs: 5857-5858 show protospacer adjacent motifs associated with MG49 nuclease.

[0183] SEQ ID NOs: 5862-5873 show the nucleotide sequences of sgRNAs engineered to function with MG40 nuclease.

[0184] SEQ ID NOs: 5876-5877 show the nucleotide sequences of sgRNAs engineered to function with MG49 nuclease.

[0185] MG50

[0186] SEQ ID NOs: 5824 to 5826 show the full-length peptide sequences of MG50 nuclease.

[0187] SEQ ID NO: 5859 shows the protospacer adjacent motif associated with MG50 nuclease.

[0188] SEQ ID NOs: 5884-5885 show the nucleotide sequences of sgRNAs engineered to function with MG50 nuclease.

[0189] MG51

[0190] SEQ ID NOs: 5827 to 5830 show the full-length peptide sequences of MG51 nuclease.

[0191] SEQ ID NO: 5860 shows the protospacer adjacent motif associated with MG51 nuclease.

[0192] SEQ ID NOs: 5882-5883 show the nucleotide sequences of sgRNAs engineered to function with MG51 nuclease.

[0193] MG52

[0194] SEQ ID NOs: 5831 to 5846 show the full-length peptide sequences of MG52 nuclease.

[0195] SEQ ID NO: 5861 shows the protospacer adjacent motif associated with MG52 nuclease.

[0196] SEQ ID NOs: 5874-5875 show the nucleotide sequences of sgRNAs engineered to function with MG42 nuclease. DETAILED DESCRIPTION OF THE INVENTION

[0197] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0198] The practice of some methods disclosed herein employs, unless otherwise indicated, techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, e.g., Sambrook and Green, "Molecular Cloning: A Laboratory Manual," 4th Edition (2012); "The Series Current Protocols in Molecular Biology" (eds. F.M.A.usubel et al.); "The Series Methods in Enzymology" (Academic Press, Inc.); "PCR 2: A Practical Approach" (eds. M.J. MacPherson, B.D. Hames, and G.R. Taylor (1995)); Harlow and Lane (eds. 1988); "Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications," 6th Edition (ed. R.I. Freshney (2010)), which is incorporated herein by reference in its entirety.

[0199] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof, are used in either the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."

[0200] The terms "about" or "approximately" mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, in accordance with practice in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0201] As used herein, "cell" generally refers to a biological cell. A cell may be the basic structural, functional, and / or biological unit of an organism. A cell may be derived from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of single-celled eukaryotes, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, liverworts, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, and Sargassum patens), and cells from plants (e.g., plant cells, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, liverworts, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, and Sargassum patens). C. Agardh), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., Drosophila, Cnidaria, Echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). The cells may not be derived from a naturally occurring organism (e.g., the cells may be synthetically produced, sometimes referred to as artificial cells).

[0202] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides may be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates such as adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. The term "nucleotide" as used herein may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as with a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels.Fluorescent labels for nucleotides may include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer, Foster City, California; FluoroLink deoxyribonucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, available from Amersham, Arlington Heights, Illinois; and Boehringer Ingelheim, Indianapolis, Indiana. Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Mannheim; and BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green available from Molecular Probes, Eugene, Oregon. Examples of suitable nucleotides include 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled or marked by chemical modification. The chemically modified single nucleotide can be biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0203] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in either single-, double-, or multi-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can be present in a cell-free environment. A polynucleotide can be a gene or a fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three-dimensional structure and perform any function. A polynucleotide can contain one or more analogs (e.g., altered backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure can be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acids, heterologous nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queusine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides may be interrupted by non-nucleotide components.

[0204] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule can be a gene sequence encoding an entire protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, pp. 18.1-18.88.

[0205] The terms "peptide," "polypeptide," and "protein" are generally used interchangeably herein to refer to a polymer of at least two amino acid residues linked by peptide bond(s). The term does not imply a particular length of the polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical synthesis, enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers and amino acid polymers comprising at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary and / or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation, such as conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include natural amino acids and unnatural amino acids, which are chemically modified to include a non-naturally occurring group or chemical moiety on the amino acid. Amino acid analogs may refer to amino acid derivatives. The term "amino acid" includes both D- and L-amino acids.

[0206] As used herein, "non-naturally occurring" can generally refer to a nucleic acid or polypeptide sequence that is not found in a naturally occurring nucleic acid or protein. Non-naturally occurring can refer to an affinity tag. Non-naturally occurring can refer to a fusion. Non-naturally occurring can refer to a naturally occurring nucleic acid or polypeptide sequence that includes mutations, insertions, and / or deletions. A non-naturally occurring sequence can exhibit and / or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) that can also be exhibited by the nucleic acid and / or polypeptide sequence to which the non-naturally occurring sequence is fused. A non-naturally occurring nucleic acid or polypeptide sequence can be linked to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid and / or polypeptide sequence that encodes the chimeric nucleic acid and / or polypeptide.

[0207] As used herein, the term "promoter" generally refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotide at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, which promote the binding of RNA polymerase to DNA, resulting in gene transcription. A "basal promoter," also referred to as a "core promoter," may generally refer to a promoter that contains all the basic elements necessary to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically, but not necessarily, contain a TATA box and / or a CAAT box.

[0208] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0209] As used herein, "operably linked," "operable linkage," "operatively linked," or their grammatical equivalents generally refer to the juxtaposition of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., wherein the elements are in a relationship permitting them to operate in the expected manner. By way of example, a regulatory element, which may include a promoter and / or enhancer sequence, is operably linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. Intervening residues can be present between the regulatory element and the coding region so long as this functional relationship is maintained.

[0210] As used herein, a "vector" generally refers to a polymer or assembly of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally contains genetic elements, such as regulatory elements, operably linked to a gene to promote expression of the gene in a target.

[0211] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of regulatory elements and one or more genes to which they are operably linked for expression.

[0212] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) that is substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be its ability to affect expression in a manner known to be attributed to the full-length sequence.

[0213] As used herein, an "engineered" object generally refers to an object that has been altered by human intervention. By non-limiting example, a nucleic acid can be modified by changing its sequence to a sequence that does not occur in nature. A nucleic acid can be modified by ligating it to a nucleic acid with which it is not naturally associated, such that the ligated product has a function not present in the original nucleic acid. Engineered nucleic acids can be synthesized in vitro using sequences that do not occur in nature. A protein can be modified by changing its amino acid sequence to a sequence that does not occur in nature. An engineered protein can acquire a new function or property. An "engineered" system includes at least one engineered component.

[0214] As used herein, "synthetic" and "artificial" are used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.

[0215] As used herein, the term "tracrRNA" or "tracr sequence" can generally refer to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc., or SEQ ID NOs: 5476-5511). A tracrRNA can refer to a nucleic acid having up to about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.). A tracrRNA can also refer to modified forms of a tracrRNA that can include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeras. A tracrRNA can also refer to a nucleic acid that can be at least about 60% identical to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.) over a stretch of at least six contiguous nucleotides. For example, the tracrRNA sequence can be at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.) over a stretch of at least six contiguous nucleotides. Type II tracrRNA sequences can be predicted on genomic sequences by identifying regions that share complementarity with portions of repeat sequences in adjacent CRISPR arrays.

[0216] As used herein, "guide nucleic acid" can generally refer to a nucleic acid that can hybridize to another nucleic acid. A guide nucleic acid can be RNA. A guide nucleic acid can be DNA. A guide nucleic acid can be programmed to site-specifically bind to a sequence of a nucleic acid. A targeted nucleic acid, or target nucleic acid, can comprise nucleotides. A guide nucleic acid can comprise nucleotides. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to and hybridizes with a guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore may not be complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can comprise a polynucleotide strand and can be referred to as a "single guide nucleic acid." A guide nucleic acid can comprise two polynucleotide strands and can be referred to as a "double guide nucleic acid." Unless otherwise specified, the term "guide nucleic acid" can be inclusive and refer to both single guide nucleic acids and double guide nucleic acids. A guide nucleic acid may comprise a segment that may be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence." The nucleic acid targeting segment may comprise a subsegment that may be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."

[0217] The term "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are the same or have a specified percentage of identical amino acid residues or nucleotides when compared over a local or global comparison window and aligned for maximum correspondence as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP, using parameters of a word length of 3, an expectation I of 10, and a BLOSUM62 scoring matrix setting gap cost at 11, an extension of 1, and a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues. BLASTP, with parameters of word length (W) of 2, expectation (E) of 1,000,000, and the PAM30 scoring matrix setting gap, has a cost of 9 to open a gap and 1 to extend a gap for sequences of less than 30 residues (these are the default parameters for BLAST in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); the Smith-Waterman homology search algorithm with parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE with default parameters; MAFFT with parameters of 2 retree and 1,000; Novafold with default parameters; and HMMER with default parameters and CLUSTALW with parameters of hmmalign.

[0218] The present disclosure includes variants of any of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length for each other. Additionally or alternatively, conservative substitutions can be identified by comparing aligned sequences of homologous proteins from different species, by mutating amino acid residues (e.g., non-conserved residues) between species without altering the basic function of the encoded protein. Such conservatively substituted variants can include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of the endonuclease protein sequences described herein (e.g., an MG1, MG2, MG3, MG4, MG6, MG7, MG14, MG15, MG16, MG18, MG21, MG22, or MG23 family endonucleases described herein). In some embodiments, such conservative substitution variants are functional variants. Such functional variants can include sequences with substitutions that do not destroy the activity of key active site residues of the endonuclease. In some embodiments, a functional variant of any of the proteins described herein lacks at least one substitution of a conserved or functional residue designated in Figure 6, Figure 7, Figure 8, Figure 9A, Figure 9B, Figure 9C, Figure 9D, Figure 9E, Figure 9F, Figure 9G, or Figure 9H.In some embodiments, a functional variant of any of the proteins described herein lacks all of the conserved or functional residue substitutions designated in Figure 6, Figure 7, Figure 8, Figure 9A, Figure 9B, Figure 9C, Figure 9D, Figure 9E, Figure 9F, Figure 9G, or Figure 9H.

[0219] Conservative substitution tables providing functionally similar amino acids are available in various references (e.g., Creighton, "Proteins: Structures and Molecular Properties" (W.H. Freeman & Co.; 2nd ed. (December 1993)). The following eight groups each contain amino acids that are conservative substitutions for one another: (1) alanine (A), glycine (G); (2) aspartic acid (D), glutamic acid (E); (3) asparagine (N), glutamine (Q); (4) arginine (R), lysine (K); (5) isoleucine (I), leucine (L), methionine (M), valine (V); (6) phenylalanine (F), tyrosine (Y), tryptophan (W); (7) serine (S), threonine (T); and (8) Cysteine ​​(C), Methionine (M)

[0220] As used herein, the term "RuvC_III domain" generally refers to the third, non-contiguous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three non-contiguous segments, RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to known domain sequences, by structural alignment to proteins with annotated domains, or by comparison with Hidden Markov Models (HMMs) constructed based on known domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).

[0221] As used herein, the term "HNH domain" generally refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to known domain sequences, by structural alignment to proteins with annotated domains, or by comparison to a Hidden Markov Model (HMM) constructed based on known domain sequences (e.g., Pfam HMM PF01844 for domain HNH).

[0222] overview

[0223] The discovery of novel Cas enzymes with unique functionality and structure could further disrupt deoxyribonucleic acid (DNA) editing technologies, offering the potential to improve their speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microorganisms and the sheer diversity of microbial species, there are relatively few functionally characterized CRISPR / Cas enzymes in the literature. This is because the vast number of microbial species may not be easily cultivated under laboratory conditions. Metagenomic sequencing from natural environmental niches representing a large number of microbial species could dramatically increase the number of known new CRISPR / Cas systems and offer the potential to accelerate the discovery of novel oligonucleotide editing functions. A recent example of the utility of such an approach was demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.

[0224] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30–40 bp) separated by an equally short spacer sequence that encodes an RNA-based targeting element, and (ii) an ORF encoding a Cas, which encodes a nuclease polypeptide guided by the RNA-based targeting element along with accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6–8 nucleic acids of the target (the target seed) and the crRNA guide, and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM is typically a sequence not commonly represented in the host genome). Depending on the exact function and organization of the system, CRISPR-Cas systems are generally organized into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities.

[0225] Class I CRISPR-Cas systems have large multi-subunit effector complexes and include type I (Type I), type III (Type III), and type IV (Type IV).

[0226] Type I CRISPR-Cas systems are considered to have intermediate complexity in terms of components. In Type I CRISPR-Cas systems, an array of RNA targeting elements transcribes nucleic acid targets as long precursor crRNAs (pre-crRNAs) that are processed at repeat elements to release short mature crRNAs that direct the nuclease complex to the nucleic acid target when followed by a suitable short consensus sequence called a protospacer adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nuclease primarily functions as a DNA nuclease.

[0227] Type III CRISPR systems can be characterized by the presence of a central nuclease known as Cas10, along with a repeat-associated mysterious protein (RAMP) containing Csm or Cmr protein subunits. Similar to type I systems, mature crRNA is processed from pre-crRNA using a Cas6-like enzyme. Unlike type I and type II systems, type III systems appear to target and cleave DNA-RNA duplexes (such as the DNA strand that is being used as a template for RNA polymerase).

[0228] Type IV CRISPR-Cas systems possess an effector complex consisting of a highly reduced large subunit nuclease (csf1) and two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) families, and optionally a gene for a predicted small subunit. Such systems are commonly found on endogenous plasmids.

[0229] Class II CRISPR-Cas systems generally have a single polypeptide multidomain nuclease effector and include type II (Type II), type V (Type V), and type VI (Type VI).

[0230] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, processing of the CRISPR array into mature crRNA does not require the presence of a specialized endonuclease subunit, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence. This tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate the mature effector enzyme loaded with both the tracrRNA and crRNA. Cas II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-like endonuclease domain adopting an RNase H fold with an unrelated HNH nuclease domain inserted within the RuvC-like nuclease domain fold. The RuvC-like domain is responsible for cleaving the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleaving the displaced DNA strand.

[0231] Type V CRISPR-Cas systems feature a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (if not all) type V CRISPR systems use a tracrRNA to process pre-crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems can cleave pre-crRNA using the effector nuclease itself. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are also known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-strand nonspecific deoxyribonuclease activity that is activated by the initial crRNA-directed cleavage of the double-stranded target sequence.

[0232] Type VI CRIPSR-Cas systems possess an RNA-guided RNA endonuclease. Instead of a RuvC-like domain, the single polypeptide effector of type VI systems (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems also do not appear to require a tracrRNA to process pre-crRNA into crRNA. However, like type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-strand nonspecific nuclease (ribonuclease) activity that is activated by the initial crRNA-directed cleavage of the target RNA.

[0233] Due to their simpler structure, class II CRISPR-Cas are the most widely adopted in engineering and development for designer nuclease / genome editing applications.

[0234] One early application of such a system for in vitro use can be found in Jinek et al. (Science., August 17, 2012; Vol. 337 (No. 6096): pp. 816-21, incorporated herein by reference in its entirety). Jinek's work first described a system that included (i) recombinantly expressed purified full-length Cas9 (e.g., a class II, type II Cas enzyme) isolated from S. pyogenes SF370, (ii) a purified mature approximately 42-nt fragment carrying a 5' approximately 20-nt fragment complementary to the target DNA sequence where cleavage is desired, followed by a 3' tracr binding sequence (the entire crRNA is in vitro transcribed from a synthetic DNA template bearing a T7 promoter sequence), (iii) purified tracrRNA in vitro transcribed from a synthetic DNA template bearing a T7 promoter sequence, and (iv) Mg. Jinek later described an improved and engineered system in which (ii) the crRNA is linked to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can independently guide Cas9 to its target (compare the top and bottom panels of Figure 2).

[0235] Mali et al. (Science., February 15, 2013; Vol. 339(6121):823-826), incorporated herein by reference in its entirety, later adapted this system for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding a codon-optimized Cas9 (e.g., a Class II Type II Cas enzyme) under a suitable mammalian promoter with a C-terminal nuclear localization sequence (e.g., SV40NLS) and a suitable polyadenylation signal (e.g., a TKpA signal), and (ii) an ORF encoding an sgRNA (a 5' sequence starting with G followed by a 20-nt complementary targeting nucleic acid sequence linked to the 3' tracr binding sequence, a linker, and the tracrRNA sequence) under a suitable polymerase III promoter (e.g., a U6 promoter).

[0236] MG enzyme

[0237] In one aspect, the present disclosure provides engineered nuclease systems discovered by metagenomic sequencing. In some cases, metagenomic sequencing is performed on a sample. In some cases, the sample may be collected from a variety of environments. Such environments may be human microbiota, animal microbiota, high temperature environments, or low temperature environments. Such environments may include sediments. Examples of the types of environments for the engineered nuclease systems described herein can be found in Figure 45.

[0238] MG1 enzyme

[0239] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 1827-2140. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1827-2140. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 1827-2140. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 1827-1831. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1827-1831.In some cases, the endonuclease can comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 1827-1831. In some cases, the endonuclease can comprise a RuvC_III domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1827. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1828. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1829. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1830.In some cases, the endonuclease may comprise a RuvC_III domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1831.

[0240] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638-3955. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638-3955. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3638-3955. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638-3955. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638-3955. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3638-3955. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638-3641. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638-3641. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3638-3641.The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3638. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3639. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3639. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3639. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3640. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3640. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3640. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3641.In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3641. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3641.

[0241] In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-6 or SEQ ID NOs: 9-319. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1-6 or SEQ ID NOs: 9-319. In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-4. In some cases, the endonuclease may comprise a peptide motif substantially identical to any one of SEQ ID NOs: 5615, 5616, or 5617.

[0242] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 1-6 or 9-319, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 1-319. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1 below, or a combination thereof.

[0243] [Table 1]

[0244] In some cases, the endonuclease can be recombinant (e.g., cloned, expressed, and purified by suitable methods, such as expression in E. coli followed by epitope tag purification). In some cases, the endonuclease can be derived from a bacterium having a 16S rRNA gene having at least about 90% identity to any one of SEQ ID NOs: 5592-5595. The endonuclease can be derived from a species having a 16S rRNA gene at least about 80%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of SEQ ID NOs: 5592-5595. The endonuclease can be derived from a species having a 16S rRNA gene substantially identical to any one of SEQ ID NOs: 5592-5595. The endonuclease can be derived from a bacterium belonging to the phylum Verrucomicrobia or Candidatus Peregrinibacteria.

[0245] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0246] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0247] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5476-5489. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5476-5489. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5476-5489. The tracrRNA may comprise any of SEQ ID NOs: 5476-5489.

[0248] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to any one of SEQ ID NOs: 5461-5464. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5461-5464. The sgRNA may comprise a sequence substantially identical to any one of SEQ ID NOs: 5461-5464.

[0249] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0250] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus, it may modify the target nucleic acid locus. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system in a buffer with a nucleic acid comprising the locus of interest. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0251] Where the target nucleic acid locus can be within a cell, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 1827-2140. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to any one of SEQ ID NOS: 5572-5575, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 5572-5575. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0252] In some cases, the disclosure may provide an expression cassette comprising a system disclosed herein or a nucleic acid described herein. In some cases, the expression cassette or nucleic acid may be provided as a vector. In some cases, the expression cassette, nucleic acid, or vector may be provided intracellularly. In some cases, the cell is a bacterial cell having a 16S rRNA gene having at least about 90% (e.g., at least about 99%) identity to any one of SEQ ID NOs: 5592-5595.

[0253] MG2 enzyme

[0254] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2141-2241. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2141-2241. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2141-2142. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2141-2142. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2141-2142.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2141-2142.

[0255] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3955-4055. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3955-4055. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3955-4055. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3955-3956. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3955-3956. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 3955-3956.

[0256] In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 320-420. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 320-420. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 320-321. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 320-321.

[0257] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 320-420, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 320-420. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0258] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0259] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0260] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5490-5494. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5490-5494. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5490-5494. The tracrRNA may comprise any of SEQ ID NOs: 5490-5494.

[0261] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5465. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5465. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5465.

[0262] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0263] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0264] Where the target nucleic acid locus may be within a cell, the enzyme may be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2141-2241. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to any one of SEQ ID NOS: 5576-5577, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 5576-5577. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0265] MG3 enzyme

[0266] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2242-2251. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2242-2251. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2242-2251. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2242-2244. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2242-2244.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2242-2244.

[0267] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4056-4066. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4056-4066. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4056-4066. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4056-4058. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4056-4058. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4056-4058.

[0268] In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 421-431. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 421-431. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 421-423. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 421-423.

[0269] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 421-431, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 421-431. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0270] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0271] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0272] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5495-5502. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5495-5502. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of any one of SEQ ID NOs: 5495-5502. The tracrRNA may comprise any of SEQ ID NOs: 5495-5502.

[0273] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to any one of SEQ ID NOs: 5466-5467. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5466-5467. The sgRNA may comprise a sequence substantially identical to any one of SEQ ID NOs: 5466-5467.

[0274] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0275] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0276] Where the target nucleic acid locus may be within a cell, the enzyme may be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2242-2251. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to any one of SEQ ID NOs: 5578-5580, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5578-5580. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0277] MG4 enzyme

[0278] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2253-2481. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2253-2481. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2253-2481. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2253 to 2481. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2253 to 2481.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2253-2481.

[0279] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4067-4295. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4067-4295. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4067-4295. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4067-4295. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4067-4295. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4067-4295.

[0280] In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 432-660. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 432-660. In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 432-660. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 432-660.

[0281] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 432-660, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 432-660. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0282] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0283] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0284] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5503. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5503. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5503. The tracrRNA may comprise SEQ ID NO: 5503.

[0285] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5468. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5468. The sgRNA may comprise a sequence substantially identical to SEQ ID NO:5468.

[0286] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0287] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0288] Where the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2253-2481. Optionally, the nucleic acid comprises a promoter operably linked to the open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism can be a eukaryote. In some cases, the organism can be a fungus. In some cases, the organism can be a human.

[0289] MG6 enzyme

[0290] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2482-2489. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2482-2489. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2482-2489.

[0291] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4296-4303. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4296-4303. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4056-4066.

[0292] In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 661-668. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 661-668.

[0293] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 661-668, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 661-668. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0294] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0295] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0296] In some cases, the tracr sequence can have a specific sequence, such as at least about 80% of at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence.

[0297] In some cases, the system may include two different guide RNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0298] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0299] Where the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2482-2489. Optionally, the nucleic acid comprises a promoter operably linked to the open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism can be a eukaryote. In some cases, the organism can be a fungus. In some cases, the organism can be a human.

[0300] MG7 enzyme

[0301] In one aspect, the disclosure provides an engineered nuclease system, including: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2490-2498. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2490-2498. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2490-2498. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2490-2498. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2490-2498.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2490-2498.

[0302] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4304-4312. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4304-4312. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4304-4312. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4304-4312. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4304-4312. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4304-4312.

[0303] In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 669-677. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 669-677. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 669-677. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 669-677.

[0304] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOs: 669-677, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 669-677. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0305] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0306] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0307] In some cases, the tracr sequence can have a specific sequence. The tracr sequence can have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracr RNA sequence. The tracr sequence can have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5504. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5504. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5504. The tracrRNA may comprise SEQ ID NO: 5504.

[0308] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0309] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0310] Where the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2490-2498. Optionally, the nucleic acid comprises a promoter operably linked to the open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism can be a eukaryote. In some cases, the organism can be a fungus. In some cases, the organism can be a human.

[0311] MG14 enzyme

[0312] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2499-2750. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2499-2750. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2499-2750. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2499-2750. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2499-2750.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2499-2750.

[0313] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4313-4564. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4313-4564. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4313-4564. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4313-4564. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4067-4295. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4313-4564.

[0314] In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 678-929. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 678-929. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 678-929. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 678-929.

[0315] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 678-929, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 678-929. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0316] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0317] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0318] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5505. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5505. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5505. The tracrRNA may comprise SEQ ID NO:5505.

[0319] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5469. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5469. The sgRNA may comprise a sequence substantially identical to SEQ ID NO:5469.

[0320] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0321] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0322] Where the target nucleic acid locus can be within a cell, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2499-2750. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to SEQ ID NO: 5581, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5581. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0323] MG15 enzyme

[0324] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2751-2913. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2751-2913. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2751-2913. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2751-2913. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2751-2913.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2751-2913.

[0325] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4565-4727. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4565-4727. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4565-4727. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4565-4727. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4565-4727. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4565-4727.

[0326] In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 930-1092. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 930-1092. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 930-1092. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 930-1092.

[0327] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOs: 930-1092, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 930-1092. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0328] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0329] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0330] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5506. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5506. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5506. The tracrRNA may comprise SEQ ID NO: 5506.

[0331] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5470. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5470. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5470.

[0332] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0333] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0334] Where the target nucleic acid locus can be within a cell, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2751-2913. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to SEQ ID NO:5582, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:5582. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0335] MG16 enzyme

[0336] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2914-3174. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2914-3174. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2914-3174. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2914-3174. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2914-3174.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 2914-3174.

[0337] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4728-4988. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4728-4988. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4728-4988. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4728-4988. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4728-4988. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4728-4988.

[0338] In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1093-1353. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1093-1353. In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1093-1353. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1093-1353.

[0339] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOs: 1093-1353, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1093-1353. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0340] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0341] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0342] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracr RNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5507. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5507. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5507. The tracrRNA may comprise SEQ ID NO: 5507.

[0343] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5471. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5471. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5471.

[0344] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0345] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0346] Where the target nucleic acid locus may be within a cell, the enzyme may be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2914-3174. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to SEQ ID NO: 5583, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5583. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0347] MG18 enzyme

[0348] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3175-3300. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3175-3300. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3175-3300. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3175-3300. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3175-3300.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 3175-3300.

[0349] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4989-5146. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4989-5146. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4989-5146. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4989-5146. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4989-5146. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 4989-5146.

[0350] In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1354-1511. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1354-1511. In some cases, the endonuclease may comprise a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1354-1511. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1354-1511.

[0351] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 1354-1511, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 1354-1511. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0352] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0353] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0354] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5508. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5508. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5508. The tracrRNA may comprise SEQ ID NO: 5508.

[0355] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5472. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5472. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5472.

[0356] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0357] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0358] Where the target nucleic acid locus can be within a cell, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 3175-3300. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to SEQ ID NO: 5584, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5584. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0359] MG21 enzyme

[0360] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3331-3474. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3331-3474. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3331-3474. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3331-3474. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3331-3474.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 3331-3474.

[0361] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 5147-5290. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5147-5290. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 5147-5290. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 5147-5290. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5147-5290. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 5147-5290.

[0362] In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1512-1655. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 1512-1655. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1512-1655. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 1512-1655.

[0363] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of any one of SEQ ID NOS: 1512-1655, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOS: 1512-1655. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0364] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0365] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0366] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5509. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5509. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5509. The tracrRNA may comprise SEQ ID NO: 5509.

[0367] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5473. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5473. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5473.

[0368] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0369] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0370] Where the target nucleic acid locus may be intracellular, the enzyme may be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 3331-3474. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to SEQ ID NO: 5585, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5585. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0371] MG22 enzyme

[0372] In one aspect, the disclosure provides an engineered nuclease system, including: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3475-3568. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3475-3568. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3475-3568. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3475-3568. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3475-3568.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 3475-3568.

[0373] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 5291-5389. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5291-5389. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 5291-5389. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 5291-5389. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5291-5389. The endonuclease can comprise an HNH domain substantially identical to any one of SEQ ID NOs: 5291-5389.

[0374] In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1656-1755. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 1656-1755. In some cases, the endonuclease can include a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1656-1755. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 1656-1755.

[0375] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located near the N- or C-terminus of the endonuclease. The NLS may be added to the N- or C-terminus of a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 432-660 or any one of SEQ ID NOs: 1656-1755. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may comprise any of the sequences in Table 1, or a combination thereof.

[0376] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Sequence identity may be determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and with the BLOSUM62 scoring matrix setting gap cost at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.

[0377] In some cases, the system may include (b) at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, the majority of nucleotides 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA may comprise, from 5' to 3', a target sequence in a cell and a non-native guide nucleic acid sequence capable of hybridizing to a tracr sequence. In some cases, the non-native guide nucleic acid sequence and the tracr sequence are covalently linked.

[0378] In some cases, the tracr sequence may have a specific sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of a naturally occurring tracrRNA sequence. The tracr sequence may have at least about 80% sequence identity over at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO:5510. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5510. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5510. The tracrRNA may comprise SEQ ID NO: 5510.

[0379] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of complexing with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5474. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5474. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5474.

[0380] In some cases, the system may include two different sgRNAs targeting a first region and a second region for cleavage in the target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template including, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides for a first homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 5' to the first region, and a second homology arm comprising a sequence of at least about 20 nucleotides (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) 3' to the second region.

[0381] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method may include delivering any of the non-natural systems disclosed herein, including an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA), to the target nucleic acid locus. The enzyme may form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it may modify the target nucleic acid locus of interest. Delivering the enzyme to the locus may include transfecting a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include electroporating a cell with the system or a nucleic acid encoding the system. Delivering a nuclease to the locus may include incubating the system with a nucleic acid comprising the locus of interest in a buffer. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus may comprise genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus may be in a cell. The target nucleic acid locus may be in vitro. The target nucleic acid locus can be in a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single- or double-stranded break at or adjacent to the target locus of interest.

[0382] Where the target nucleic acid locus can be intracellular, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having a RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 3475-3568. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can comprise a sequence substantially identical to SEQ ID NO: 5586, or a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5586. Optionally, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease may be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease may be provided as a translated polypeptide. At least one engineered sgRNA may be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the organism may be a eukaryote. In some cases, the organism may be a fungus. In some cases, the organism may be a human.

[0383] MG23 enzyme

[0384] In one aspect, the disclosure provides an engineered nuclease system, including: (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease can include a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3569-3637. In some cases, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3569-3637. In some cases, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3569-3637. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3569-3637. In some cases, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3569-3637.In some cases, the endonuclease may comprise a RuvC_III domain substantially identical to any one of SEQ ID NOs: 3569-3637.

[0385] The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 5390-5460. In some cases, the endonuclease may comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5390-5460. The endonuclease may comprise an HNH domain substantially identical to any one of SEQ ID NOs: 5390-5460. The endonuclease may comprise an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 5390-5460. In some cases, the endonuclease can comprise an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5390-5460. The endonuclease can comprise an HNH domain substantia...

Claims

1. 1. A method for editing a TRAC locus in a cell, comprising administering to the cell: (a) a ribonucleic acid-guided endonuclease or a nucleic acid encoding said ribonucleic acid-guided endonuclease, wherein said ribonucleic acid-guided endonuclease comprises a sequence having at least 90% sequence identity to SEQ ID NO: 421; and (b) an engineered guide ribonucleic acid or a nucleic acid encoding said engineered guide ribonucleic acid, wherein said engineered guide ribonucleic acid is configured to form a complex with said ribonucleic acid guide endonuclease, and wherein said engineered guide ribonucleic acid comprises a spacer sequence configured to hybridize to a region of said TRAC locus. contacting the wherein the spacer sequence comprises a sequence having at least 90% sequence identity to at least 18 consecutive nucleotides of SEQ ID NO: 5950; method.

2. 2. The method of claim 1, wherein the ribonucleic acid guided endonuclease is a Class II Type II Cas endonuclease.

3. 3. The method of claim 1 or 2, wherein the ribonucleic acid-guided endonuclease comprises a sequence having at least 95% sequence identity to SEQ ID NO:

421.

4. 4. The method of claim 3, wherein the ribonucleic acid guide endonuclease comprises the sequence of SEQ ID NO:

421.

5. 3. The method of claim 1 or 2, wherein the ribonucleic acid-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 90% sequence identity to SEQ ID NO: 2242 or SEQ ID NO: 2244.

6. 6. The method of claim 5, wherein the RuvCIII domain comprises the sequence of SEQ ID NO: 2242 or SEQ ID NO: 2244.

7. 7. The method of claim 6, wherein the RuvCIII domain comprises the sequence of SEQ ID NO: 2242.

8. The method of any one of claims 1 to 3, wherein the ribonucleic acid-guided endonuclease comprises an HNH domain.

9. 9. The method of claim 8, wherein the HNH domain comprises a sequence having at least 90% sequence identity to SEQ ID NO: 4056.

10. 10. The method of claim 9, wherein the HNH domain comprises the sequence of SEQ ID NO: 4056.

11. 11. The method of any one of claims 1 to 10, wherein the engineered guide ribonucleic acid comprises a sequence configured to bind to the ribonucleic acid guide endonuclease.

12. 12. The method of claim 11, wherein the sequence configured to bind to the ribonucleic acid guide endonuclease comprises a sequence having at least 90% sequence identity to SEQ ID NO: 5466.

13. 13. The method of claim 12, wherein the sequence configured to bind to the ribonucleic acid guide endonuclease comprises the sequence of SEQ ID NO: 5466.

14. The method of any one of claims 1 to 13, wherein the cells are human cells.

15. 15. The method of any one of claims 1 to 14, wherein the cells are peripheral blood mononuclear cells (PBMC), T cells, NK cells, hematopoietic stem cells (HSCT), or B cells, or a combination thereof.

16. 16. The method of any one of claims 1 to 15, wherein said contacting further comprises transfecting said cell with said nucleic acid encoding said ribonucleic acid guide endonuclease and said nucleic acid encoding said engineered guide ribonucleic acid.

17. 17. The method of any one of claims 1 to 16, wherein the spacer sequence comprises a sequence having at least 95% sequence identity to SEQ ID NO: 5950.

18. 18. The method of any one of claims 1 to 17, wherein the spacer sequence comprises a sequence having at least 18 consecutive nucleotides of SEQ ID NO: 5950.

19. 19. The method of claim 18, wherein the spacer sequence comprises the sequence of SEQ ID NO: 5950.

20. The method of any one of claims 1 to 19, wherein the cells are PBMCs or T cells.

21. 21. The method of any one of claims 1 to 20, wherein the ribonucleic acid guide endonuclease comprises the sequence of SEQ ID NO: 421 and the spacer sequence comprises the sequence of SEQ ID NO: 5950.

Citation Information

Patent Citations

  • CRISPR-CPF1-Related Methods, Compositions, and Components for Cancer Immunotherapy

    JP2019507599A

  • Novel CAS9 orthologs

    US20190264232A1

  • Materials and methods for engineering cells and uses thereof in immuno-oncology

    WO2019097305A2

  • Modification of immune-related genomic LOCI using paired crispr nickase ribonucleoproteins

    WO2019200306A1

  • Compositions and methods for immunotherapy

    WO2020081613A1