Class II Type II CRISPR system

By developing a small-sized SMART nuclease system and combining it with an engineered guide RNA structure, the challenges of delivering large-sized class 2 Cas effectors have been overcome, enabling efficient gene editing and therapeutic applications.

CN116096877BActive Publication Date: 2026-04-28METAGENOMICS THERAPEUTICS CO
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
METAGENOMICS THERAPEUTICS CO
Filing Date
2021-03-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The large size of the two existing classes of Cas effectors presents challenges for delivery in therapeutic applications, making it difficult to effectively utilize their potential in gene editing and therapy.

Method used

The SMART (SMall ARchaeal-associated) nuclease system was developed, which includes a small-sized endonuclease and an engineered guide ribonucleic acid structure defined by the RuvC and HNH catalytic domains, and is suitable for specific binding and cleavage of target deoxyribonucleic acid sequences.

Benefits of technology

A small molecular weight nuclease system has been developed, which can efficiently bind to and cleave target sequences, making it suitable for gene editing and therapeutic applications and overcoming the challenge of delivering large-sized nucleases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116096877B_ABST
    Figure CN116096877B_ABST
Patent Text Reader

Abstract

The present disclosure provides endonucleases and methods of using such enzymes or variants thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims the benefits of U.S. Provisional Application No. 63 / 116,149, filed November 19, 2020, entitled “CLASS II, TYPE II CRISPR SYSTEMS”, and U.S. Provisional Application No. 63 / 003,159, filed March 31, 2020, entitled “CLASS II, TYPE II CRISPR SYSTEMS”, both of which are incorporated herein by reference in their entirety.

[0003] sequence list

[0004] This application contains a sequence list submitted electronically in ASCII format and hereby incorporated in its entirety by reference. The ASCII copy created on March 27, 2021, is named 55921-711_601_SL.txt and is 2,235,526 bytes in size. Background Technology

[0005] Cas enzymes and their associated clustered regularly interspaced short palindromic repeats (CRISPR)-guided ribonucleic acid (RNA) appear to be ubiquitous components of the prokaryotic immune system (approximately 45% of bacteria and 84% of archaea), protecting these microorganisms from non-self nucleic acids such as infectious viruses and plasmids through CRISPR-RNA-guided nucleic acid cleavage. While the deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins exhibit high diversity, containing a variety of nucleic acid interaction domains. Although CRISPR DNA elements were observed as early as 1987, the programmable endonuclease cleavage capability of CRISPR / Cas complexes has only recently been recognized, enabling the application of recombinant CRISPR / Cas systems in various DNA manipulation and gene editing applications. Due to the uses of these enzymes, they are being repurposed for a wide range of biotechnological, gene editing, and therapeutic applications. Due to their single-effects architecture, most of these systems are now being reused for genome engineering belonging to CRISPR Class II and Class V categories. Summary of the Invention

[0006] The large size (greater than approximately 1200 amino acids) of many class 2 Cas effectors makes their delivery for therapeutic applications challenging. Therefore, this paper describes methods, compositions, and systems relating to novel putatively directed dsDNA nucleases known as the SMART (SMall ARchaeal-associated) nuclease system. These endonuclease effectors are defined by their small size (400-1050 aa), the presence of RuvC and HNH catalytic domains, and other predicted protein features that collectively suggest novel biochemical mechanisms.

[0007] In some aspects, this disclosure provides an engineered nuclease system comprising: (a) a nuclease comprising a RuvC domain and an HNH domain, wherein the nuclease is derived from an uncultured microorganism; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the nuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize with a target deoxyribonucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to the nuclease; wherein the nuclease has a molecular weight of about 96 kDa or less. In some embodiments, the nuclease is an archaeal nuclease. In some embodiments, the nuclease is a type II Cas nuclease. In some embodiments, the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity with any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises an arginine-rich region containing an RRxRR motif or a domain homologous to PF14239. In some embodiments, the arginine-rich region or the domain homologous to PF14239 has at least 85%, at least 90%, or at least 95% identity with the arginine-rich region or the domain homologous to PF14239 of any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the REC domain has at least 85%, at least 90%, or at least 95% identity with the REC domain of any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises a BH (bridge helix) domain, a WED (wedge) domain, and a PI (PAM interaction) domain. In some embodiments, the BH domain, the WED domain, or the PI domain has at least 85%, at least 90%, or at least 95% identity with the BH domain, WED domain, and / or PI domain of any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668.

[0008] In some aspects, this disclosure provides an engineered nuclease system comprising: (a) a nuclease comprising a RuvC-I domain and an HNH domain; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the nuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize with a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to the nuclease, wherein the nuclease comprises a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity with any one of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the nuclease is an archaeal nuclease. In some embodiments, the nuclease is a type II Cas nuclease. In some embodiments, the endonuclease further comprises an arginine-rich region containing an RRxRR motif or a domain homologous to PF14239. In some embodiments, the arginine-rich region or the domain homologous to PF14239 has at least 85%, at least 90%, or at least 95% identity with the arginine-rich region of any one of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the REC domain has at least 85%, at least 90%, or at least 95% identity with the REC domain of any one of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain. In some embodiments, the BH domain, the WED domain, or the PI domain has at least 85%, at least 90%, or at least 95% identity with the BH domain, WED domain, and / or PI domain of any one of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease is derived from uncultured microorganisms. In some embodiments, the ribonuclease-configured ribonuclease-binding sequence comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 199-200, 460-461, or 669-673, or a sequence having at least 80% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NO: 201-203 or 613-616. In some embodiments, the guide nucleic acid structure comprises a sequence having at least 80% identity with a non-degenerate nucleotide of any one of SEQ ID NO:201-203, 613-616.

[0009] In some aspects, this disclosure provides an engineered nuclease system comprising: (a) an engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize with a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to a nuclease, wherein the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO:199-200, 460-461, or 669-673, or a sequence having at least 80% sequence identity with a non-variable nucleotide of any one of SEQ ID NO:201-203 or 613-616; and (b) an RNA-directed nuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the RNA-directed nuclease is an archaeal nuclease. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide comprising the guide ribonucleic acid sequence and the tracr ribonucleic acid sequence. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaea, eukaryotic, fungal, plant, mammalian, or human genome sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) located near the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises sequences selected from SEQ ID NO:205-220. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template, the single-stranded or double-stranded DNA repair template comprising, from 5' to 3': a first homologous arm comprising a sequence having at least 20 nucleotides at the 5' of the target deoxyribonucleic acid sequence; a synthetic DNA sequence having at least 10 nucleotides; and a second homologous arm comprising a sequence having at least 20 nucleotides at the 3' of the target sequence. In some embodiments, the first or second homologous arm comprises a sequence having at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the system further comprises Mg 2+Source. In some embodiments, the endonuclease and the tracr ribonucleic acid sequence are derived from different bacterial species within the same phylum. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity with any one of SEQ ID NO:2-24, and the guide RNA structure comprises an RNA sequence predicted to contain hairpins, the hairpins comprising a stem and a loop, wherein the stem comprises at least 12 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises a second stem and a second loop, wherein the second stem comprises at least 5 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises an RNA structure containing at least two hairpins. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity with SEQ ID NO:1, and the guide RNA structure comprises an RNA sequence predicted to contain at least four hairpins, the hairpins comprising a stem and a loop. In some embodiments: a) the endonuclease comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO: 1, 2, 10, 17, or 613-616; and b) the guide RNA structure comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO: 199-200, or 669-673, or any one of SEQ ID NO: 201-203, or 613-616. In some embodiments: a) the endonuclease comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO: 1-24, 462-488, or 501-612; and b) the guide RNA structure comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO: 199-200, or 669-673, or any one of SEQ ID NO: 201-203, or 613-616. In some embodiments: a) the endonuclease comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO: 2, 10, or 17; and b) the guide RNA structure comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO: 202-203, or 613-614. In some embodiments: a) the endonuclease comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO:25-198, 221-459, or 489-580; and b) the guide RNA structure comprises at least 70%, at least 80%, or at least 90% of the sequence identical to a type II sgRNA or tracr sequence.In some embodiments, sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using Smith-Waterman homology search algorithm parameters. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm, which uses the following parameters: word length (W) of 3, expected value (E) of 10, and a BLOSUM62 scoring matrix with a gap cost of 11, an extension value of 1, and conditional composition of the scoring matrix adjustment. In some embodiments, the endonuclease is not a Cas9, Cas14, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, or Cas13d endonuclease. In some implementations, the endonuclease has less than 80% identity with the Cas9 endonuclease.

[0010] In some aspects, this disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: a) a DNA targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and b) a protein-binding segment comprising two complementary nucleotide segments that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary nucleotide segments are covalently linked to each other by intercalation nucleotides, and wherein the engineered guide ribonucleic acid polynucleotide is configured to form a complex with a nuclease comprising a variant having at least 75% sequence identity with any one of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the DNA targeting segment is located at the 5' position of both of the two complementary nucleotide segments. In some embodiments: a) the protein-binding segment comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO:199-200 or 669-673; b) the protein-binding segment comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO:201-203 or 613-616. In some embodiments: a) the endonuclease comprises at least 70%, at least 80%, or at least 90% of the sequence identical to any one of SEQ ID NO:2, 10, or 17; and b) the guide RNA structure comprises at least 70%, at least 80%, or at least 90% of the sequence identical to at least one of the sequence identical to any one of the sequence identical to any one of SEQ ID NO:200 or SEQ ID NO:202-203 or 613-614. In some embodiments: a) the endonuclease comprises at least 70%, at least 80%, or at least 90% identical sequence to any one of SEQ ID NO: 25-198, 221-459, or 489-580; and b) the guide RNA structure comprises at least 70%, at least 80%, or at least 90% identical sequence to type II sgRNA. In some embodiments, the endonuclease further comprises a base editor or histone editor coupled to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.

[0011] In some respects, this disclosure provides a deoxyribonucleic acid polynucleotide that encodes any one of the engineered guide ribonucleic acid polynucleotides described herein.

[0012] In some aspects, this disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a type II Cas endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, and wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, 60 kDa or less, or 30 kDa or less. In some embodiments, the endonuclease comprises SEQ ID NO: 1-198, 221-459, 463-612, or 617-668, or variants thereof having at least 70% sequence identity with them. In some embodiments, the endonuclease further comprises a sequence encoding one or more nuclear localization sequences (NLS) adjacent to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NO: 205-220. In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human. In some embodiments, the organism is a prokaryote or bacterium, and the organism is different from the organism from which the endonuclease originates. In some embodiments, the organism is not the uncultured microorganism.

[0013] In some aspects, this disclosure provides a vector comprising a nucleic acid sequence encoding an RNA-directed endonuclease comprising a RuvC-I domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism and wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less, and wherein the RNA-directed endonuclease is optionally an archaea. In some embodiments, the endonuclease further comprises an arginine-rich region containing an RRxRR motif or a domain having PF14239 homology. In some embodiments, the endonuclease further comprises a REC (recognition) domain. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain.

[0014] In some aspects, this disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: a) a guide ribonucleic acid sequence configured to hybridize with a target deoxyribonucleic acid sequence; and b) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, microcircle, CELiD, adeno-associated virus (AAV)-derived viral particle, or lentivirus.

[0015] In some aspects, this disclosure provides a cell comprising any of the vectors described herein. In some embodiments, the cell is a bacterial, archaea, fungus, eukaryotic, mammalian, or plant cell. In some embodiments, the cell is a bacterial cell.

[0016] In some respects, this disclosure provides a method for producing a nucleic acid endonuclease, the method comprising culturing any of the cells described herein.

[0017] In some aspects, this disclosure provides a method for binding, cleaving, labeling, or modifying a double-stranded deoxyribonucleic acid (DDNA) polynucleotide, the method comprising: (a) contacting the DDNA polynucleotide with a type II Cas endonuclease in the form of a complex of an engineered guide ribonuclease structure configured to bind to the endonuclease and the DDNA polynucleotide; (b) wherein the DDNA polynucleotide comprises a protospacer adjacent motif (PAM); wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the endonuclease cleaves the DDNA polynucleotide, wherein the PAM comprises NGG. In some embodiments, the endonuclease cleaves 6-8 or 7 nucleotides derived from the PAM in the DDNA polynucleotide. In some embodiments, the endonuclease comprises a variant having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity with any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668.

[0018] In some aspects, this disclosure provides a method for binding, cleaving, labeling, or modifying a double-stranded deoxyribonucleic acid (DDNA) polynucleotide, the method comprising: (a) contacting the DDNA polynucleotide with an RNA-guided archaeal endonuclease in the form of a complex of an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the DDNA polynucleotide; wherein the DDNA polynucleotide comprises a protospacer adjacent motif (PAM); and wherein the endonuclease comprises a variant having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease cleaves the DDNA polynucleotide, wherein the PAM comprises NGG. In some embodiments, the endonuclease cleaves 6-8 or 7 nucleotides from the PAM in the double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the type II Cas endonuclease is not a Cas9, Cas14, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, or Cas13d endonuclease. In some embodiments, the type II Cas endonuclease is derived from uncultured microorganisms. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a double-stranded deoxyribonucleic acid polynucleotide from a prokaryote, archaea, or bacterium other than the species from which the endonuclease originates.

[0019] In some aspects, this disclosure provides a method for modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus any engineered nuclease system described herein, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, and wherein the complex is configured such that when the complex binds to the target nucleic acid locus, the complex modifies the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus includes binding, cleaving, splitting, or labeling the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic eukaryotic DNA, archaea DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid comprises bacterial DNA, wherein the bacterial DNA is derived from a bacterial or archaea species different from the species from which the endonuclease originates. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is intracellular. In some embodiments, the endonuclease and the engineered guide nucleic acid structure are encoded by separate nucleic acid molecules. In some embodiments, the cell is a prokaryotic cell, bacterial cell, archaea cell, eukaryotic cell, fungal cell, plant cell, animal cell, mammalian cell, rodent cell, primate cell, or human cell. In some embodiments, the cell is derived from a species different from the species from which the endonuclease originates. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus includes delivering any nucleic acid or any vector described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a nucleic acid containing an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid contains a promoter operatively linked to the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a translated polypeptide. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus includes delivering deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid structure, the engineered guide ribonucleic acid structure being operatively linked to the ribonucleic acid (RNA) polIII promoter. In some embodiments, the endonuclease induces single-strand or double-strand breaks at or near the target locus. In some embodiments, the endonuclease induces a double-strand break near the target locus at the 5' position of the protospacer adjacent motif (PAM).In some embodiments, the endonuclease induces a double-strand break at 6-8 or 7 nucleotides at the 5' of the PAM. In some embodiments, the engineered nuclease system induces chemical modification of nucleotide bases within or near the target locus, or induces chemical modification of histones within or near the target locus. In some embodiments, the chemical modification is the deamination of adenosine or cytosine nucleotides. In some embodiments, the endonuclease further comprises a base editor coupled to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase includes ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase includes APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.

[0020] Further aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of this disclosure are shown and described. As will be understood, this disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious respects, all without departing from this disclosure. Therefore, the drawings and description are to be regarded in an illustrative rather than restrictive manner.

[0021] Incorporation

[0022] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent or patent application is explicitly and individually incorporated by reference. Attached Figure Description

[0023] The novel features of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description and accompanying drawings (also referred to herein as “Figure” and “FIG”) illustrating illustrative embodiments utilizing the principles of the invention, in which:

[0024] Figure 1A phylogenetic tree was constructed to illustrate the homology relationships of different classes and types of CRISPR / Cas loci. The SMART I and IICas enzyme classes described herein are shown relative to the two types II-A, II-B, and II-C Cas systems, demonstrating that these systems are classified into categories distinct from II-A, II-B, and II-C. (A) shows a SMART phylogenetic tree against the background of a Cas9 reference sequence, where SMART effectors cluster away from the Cas9 reference sequence (types II-A, II-B, and II-C); (B) shows a SMART phylogenetic tree illustrating the subgroups of SMART enzymes.

[0025] Figure 2 The length distribution of the SMART effectors described herein is shown, revealing that SMART I and II enzymes cluster at lower molecular weights than Cas9-like enzymes. SMART nucleases exhibit a bimodal distribution, with one peak around 400 aa (SMART II) and a second peak around 750 aa (SMART II). Cas9 nucleases also show a bimodal distribution, with peaks around 1,100 aa (e.g., SaCas9) and around 1,300 aa (e.g., SpCas9).

[0026] Figure 3 The genomic background of the 'small' type II nucleases MG33-1 and MG35-236 is depicted. SMART nucleases and CRISPR accessory proteins are shown with dark gray arrows, while other genes are depicted with light gray arrows. The domains of all genes in the predicted genomic fragment are shown with gray boxes under the arrows. The following are shown: (A) Genomic background of the SMART I MG33-1 nuclease and the CRISPR locus upstream of the SMART II nuclease MG35-236, with the predicted insertion sequence carrying transposases TnpA and TnpB shown downstream of SMART II; (B) Genomic background of the SMART I nuclease MG34-1, showing alignment of environmental expression sequencing reads with the CRISPR array and the predicted tracrRNA, and transcriptome coverage of this region shown on contig sequences; (C) Genomic background of the SMART I nuclease MG34-16, showing alignment of environmental expression sequencing reads with the CRISPR array and the predicted tracrRNA, and transcriptome coverage of this region shown on contig sequences; and (D) MG34-16 from (D). The CRISPR array targets a genomic fragment by space 7, where phage-derived genomic fragments are identified based on virus-specific gene annotation termination enzymes and entry enzymes; the inset shows the location of space 7 for MG34-16, which targets the C-terminus of a viral gene with unknown function—the putative NGG PAM of MG34-16 is highlighted in gray boxes downstream of the spacer match.

[0027] Figure 4 Exemplary SMART endonucleases (MG33-1 (SEQ ID NO:1), MG33-2 (SEQ ID NO:463), MG33-3 (SEQ ID NO:464), MG34-1 (SEQ ID NO:2), MG34-9 (SEQ ID NO:10), MG34-16 (SEQ ID NO:17), MG102-1 (SEQ ID NO:581), MG102-2 (SEQ ID NO:582), MG35-1 (SEQ ID NO:25), MG35-2 (SEQ ID NO:26), MG35-3 (SEQ ID NO:27), MG35-102 (SEQ ID NO:126), MG35-236 (SEQ ID NO:284), MG35-419 (SEQ ID NO:222), MG35-420 (SEQ ID NO:223), and MG35-421 (SEQ ID NO:1) are shown. Multiple sequence alignments of NO:224) are shown, where the SaCas9 sequence used as the reference domain is shown as a rectangle below the reference sequence, and the catalytic residues are shown as squares above each sequence. The alignments are shown as follows: (A) alignment of the endonuclease region containing the RuvC-I and bridging helical domains; (B) alignment of the region containing the RuvC-III domain; and (C) alignment of the region containing the RuvC-II and HNH domains.

[0028] Figure 5 An exemplary domain organization of the SMART I nuclease is depicted using MG34-1 as an example. (A) A diagram showing the predicted domain architecture of the SMART I nuclease, which comprises: three RuvC domains, a bridging helix (“BH”), a domain homologous to Pfam PF14239 (including the interrupted recognition domain (“REC”)), an HNH endonuclease domain (“HNH”), a wedge-shaped domain (“WED”), and a PAM interaction domain (PI); and (B) A multiple sequence alignment profile of two SMART I nucleases relative to a reference Cas9 nuclease sequence, where RuvC and HNH catalytic residues are indicated by black bars on each sequence, regions aligned to the SaCas crystal structure in 3D space are indicated by circular boxes, and dashed lines indicate poorly aligned or unaligned regions in 3D space between the SMART and SaCas9 3D structure predictions.

[0029] Figure 6An exemplary domain organization of SMART II endonucleases is depicted using MG35 family enzymes (MG35-3, MG35-4) as examples. (A) A diagram showing the predicted domain architecture of a SMART II nuclease composed of: three RuvC domains, a domain homologous to Pfam PF14239, an HNH endonuclease domain, an unknown domain, and a recognition domain (REC); and (B) a multiple sequence alignment overview of two SMART II nucleases relative to a reference Cas9 nuclease sequence, where RuvC and HNH catalytic residues are indicated by black bars on each sequence, regions aligned to the SaCas crystal structure in 3D space are indicated by circular boxes, and residues identified from the 3D structure prediction that may be involved in the recognition guide / target / PAM sequence are indicated by dark gray boxes above the MG35-419 sequence (within the RRXRR and REC domains).

[0030] Figure 7 Various characteristics of SMART enzymes are shown. A dot plot (A) showing the identity of the SMART I domains of the various enzymes described herein compared to spCas9 is shown, indicating a maximum sequence identity of approximately 35%; (B) a dot plot of the lengths of individual SMART I domains of the enzymes described herein.

[0031] Figure 8 The count distribution of various SMART-specific motifs relative to predicted motifs in Cas9 nuclease sequences is shown, indicating that these motifs are more common in SMART enzymes; motifs were predicted on 803 reference Cas9 sequences (types II-A, II-B, and II-C), 84 SMART sequences, and 471 SMART II sequences. (A) shows the Zn-binding band motif (CX) in various types of class 2 Cas enzymes. [2-4] C and CX [2-4] (H) is a box plot of the counting frequency; and (B) is a histogram of the counting frequency of the RRXRR motif in various types of Cas enzymes. The average count values ​​are tracked on the lines in (A) and (B), while outliers are represented by points.

[0032] Figure 9 The predicted guide RNA structures for the designed single guide RNA (sgRNA) against the cleavage activity of the SMARTI endonuclease are shown. (A) MG34-1 sgRNA 1; (B) MG34-1 sgRNA 2; (C) MG34-9 sgRNA 1; and (D) MG34-16 sgRNA 1 are shown.

[0033] Figure 10The lysis characterization of the SMART I nuclease as described in Example 1 is depicted. (A) An Agilent TapeStation gel shows the ligation products of MG34-1 designed with two sgRNAs compared to the negative control. Lane L3: Ladder diagram. Lane A4: Apo, no sgRNA. Lanes B4 and C4: MG34-1 sgRNAs tested (sg1: SEQ ID No. 613, sg2: 614). The lysis product bands are marked with arrows. Lanes G3 and H3: Gray, not relevant to this experiment. (B) A PCR gel shows the ligation products, demonstrating the activity of MG34-1, 34-9, and 34-16. Lane 1: Ladder diagram. Lanes 2-7: sgRNA designs for MG34-1 with six spacers. Lanes 8 and 9: sgRNA designs for 34-9 and 34-16, respectively. Arrows indicate lysis confirmation bands.

[0034] Figure 11 The sequence cleavage preferences of the MG34 nuclease are shown. (A) SeqLogo representation of the common PAM sequence (NGGN) of MG34-1, including sgRNA 1 (top, SEQ ID NO: 613) and sgRNA 2 (bottom, SEQ ID NO: 614). (B) A histogram showing the cleavage site locations of MG34-1 indicates that MG34-1 prefers cleavage from approximately position 7 of the PAM. (C) A Sanger sequencing chromatogram showing the preferred NGG PAM for MG34-9 (highlighted in a box). Arrows indicate the cleavage site from position 7 of the PAM.

[0035] Figure 12 Results of plasmid targeting assays for MG34-1 in *E. coli* are shown. (A) shows the plating of *E. coli* strains, demonstrating plasmid cleavage; *E. coli* expressing MG34-1 and sgRNA were transformed with a kanamycin-resistant plasmid containing the sgRNA target (+sp). The quadrants showing growth defects (+sp) compared to the negative control (no target and PAM (-sp)) indicate successful enzyme targeting and cleavage. The experiments were repeated twice and performed in triplicate. (B) shows a graph of colony-forming units (CFU) measurements in the plating assays in A, showing growth inhibition under the target condition (+sp) compared to the non-target control (-sp), indicating plasmid cleavage.

[0036] Figure 13An exemplary genomic background of the SMART system for MG35-419 is shown. SMART nucleases are indicated by dark gray arrows, and other genes by light gray arrows. Domains of all genes in the predicted genomic fragment are indicated by gray boxes under the arrows. Alignment of environmental expression sequencing reads is shown below the CRISPR array in (A) and upstream of the effector in (B). Transcriptome coverage of the expressed regions is shown on contig sequences. (A) shows the genomic background of the SMART II MG35-419 effector and the nearby CRISPR loci. (B) shows the genomic background of the SMART II effector MG35-3, displaying the transcribed 5' UTR.

[0037] Figure 14 The 3D structure prediction of SMART II MG35-419 is shown. This 3D model is well compared to the SaCas9 crystal structure, although it is less than half its size. The regions compared to the SaCas9 template include short regions of the catalytic leaflets (RuvC-I, HNH, and RuvC-III domains) and the recognition leaflets (REC). SMART II-specific domains include those containing the RRXRR motif and those homologous to Pfam PF14239, as well as domains with unknown functions.

[0038] Figure 15 Results of preliminary cleavage assays of SMART II effectors were depicted. The cleavage activity of the MG35-420 (SEQ ID NO:223) protein formulation in TXTL extracts expressing the entire locus was tested. The protein formulation was incubated with a PAM library (dsDNA target, predicted repeat regions in the forward and reverse (fw and rv) loci (cr1)) and intergenic regions that may encode the desired cofactors. Lanes 2-9 (cr-free array): control experiments with no repeat regions. Apo: protein formulation only with the target PAM library. Labels 1-2.5 indicate seven distinct intergenic regions. -IG: no intergenic regions included as a control. PCR gels of the ligation products show putative cleavage bands (arrows) indicating dsDNA cleavage.

[0039] A brief description of the sequence list

[0040] The accompanying sequence listing provides exemplary polynucleotide and polypeptide sequences used in the methods, compositions, and systems according to this disclosure. Exemplary descriptions of the sequences are given below.

[0041] MG33 nuclease

[0042] SEQ ID NO:1 and 463-486 show the full-length peptide sequences of the MG33 nuclease.

[0043] SEQ ID NO:199 and 669-670 show the nucleotide sequences of tracrRNAs predicted to act in conjunction with the MG33 nuclease.

[0044] SEQ ID NO:201 shows the nucleotide sequence of the predicted single guide RNA (sgRNA) sequence that is predicted to act in conjunction with the MG33 nuclease. “N” indicates a variable residue and non-N residues indicate a scaffold sequence.

[0045] MG34 nuclease

[0046] SEQ ID NO:2-24 and 487-488 show the full-length peptide sequences of the MG34 nuclease.

[0047] SEQ ID NO:200 shows the nucleotide sequence of the tracrRNA that is predicted to act in conjunction with the MG34 nuclease.

[0048] SEQ ID NO: 202, 203 and 613-616 show the nucleotide sequences of the predicted single guide RNA (sgRNA) sequences that are predicted to act in conjunction with the MG34 nuclease. “N” indicates a variable residue and non-N residues indicate a scaffold sequence.

[0049] MG35 nuclease

[0050] SEQ ID NO:25-198, 221-459, 489-580 and 617-668 show the full-length peptide sequences of the MG35 nuclease.

[0051] SEQ ID NO:460-461 shows the nucleotide sequence of MG35 tracrRNA derived from the same locus as the MG35 nuclease.

[0052] SEQ ID NO:462 shows the repeat sequence of the MG35 nuclease described herein.

[0053] MG102 nuclease

[0054] SEQ ID NO:581-612 shows the full-length peptide sequence of the MG102 nuclease.

[0055] SEQ ID NO:672-673 shows the nucleotide sequence of MG102tracrRNA derived from the same locus as the MG102 nuclease.

[0056] SEQ ID NO:205-220 illustrates an exemplary nuclear localization sequence (NLS) that can be attached to a nuclease according to the present disclosure. Detailed Implementation

[0057] Although various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives may be employed with respect to the embodiments of the invention described herein.

[0058] Unless otherwise stated, the practice of some of the methods disclosed herein utilizes techniques from immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th ed. (2012); the series Current Protocols in Molecular Biology (edited by F.M.A. Susubel et al.); the series Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (edited by M.J. MacPherson, B.D. Hames, and G.G. Taylor (1995)); Harlow and Lane (edited by Harlow and Lane (1988); Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th ed. (edited by R.R. Freshney (2010)) (which are incorporated herein by reference in their entirety).

[0059] Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein are intended to also include the plural forms. Furthermore, with regard to the terms “including,” “includes,” “having,” “has,” “with,” or variations thereof used in the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0060] The terms “about” or “approximately” mean that a specific value, as determined by a person skilled in the art, is within an acceptable range of error, which will depend in part on how the value is measured or determined, for example, the limitations of the measurement system. For example, according to practice in the art, “about” can mean within one or more standard deviations. “About” can also mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0061] As used herein, “cell” generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of a living organism. Cells can originate from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of unicellular eukaryotic organisms, protozoan cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, lycophytes, hornwort, liverwort, mosses), and algal cells (e.g., *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsisgaditana*, *Chlorella pyrenoidosa*, *Sargassum*). Cells can be derived from various organisms, including: algae (e.g., kelp), fungal cells (e.g., yeast cells, mushroom cells), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes, cells are not derived from natural organisms (e.g., cells can be synthetic, sometimes called artificial cells).

[0062] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. Nucleotides can include synthetic nucleotides. Nucleotides can include synthetic nucleotide analogs. Nucleotides can be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates (ATP), uridine triphosphates (UTP), cytosine triphosphates (CTP), guanosine triphosphates (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled, such as using portions containing optically detectable parts (e.g., fluorophores). Quantum dots can also be used for labeling. Detectable labels can include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides can include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, Cyan, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides may include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, obtained from Perkin Elmer, Foster City, Calif; FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, obtained from Amersham, Arlington. Heights, Il.; luciferin-15-dATP, luciferin-12-dUTP, tetramethyl-rhodamine-6-dUTP, TR770-9-dATP, luciferin-12-ddUTP, luciferin-12-UTP and luciferin-15-2'-dATP, obtained from Boehringer Mannheim, Indianapolis, Ind.; and chromosome-marked nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, luciferin-12-UTP, luciferin-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, obtained from Molecular Probes, Eugene, Oreg. Nucleotides can also be labeled or marked by chemical modifications. Chemically modified mononucleotides can be biotin-dNTPs.Some non-limiting examples of biotinylated dNTPs may include biotin-dATP (e.g., biotin-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP). Nucleotides may include nucleotide analogs. In some embodiments, nucleotide analogs may include native nucleotide structures modified at any position to alter certain chemical properties of the nucleotide while retaining the nucleotide analog's ability to perform its intended function (e.g., hybridization with other nucleotides in RNA or DNA). Examples of derivatizable nucleotide positions include the 5-position, such as 5-(2-amino)propyluridine, 5-bromouridine, 5-propynyluridine, 5-propenyluridine, etc.; the 6-position, such as 6-(2-amino)propyluridine; the 8-position of adenosine and / or guanosine, such as 8-bromoguanosine, 8-chloroguanosine, 8-fluoroguanosine, etc. Nucleotide analogs also include denitronucleotides, such as 7-denitro-adenosine; O- and N-modified (e.g., alkylated, such as N6-methyladenosine, or otherwise known in the art) nucleotides; and other heterocyclic modified nucleotide analogs, such as those described in Herdewijn, Antisense Nucleic Acid Drug Dev., 2000 Aug. 10(4):297-310. Nucleotide analogs may also include modifications to the sugar moiety of the nucleotide. For example, the 2'OH- group can be replaced by a group selected from H, OR, R, F, Cl, Br, I, SH, SR, NH2, NHR, NR2, COOR, or OR, wherein R is a substituted or unsubstituted C1-C6 alkyl, alkenyl, alkynyl, aryl, etc. Other possible modifications include those described in U.S. Patent Nos. 5,858,988 and 6,291,438. Examples of derivatizable nucleotide positions include the 5-position, such as 5-(2-amino)propyluridine, 5-bromouridine, 5-propynyluridine, 5-propenyluridine, etc.; the 6-position, such as 6-(2-amino)propyluridine; and the 8-position of adenosine and / or guanosine, such as 8-bromoguanosine, 8-chloroguanosine, 8-fluoroguanosine, etc. Nucleotide analogs also include denitronucleotides, such as 7-denitro-adenosine: O- and N-modified (e.g., alkylated, such as N6-methyladenosine, or otherwise known in the art) nucleotides; and other heterocyclic modified nucleotide analogs, such as those described in Herdewijn, Antisense Nucleic Acid Drug Dev., Aug. 2000. 10(4):297-310. Nucleotide analogs may also include modifications to the sugar moiety of the nucleotide.For example, the 2'OH- group can be replaced by a group selected from H, OR, R, F, Cl, Br, I, SH, SR, NH2, NHR, NR2, COOR, or OR, wherein R is a substituted or unsubstituted C1-C6 alkyl, alkenyl, alkynyl, aryl, etc. Other possible modifications include those described in U.S. Patent Nos. 5,858,988 and 6,291,438.

[0063] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably and generally refer to polymeric forms of nucleotides (deoxyribonucleotides or ribonucleotides) of any length, whether single-stranded, double-stranded, or multi-stranded. Polynucleotides can be exogenous or endogenous to cells. Polynucleotides can exist in cell-free environments. Polynucleotides can be genes or segments thereof. Polynucleotides can be DNA. Polynucleotides can be RNA. Polynucleotides can have any three-dimensional structure and can perform any function. Polynucleotides can contain one or more analogues (e.g., modified backbones, sugars, or nucleic acid bases). If present, modifications to the nucleotide structure can be imparted before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acids, heteronucleic acids, morpholino, locked nucleic acids, glycol nucleic acids, thioarabinonucleotides, dideoxynucleotides, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, piracetin, and woyoside. Non-restricted examples of polynucleotides include coding or non-coding regions of genes or gene segments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched-chain polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of a nucleotide may be interrupted by non-nucleotide components.

[0064] The term “transfection” or “transfected” generally refers to the introduction of nucleic acids into cells via non-viral or virus-based methods. Nucleic acid molecules can be gene sequences encoding complete proteins or functional portions thereof. See, for example, Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1–18.88 (which is incorporated herein by reference in its entirety).

[0065] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein and generally refer to a polymer consisting of at least two amino acid residues linked by peptide bonds. This term does not imply a polymer of a specific length, nor is it intended to suggest or distinguish whether a peptide is produced using recombinant technology, chemical or enzymatic synthesis, or is naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acid chains. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary and / or tertiary structures (e.g., domains). The term also covers amino acid polymers that have been modified, for example, by disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, oxidation, and any other manipulation, such as conjugation with a labeled component. As used herein, the terms “amino acid” and “amino acids” generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include both natural and non-natural amino acids that have been chemically modified to include groups or chemical moieties not naturally present on the amino acid. Amino acid analogs may refer to amino acid derivatives. The term "amino acid" includes both D-amino acids and L-amino acids.

[0066] As used herein, the term "non-natural" generally refers to a nucleic acid or polypeptide sequence not found in native nucleic acids or proteins. Non-natural can refer to an affinity tag. Non-natural can refer to a fusion. Non-natural can refer to a naturally occurring nucleic acid or polypeptide sequence containing mutations, insertions, and / or deletions. Non-natural sequences may present and / or encode an activity (such as enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.), and the nucleic acid and / or polypeptide sequence fused to the non-natural sequence may also present this activity. Non-natural nucleic acid or polypeptide sequences can be genetically engineered to link with naturally occurring nucleic acid or polypeptide sequences (or variants thereof) to generate chimeric nucleic acid and / or polypeptide sequences encoding chimeric nucleic acids and / or polypeptides.

[0067] As used herein, the term "promoter" generally refers to a DNA regulatory region that controls gene transcription or expression. This region may be located near or overlap with a nucleotide region where RNA transcription is initiated. A promoter may contain a specific DNA sequence that binds to a protein factor, often called a transcription factor, which promotes the binding of RNA polymerase to DNA, leading to gene transcription. A 'basal promoter,' also known as a 'core promoter,' generally refers to a promoter that contains all the essential elements that promote the operatively linked transcriptional expression of polynucleotides. Eukaryotic basal promoters typically, but do not necessarily, contain a TATA box and / or a CAAT box.

[0068] As used herein, the term "expression" generally refers to the process of transcribing a nucleic acid sequence or polynucleotide from a DNA template (such as into mRNA or other RNA transcripts) and / or the subsequent translation of the transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotide is derived from genomic DNA, expression in eukaryotic cells may include the splicing of mRNA.

[0069] As used herein, "operably linked," "operably connected," "operably linked," or their grammatical equivalents generally refer to the juxtaposition of gene elements such as promoters, enhancers, polyadenylated sequences, etc., in a relationship that allows them to operate in the intended manner. For example, regulatory elements may include promoter and / or enhancer sequences, and if the regulatory element helps initiate transcription of a coding sequence, then the regulatory element is operably linked to the coding region. Intercalation residues may exist between the regulatory element and the coding region, as long as this functional relationship is maintained.

[0070] As used herein, "vector" generally refers to a macromolecule or macromolecular complex containing or associated with polynucleotides that can be used to mediate the delivery of polynucleotides into cells. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery mediators. Vectors generally contain genetic elements, such as regulatory elements, that are operatively linked to genes to promote gene expression at a target.

[0071] As used herein, "expression cassette" and "nucleic acid cassette" are used interchangeably and generally refer to a combination of nucleic acid sequences or elements that are expressed together or operatively linked together for expression. In some cases, an expression cassette refers to a combination of a regulatory element with one or more genes operatively linked together for expression.

[0072] A "functional segment" of a DNA or protein sequence generally refers to a segment that retains biological activity (function or structure) and whose biological activity is broadly similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be its ability to influence expression in a manner known to be attributable to the full-length sequence.

[0073] As used herein, an "engineered" object generally indicates that the object has been modified through human intervention. By way of non-limiting examples: nucleic acids can be modified by altering their sequences to sequences not found in nature; nucleic acids can be modified by linking them to nucleic acids that do not associate with them in nature, so that the linked product has a function not found in the original nucleic acid; engineered nucleic acids can be synthesized in vitro, and their sequences do not exist in nature; proteins can be modified by altering their amino acid sequences to sequences not found in nature; engineered proteins can acquire new functions or properties. An "engineered" system contains at least one engineered component.

[0074] As used in this article, the term "best alignment" generally refers to an alignment of two amino acid sequences that gives the highest percentage of identity score or the largest number of matching residues.

[0075] As used herein, "synthetic" and "artificial" are used interchangeably and generally refer to proteins or their domains that have low sequence identity with naturally occurring human proteins (e.g., less than 50%, less than 25%, less than 10%, less than 5%, less than 1%). For example, the VPR and VP64 domains are synthetic transactivation domains.

[0076] As used herein, the term "tracrRNA" or "tracr sequence" generally refers to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc., or SEQ ID NO:199-203). tracrRNA can refer to a nucleic acid having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc.). tracrRNA can refer to a modified form of tracrRNA, which may include nucleotide changes such as deletions, insertions or substitutions, variants, mutations, or chimeras. tracrRNA can refer to a nucleic acid that is at least about 60% identical to the sequence of a wild-type exemplary tracrRNA (e.g., tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc.) across at least six consecutive nucleotide segments. For example, the tracrRNA sequence may be at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or 100% identical to the sequence of a wild-type exemplary tracrRNA (e.g., tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc.) across at least six consecutive nucleotide segments. Type II tracrRNA sequences can be predicted on the genome sequence by identifying regions complementary to partially repetitive sequences in adjacent CRISPR arrays.

[0077] As used herein, "guide nucleic acid" generally refers to a nucleic acid that can hybridize with another nucleic acid. A guide nucleic acid can be RNA. A guide nucleic acid can be DNA. A guide nucleic acid can be programmed to specifically bind to a nucleic acid sequence site. The target nucleic acid or the nucleic acid to be targeted can contain nucleotides. A portion of the target nucleic acid can be complementary to a portion of the guide nucleic acid. The strand of a double-stranded target polynucleotide that is complementary to and hybridizes with the guide nucleic acid can be called the complementary strand. The strand of a double-stranded target polynucleotide that is complementary to the complementary strand and therefore may not be complementary to the guide nucleic acid can be called the non-complementary strand. A guide nucleic acid can contain one polynucleotide chain and can be called a "single guide nucleic acid". A guide nucleic acid can contain two polynucleotide chains and can be called a "double guide nucleic acid". Unless otherwise specified, the term "guide nucleic acid" can be inclusive, referring to both single and double guide nucleic acids. A guide nucleic acid can contain a segment that can be called a "nucleic acid targeting region" or "nucleic acid targeting sequence". A nucleic acid targeting region can contain a sub-segment that can be called a "protein-binding region", "protein-binding sequence", or "Cas protein-binding region".

[0078] The terms “sequence identity” or “identity percentage” in the case of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., in paired alignments) or more (e.g., in multiple sequence alignments) sequences that are identical or have a specified percentage of identical amino acid residues or nucleotides, as measured using sequence comparison algorithms, when performing maximum correspondence comparisons and alignments on a local or global comparison window. Suitable sequence comparison algorithms for peptide sequences include, for example, BLASTP with the following parameters: word length (W) of 3, expectation value (E) of 10, and BLOSUM62 scoring matrix with gap cost of 11 and extension value of 1, adjusted using conditional composition scoring matrix for peptide sequences greater than 30 residues; BLASTP with the following parameters: word length (W) of 2, expectation value (E) of 1,000,000, and PAM30 scoring matrix with gap cost of 9 for open gaps and 1 for extension gaps, for sequences less than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); or CLUSTALW, which uses the Smith-Waterman homology search algorithm parameters: match value of 2, mismatch value of -1, and gap value of -1; MUSCLE with default parameters; MAFFT with the following parameters: retree of 2 and maximum iterations of 1000; Novafold with default parameters; and HMMER with default parameters. hmmalign.

[0079] As used herein, the term "RuvC_III domain" generally refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain comprises three discontinuous segments: RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof (e.g., RuvC_I, RuvC_II, or RuvC_III) can generally be identified by alignment with known domain sequences, by structural alignment with proteins having annotated domains, or by comparison with a Hidden Markov Model (HMM) constructed based on known domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).

[0080] As used herein, the term "HNH domain" generally refers to a nuclease domain containing characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment with known domain sequences, by alignment with the structure of proteins with annotated domains, or by comparison with a Hidden Markov Model (HMM) constructed based on known domain sequences (e.g., Pfam HMM PF01844 for the HNH domain).

[0081] As used in this article, the term "bridged helical domain" or "BH domain" generally refers to an arginine-rich helical domain present in Cas enzymes and playing an important role in the initiation of cleavage activity after target DNA binding.

[0082] As used herein, the term “recognition domain” or “REC domain” generally refers to the domain that is believed to interact with the repeat:anti-repeat double helix of gRNA and mediate the formation of the Cas endonuclease / gRNA complex.

[0083] As used herein, the term "wedge domain" or "WED domain" generally refers to a fold containing a twisted five-stranded β-sheet and four α-helices on either side, which is typically responsible for the recognition of the modified repeating antiduplex of the Cas enzyme. The WED domain can also be responsible for the recognition of unidirectional guide RNA scaffolds.

[0084] As used herein, the term “PAM interaction domain” or “PI domain” generally refers to the domain found in the Cas enzyme located in the endonuclease-DNA complex, which is used to recognize the PAM sequence on the non-complementary DNA strand of the guide RNA.

[0085] Overview

[0086] The discovery of novel Cas enzymes with unique functions and structures offers the potential to further disrupt DNA editing technologies, improving speed, specificity, functionality, and ease of use. Compared to the predicted prevalence and sheer diversity of clustered regular-spaced short palindromic repeat (CRISPR) systems in microorganisms, relatively few functionally characterized CRISPR / Cas enzymes exist in the literature. This is partly because a large number of microbial species may not be easily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches containing a large number of microbial species could significantly increase the number of known novel CRISPR / Cas systems and accelerate the discovery of new oligonucleotide editing functions. The discovery of the CasX / CasY CRISPR system through metagenomic analysis of natural microbial communities in 2016 is a recent example demonstrating the success of this approach.

[0087] The CRISPR / Cas system is an RNA-directed nuclease complex described as serving as an adaptive immune system in microorganisms. In its natural context, the CRISPR / Cas system occurs within CRISPR (clustered regularly spaced short palindromic repeats) operators or loci, typically comprising two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences encoding RNA-based targeting elements; and (ii) an ORF encoding Cas, which encodes a nuclease polypeptide directed by both the RNA-based targeting element and a helper protein / enzyme. Effective nuclease targeting of a specific target nucleic acid sequence generally requires: (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence (which is typically an uncommon sequence in the host genome) within the defined area of ​​the target seed. Depending on the exact function and organization of the system, CRISPR-Cas systems are generally classified into 2 classes, 5 types, and 16 subtypes based on shared functional characteristics and evolutionary similarity.

[0088] Class I CRISPR-Cas systems have large multi-subunit effector complexes and include types I, III, and IV.

[0089] Type I CRISPR-Cas systems are considered moderately complex in terms of their components. In a type I CRISPR-Cas system, an array of RNA-targeting elements is transcribed into long precursor crRNAs (crRNA precursors), which are processed at repeat elements to release short, mature crRNAs. These crRNAs then direct the nuclease complex toward the nucleic acid target when followed by a suitable short contiguous sequence called a protospacer adjacent motif (PAM). This processing occurs via a ribonuclease subunit (Cas6) of a large endonuclease complex called a cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nucleases are primarily used as DNA nucleases.

[0090] Type III CRISPR systems are characterized by the presence of a central nuclease called Cas10, and repeat-associated mystery proteins (RAMPs) containing Csm or Cmr protein subunits. As in Type I systems, mature crRNA is processed from crRNA precursors using a Cas6-like enzyme. Unlike Type I and Type II systems, Type III systems appear to target and cleave the DNA-RNA duplex (where the DNA strand is used as a template for RNA polymerase).

[0091] The type IV CRISPR-Cas system has an effector complex consisting of two genes of RAMP proteins from the highly reduced large subunit nuclease (csf1), Cas5 (csf3), and Cas7 (csf2) groups, and in some cases, the predicted small subunit gene; this system is typically found on endogenous plasmids.

[0092] Class II CRISPR-Cas systems typically have single-peptide multi-domain nuclease effectors and include types II, V, and VI.

[0093] Type II CRISPR-Cas systems are considered the simplest in terms of components. In Type II CRISPR-Cas systems, processing the CRISPR array into mature crRNA does not require the presence of specific endonuclease subunits; instead, it requires a small trans-coding crRNA (tracrRNA) whose region is complementary to the array's repetitive sequences. The tracrRNA interacts with the corresponding effector nuclease (e.g., Cas9) and the repetitive sequences to form a precursor dsRNA structure, which is cleaved by endogenous RNase III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. CasII nucleases are called DNA nucleases. Type II effectors generally exhibit a structure consisting of RuvC-like endonuclease domains that employ RNase H folds, with an unrelated HNH nuclease domain inserted within the fold of the RuvC-like endonuclease domain. The RuvC-like domain is responsible for the cleavage of the target DNA strand (e.g., complementary to the crRNA), while the HNH domain is responsible for the cleavage of the trans DNA strand.

[0094] Type V CRISPR-Cas systems are characterized by nuclease effectors (e.g., Cas12) with structures similar to type II effectors, containing a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process the pre-crRNA into mature crRNA; however, unlike type II systems which require RNase III to cleave the pre-crRNA into multiple crRNAs, type V systems can use the effector nuclease itself to cleave the pre-crRNA. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are also known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-stranded, non-specific deoxyribonuclease activity, which can be activated by targeted cleavage of the first crRNA in a double-stranded target sequence.

[0095] Type VI CRISPR-Cas systems possess RNA-directed RNA endonucleases. The single polypeptide effector of a type VI system (e.g., Cas13) contains two HEPN ribonuclease domains, rather than a RuvC-like domain. Unlike both type II and type V systems, type VI systems do not appear to require tracrRNA to process the crRNA precursor into crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-stranded nonspecific nuclease (ribonuclease) activity, activated by the first crRNA-directed cleavage of the target RNA.

[0096] Due to its simpler architecture, class II CRISPR-Cas has been most widely used in the engineering and development of nuclease / genome editing applications.

[0097] One of the early adaptations of this system for in vitro use can be found in Jinek et al. (Science. 2012 Aug 17; 337(6096):816-21, which is incorporated herein by reference in its entirety). Jinek’s study first describes a system involving (i) recombinantly expressed, purified full-length Cas9 (e.g., type II Cas enzyme), isolated from Streptococcus pyogenes SF370; (ii) purified mature approximately 42 nt crRNA carrying approximately 20 nt 5' sequence complementary to the target DNA sequence, which needs to be cleaved after the 3' tracr-binding sequence (the entire crRNA is transcribed in vitro from a synthetic DNA template carrying the T7 promoter sequence); (iii) purified tracrRNA, which is transcribed in vitro from a synthetic DNA template carrying the T7 promoter sequence; and (iv) Mg 2+ Jinek later described an improved engineered system in which the crRNA of (ii) binds to the 5' end of (iii) via a linker (e.g., GAAA) to form a single fused synthetic guide RNA (sgRNA) capable of autonomously guiding Cas9 to its target (compare). Figure 2 (The top and bottom images).

[0098] Mali et al. (Science. 2013 Feb 15; 339(6121):823-826.) (incorporated herein by reference in its entirety) later adapted this system for use in mammalian cells by providing DNA vectors encoding the following: (i) an ORF encoding a codon-optimized Cas9 (e.g., type II Cas enzyme) with a suitable mammalian promoter having a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal); and (ii) an ORF encoding an sgRNA (having a 5' sequence starting with G, followed by a 20nt complementary target nucleic acid sequence that binds to a 3' tracr-binding sequence, an adapter, and a tracrRNA sequence) with a suitable polymerase III promoter (e.g., U6 promoter).

[0099] MG enzyme

[0100] On one hand, this disclosure provides an engineered nuclease system. The engineered nuclease system may comprise (a) a nuclease. In some cases, the nuclease comprises a RuvC domain and an HNH domain. The nuclease may be derived from uncultured microorganisms. The nuclease may be a Cas nuclease. The nuclease may be a type II nuclease. The nuclease may be a type II Cas nuclease. The engineered nuclease system may comprise (b) an engineered guide ribonucleic acid (DRRNA) structure. The engineered guide ribonucleic acid structure may be configured to form a complex with the nuclease. In some cases, the engineered guide ribonucleic acid structure configured to form a complex with the nuclease comprises a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize with a target deoxyribonucleic acid (DNA) sequence. In some cases, the engineered guide ribonucleic acid structure configured to form a complex with the nuclease comprises a tracr ribonucleic acid (TRACR) sequence. The TRACR ribonucleic acid (TRACR) sequence may be configured to bind to the nuclease. In some cases, endonucleases have the following molecular weights: about 120 kDa or less, about 110 kDa or less, about 100 kDa or less, about 90 kDa or less, about 80 kDa or less, about 70 kDa or less, about 60 kDa or less, about 50 kDa or less, about 40 kDa or less, about 30 kDa or less, about 20 kDa or less, or about 10 kDa or less.

[0101] In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668.

[0102] On one hand, this disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) a nuclease. The nuclease may include a RuvC-1 domain or a RuvC domain. The nuclease may include an HNH domain. The nuclease may include a RuvC-1 domain or an HNH domain. The nuclease may be a Cas nuclease. The nuclease may be a type II nuclease. The nuclease may be a type II Cas nuclease. The engineered nuclease system may include (b) an engineered guide ribonucleic acid (BRNA). The engineered guide BRNA structure may be configured to form a complex with the nuclease. The guide BRNA structure configured to form a complex with the nuclease contains a guide BRNA sequence. The guide BRNA sequence may be configured to hybridize with a target deoxyribonucleic acid (DNA) sequence. The engineered guide BRNA structure configured to form a complex with the nuclease may contain a tracr BRNA sequence. The tracr ribonucleic acid sequence can be configured to bind to an endonuclease. The endonuclease may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the sequences 1-198, 221-459, 463-612, or 617-668. The endonuclease may be an archaeological endonuclease. The endonuclease may be a type II Cas endonuclease. Nucleotide endonucleases may contain arginine-rich regions with RRxRR motifs or domains homologous to PF14239. The arginine-rich region or the domain homologous to PF14239 may contain a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the arginine-rich regions or the domain homologous to PF14239 in any of SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. Domain boundaries of arginine-rich domains or domains homologous to PF14239 can be identified by optimal alignment with MG34-1 or MG34-9. Endonucleases may contain REC domains.The REC domain may contain a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the REC domains in SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the REC domain can be identified by optimal alignment with MG34-1 or MG34-9. The endonuclease may contain a BH (bridged helix) domain. The BH domain may contain a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the BH domains in SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the BH domain can be identified by best alignment with MG34-1 or MG34-9.

[0103] Endonucleases may contain a WED (wedge-shaped) domain. The WED domain may contain a sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the WED domains in SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the WED domain can be identified by optimal alignment with MG34-1 or MG34-9. Endonucleases may contain a PI (PAM interaction) domain. The PI domain may contain a sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the PI domains in SEQ ID NO: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the PI domain can be identified by optimal alignment with MG34-1 or MG34-9.

[0104] In some cases, the endonuclease is derived from uncultured microorganisms. In some cases, the tracr ribonuclease sequence comprises at least 50, at least 60, at least 70, or at least 80 consecutive nucleotides having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of SEQ ID NO:199-200, 460-461, or 669-673, or a sequence with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO:199-200, 460-461, or 669-673. A sequence of at least 50, at least 60, at least 70, or at least 80 consecutive nucleotides of any one of NO:201-203 or 613-616 having sequence identity of at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.

[0105] In some cases, the guide nucleic acid structure includes SEQ ID NO:201. In some cases, the guide nucleic acid structure includes SEQ ID NO:202. In some cases, the guide nucleic acid structure includes SEQ ID NO:203. In some cases, the guide nucleic acid structure includes SEQ ID NO:201-203. In some cases, the guide nucleic acid structure includes SEQ ID NO:613. In some cases, the guide nucleic acid structure includes SEQ ID NO:614. In some cases, the guide nucleic acid structure includes SEQ ID NO:615. In some cases, the guide nucleic acid structure includes SEQ ID NO:616.

[0106] On one hand, this disclosure provides an engineered nuclease system. The engineered nuclease system may comprise (a) an engineered guide ribonucleic acid (BRRNA) structure. The engineered guide BRRNA structure may comprise a guide ribonucleic acid (BRRNA) sequence. The guide BRRNA sequence may be configured to hybridize with a target deoxyribonucleic acid (DNA) sequence. The engineered guide BRRNA structure may comprise a tracr ribonucleic acid (TRACR) sequence. The TRACR BRRNA sequence may be configured to bind to a nuclease endonuclease. In some cases, the tracr ribonucleic acid sequence comprises a sequence having at least 50, 60, 70, or 80 consecutive nucleotides with at least 50%, 55%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NO: 199-200, 460-461, or 669-673, or a sequence with at least 50%, 60%, 70%, 80%, 95%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 199-200, 460-461, or 669-673, or a sequence with at least 50%, 60%, 70%, 80 ...7%, 98%, or 99% sequence identity with SEQ ID NO: 199-200, 460-461, or 669-673, or a sequence with at least 50%, 60%, 70%, 80%, 97%, 98%, or 99 A sequence of at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, or at least 80 consecutive nucleotides of any one of NO:201-203 or 613-616 having sequence identity of at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.

[0107] In some cases, engineered nuclease systems contain endonucleases. These endonucleases can be type II endonucleases. They can be Cas endonucleases. They can be type II Cas endonucleases.

[0108] In some cases, endonucleases have specific molecular weight ranges. In some embodiments, endonucleases have the following molecular weights: about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, about 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, engineered guide ribonucleic acid structures contain at least two ribonucleic acid polynucleotides. In some cases, endonucleases contain a specific number of residues. Endonucleases can contain 1,100 or less, 1,000 or less, 950 or less, 900 or less, 850 or less, 800 or less, 750 or less, 700 or less, 650 or less, 600 or less, 550 or less, 500 or less, 450 or less, 400 or less, or 350 or less. Endonucleases can contain from about 700 to about 1,100 residues. Endonucleases can contain from about 400 to about 600 residues. In some cases, engineered guide ribonucleic acid (BRRNA) structures comprise a single ribonucleic acid polynucleotide (RNAP). A single RNAP may comprise a guide ribonucleic acid sequence and a TRACR RNA sequence.

[0109] In some cases, the guide RNA sequence is complementary to the genome sequences of prokaryotes, bacteria, archaea, eukaryotes, fungi, plants, mammals, or humans. Specifically, it is complementary to the genome sequences of prokaryotes, bacteria, archaea, eukaryotes, eukaryotes, fungi, fungi, plants, plants, mammals, and humans.

[0110] In some cases, the guide ribonucleic acid targeting sequence or spacer is 10-30 nucleotides, 12-28 nucleotides, or 15-24 nucleotides in length. In some cases, the endonuclease contains one or more nuclear localization sequences (NLS) located near the N-terminus or C-terminus of the endonuclease. In some cases, the NLS contains sequences selected from SEQ ID NO: 205-220.

[0111] Table 1: Exemplary NLS sequences that can be used with Cas effectors according to this disclosure

[0112]

[0113]

[0114] This disclosure includes variants of any of the enzymes described herein having one or more conserved amino acid substitutions. Such conserved substitutions can occur in the amino acid sequence of a polypeptide without disrupting its three-dimensional structure or function. Conservative substitutions can be achieved by substituting amino acids of similar hydrophobicity, polarity, and R-chain length. Furthermore or alternatively, conserved substitutions can be identified by locating mutated amino acid residues (e.g., non-conserved residues) between species without altering the fundamental function of the encoded protein, through comparison of aligned sequences of homologous proteins from different species. Such conserved substitution variants may include variants having the following identity with any of the endonuclease protein sequences described herein: at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%. In some embodiments, such conserved substitution variants are functional variants. Such functional variants may cover sequences with substitutions such that the activity of one or more key active site residues of the endonuclease or guide RNA binding residues is not disrupted. In some implementations, any functional variant of the protein described herein is lacking. Figure 4 The substitution of at least one conserved or functional residue shown herein. In some embodiments, any functional variant of the protein described herein lacks... Figure 4 Substitution of all conserved or functional residues shown herein. The disclosure also provides variants of any nuclease activity alterations described herein. Such variants of activity alterations can be found in those identified herein (e.g., in…). Figure 4The active variant may contain an inactivation mutation in one or more catalytic residues of the RuvC domain (or generally the RuvC domain). Such an activity-altering variant may contain a switching mutation in the catalytic residues of the RuvCI, RuvCII, or RuvCIII domains.

[0115] Conserved substitutions of functionally similar amino acids can be obtained from various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conserved substitutions for each other:

[0116] 1) Alanine (A), glycine (G);

[0117] 2) Aspartic acid (D), glutamic acid (E);

[0118] 3) Asparagine (N), glutamine (Q);

[0119] 4) Arginine (R), Lysine (K);

[0120] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);

[0121] 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);

[0122] 7) Serine (S), threonine (T); and

[0123] 8) Cysteine ​​(C), Methionine (M).

[0124] This disclosure includes variants of any endonuclease described herein that have sequence identity with a particular domain. The domain may be an arginine-rich domain (e.g., a domain homologous to PF14239), a REC (recognition) domain, a BH (bridged helix) domain, a WED (wedge) domain, a PI (PAM interaction) domain, a PF14239 homology domain, or any other domain described herein. In some embodiments, residues containing one or more of these domains are identified in the protein by alignment with a protein described, for example, when the residue boundaries of the domain are described.

[0125] Table 2: Exemplary domain boundaries of the endonucleases described in this paper

[0126]

[0127] In some cases, engineered nuclease systems also include a single-stranded DNA repair template. In some cases, engineered nuclease systems also include a double-stranded DNA repair template. In some cases, the single-stranded or double-stranded DNA repair template includes a first homologous arm from 5' to 3', the first homologous arm containing a sequence of at least 20 nucleotides at the 5' position of the target deoxyribonucleic acid sequence. In some cases, the single-stranded or double-stranded DNA repair template includes a synthetic DNA sequence having at least 10 nucleotides from 5' to 3'. In some cases, the single-stranded or double-stranded DNA repair template includes a second homologous arm from 5' to 3', the second homologous arm containing a sequence of at least 20 nucleotides at the 3' position of the target sequence. In some cases, the single-stranded or double-stranded DNA repair template from 5' to 3' includes: a first homologous arm containing a sequence of at least 20 nucleotides at the 5' position of the target deoxyribonucleic acid sequence; a synthetic DNA sequence having at least 10 nucleotides; or a second homologous arm containing a sequence of at least 20 nucleotides at the 3' position of the target sequence.

[0128] In some cases, the first homologous arm contains a sequence of at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. In some cases, the engineered nuclease system also contains Mg. 2+ Source. In some cases, the endonuclease and tracr ribonucleic acid sequences originate from different bacterial species. In other cases, the endonuclease and tracr ribonucleic acid sequences originate from different bacterial species within the same phylum.

[0129] In some cases, the endonuclease comprises a sequence having at least 50%, 55%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NO: 1-24 or 462-488. In some cases, the guide RNA structure comprises an RNA sequence predicted to contain a hairpin. In some cases, the hairpin comprises a stem and a loop. In some cases, the stem comprises at least 12, 14, 16, or 18 pairs of ribonucleotides.

[0130] In some cases, the guide RNA structure also includes a second stem and a second loop. In some cases, the second step includes at least 5, 6, 7, 8, 9, or 10 pairs of ribonucleotides. In some cases, the guide RNA structure also includes an RNA structure containing at least two hairpins. In some cases, the endonuclease includes a sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:1, and the guide RNA structure includes an RNA sequence predicted to contain at least four hairpins. In some cases, each of these four hairpins includes a stem and a loop.

[0131] In some cases, the engineered nuclease system contains at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the same sequence as SEQ ID NO:1. In some cases, the engineered nuclease system comprises a guide RNA structure containing at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the sequence identical to at least one of the non-variable nucleotides in SEQ ID NO:199 or SEQ ID NO:201.

[0132] In some cases, the engineered nuclease system contains at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the sequence identical to any one of SEQ ID NO:1-24 or 462-488. In some cases, the engineered nuclease system contains at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the same sequence as any one of SEQ ID NO:199-200 or 669-673 or any one of SEQ ID NO:201-203 or 613-616.

[0133] In some cases, sequence identity is determined using BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW with Smith-Waterman homology search algorithm parameters. In other cases, sequence identity is determined using the BLASTP homology search algorithm, which uses the following parameters: word length (W) of 3, expected value (E) of 10, and a BLOSUM62 scoring matrix with a gap cost of 11, an extension value of 1, and conditional composition of the scoring matrix adjustment.

[0134] In some cases, the endonuclease is not a Cas9 endonuclease, Cas14 endonuclease, Cas12a endonuclease, Cas12b endonuclease, Cas12c endonuclease, Cas12d endonuclease, Cas12e endonuclease, Cas13a endonuclease, Cas13b endonuclease, Cas13c endonuclease, or Cas13d endonuclease. In some cases, the endonuclease has less than 80%, less than 75%, less than 70%, less than 65%, less than 60%, less than 55%, or less than 50% identity with the Cas9 endonuclease.

[0135] On one hand, this disclosure provides an engineered guide RNA comprising (a) a DNA targeting segment. In some cases, the DNA targeting segment comprises a nucleotide sequence complementary to a target sequence in a target DNA molecule. In some cases, the engineered single-stranded guide RNA polynucleotide comprises a protein-binding segment. The protein-binding segment comprises two complementary nucleotide segments that hybridize to form a double-stranded RNA (dsRNA) duplex. In some cases, the two complementary nucleotide segments are covalently linked to each other by intercalation nucleotides. In some cases, engineered guide ribonucleic acid polynucleotides are configured to form a complex with a nuclease comprising a variant of any one of SEQ ID NO: 1-198, 221-459, 463-612 or 617-668 having a sequence identity of at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.

[0136] In some cases, the DNA targeting segment is located at the 5' of both complementary nucleotide segments. In some cases, the protein-binding segment contains at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the non-variable nucleotides identical to any of SEQ ID NO:199-200 or 669-673 or SEQ ID NO:201-203 or 613-616. In some cases, the deoxyribonucleic acid polynucleotide encodes the engineered guide ribonucleic acid polynucleotide described herein.

[0137] On one hand, this disclosure provides a nucleic acid comprising an engineered nucleic acid sequence. In some cases, the engineered nucleic acid sequence is optimized for expression in organisms. In some cases, the nucleic acid encodes a nuclease. The nuclease may be a Cas nuclease. The nuclease may be a class II nuclease. The nuclease may be a class II type II Cas nuclease. In some cases, the nuclease comprises a RuvC domain and an HNH domain. In some cases, the nuclease is derived from uncultured microorganisms. In some cases, the nuclease has a specific molecular weight range. In some embodiments, the endonuclease has the following molecular weights: about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, about 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, the engineered guide ribonucleic acid structure contains at least two ribonucleic acid polynucleotides. In some cases, the endonuclease contains a specific number of residues. Endonucleases can contain 1,100 or less, 1,000 or less, 950 or less, 900 or less, 850 or less, 800 or less, 750 or less, 700 or less, 650 or less, 600 or less, 550 or less, 500 or less, 450 or less, 400 or less, or 350 or less. Endonucleases can contain approximately 700 to approximately 1,100 residues. Endonucleases can contain approximately 400 to approximately 600 residues. In some cases, the endonuclease comprises SEQ ID NO:1-198, 221-459, 463-612 or 617-668 or variants thereof, said variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with it.In some cases, the endonuclease also includes a sequence encoding one or more nuclear localization sequences (NLS) located near the N-terminus or C-terminus of the endonuclease. In some cases, the NLS includes sequences selected from SEQ ID NO: 205-220.

[0138] In some cases, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human. In some cases, the organism is a prokaryote. In some cases, the organism is a bacterium. In some cases, the organism is a eukaryote. In some cases, the organism is a fungus. In some cases, the organism is a plant. In some cases, the organism is a mammal. In some cases, the organism is a rodent. In some cases, the organism is a human. If the organism is a prokaryote or bacterium, it can be an organism different from the organism from which the endonuclease originated. In some cases, the organism is not an uncultured microorganism.

[0139] On one hand, this disclosure provides a vector containing a nucleic acid sequence. In some cases, the nucleic acid sequence encodes a nuclease. In some cases, the nuclease is a Cas nuclease. In some cases, the nuclease is a type 2 nuclease. In some cases, the nuclease is a type 2 type II Case nuclease. The nuclease may contain a RuvC-I domain and an HNH domain. In some cases, the nuclease is derived from uncultured microorganisms. In some cases, the nuclease has a specific molecular weight range. In some embodiments, the endonuclease has the following molecular weights: about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, about 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, the engineered guide ribonucleic acid structure contains at least two ribonucleic acid polynucleotides. In some cases, the endonuclease contains a specific number of residues. Endonucleases can contain 1,100 or less, 1,000 or less, 950 or less, 900 or less, 850 or less, 800 or less, 750 or less, 700 or less, 650 or less, 600 or less, 550 or less, 500 or less, 450 or less, 400 or less, or 350 or less. Endonucleases can contain approximately 700 to approximately 1,100 residues. Endonucleases can contain approximately 400 to approximately 600 residues.

[0140] In some aspects, this disclosure provides a nuclease described herein configured to induce double-strand breaks at the 5' position of a protospacer adjacent motif (PAM) near the target locus. The nuclease may induce double-strand breaks at a distance of 6-8 nucleotides or 7 nucleotides from the PAM. In some aspects, this disclosure provides a nuclease described herein configured to induce single-strand breaks at the 5' position of a protospacer adjacent motif (PAM) near the target locus. The nuclease may induce single-strand breaks at a distance of 6-8 nucleotides or 7 nucleotides from the PAM. In some aspects, the nuclease configured to induce single-strand breaks contains an inactivating mutation in one or more catalytic residues of the nuclease described herein.

[0141] In some aspects, this disclosure provides a nuclease system as described herein, configured to induce chemical modification of nucleotide bases within or near a target locus targeted by the nuclease system. In this case, the chemical modification of the nucleotide bases generally refers to modification of the chemical portion involved in base pairing, rather than modification of the sugar or phosphate portion of the nucleotide. Chemical modification may include deamination of adenosine or cytosine nucleotides. In some cases, the nuclease system configured to induce chemical modification comprises a nuclease having a base editor coupled to or fused within the frame of the nuclease. The nuclease fused to or coupled to the base editor may contain an inactivation mutation in at least one catalytic residue of the nuclease (e.g., in the RuvC domain). The base editor may be fused to the N-terminus or C-terminus of the nuclease, or linked via chemical conjugation. The base editor may include any adenosine or cytosine deaminase, including but not limited to adenosine deaminase RNA-specific 1 (ADAR1), adenosine deaminase RNA-specific 2 (ADAR2), apolipoprotein B mRNA editing enzyme catalytic subunit 1 (APOBEC1), apolipoprotein B mRNA editing enzyme catalytic subunit 2 (APOBEC2), apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A), apolipoprotein B mRNA editing enzyme catalytic subunit 3B (APOBEC3B), apolipoprotein B mRNA editing enzyme catalytic subunit 3C (APOBEC3C), apolipoprotein B mRNA editing enzyme catalytic subunit 3D (APOBEC3D), apolipoprotein B mRNA editing enzyme catalytic subunit 3F (APOBEC3F), apolipoprotein B mRNA editing enzyme catalytic subunit 3G (APOBEC3G), apolipoprotein B mRNA editing enzyme catalytic subunit 3H (APOBEC3H), or apolipoprotein B mRNA editing enzyme catalytic subunit 4 (APOBEC4) or a functional fragment thereof. Base editors can include yeast, eukaryotic, mammalian, or human base editors.

[0142] In some aspects, this disclosure provides a nuclease system as described herein, configured to induce histone chemical modifications within or near a target locus targeted by the nuclease system. In some cases, the nuclease system configured to induce histone chemical modifications comprises a nuclease having a histone editor coupled to or fused within the frame of the nuclease. The histone editor may be coupled to or fused to the nuclease at the N-terminus or C-terminus. In some embodiments, the chemical modification may include methylation, acetylation, demethylation, or deacetylation. The nuclease fused to or coupled to the histone editor may contain an inactivating mutation in at least one catalytic residue of the nuclease (e.g., in the RuvC domain). Histone editors may include histone methyltransferases (e.g., ASH1L, DOT1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, PRDM2, SET, SETBP1, SETD1A, SETD1B, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETD9, SETDB1, SETDB). 2. Histone editors may include SETMAR, SMYD1, SMYD2, SMYD3, SMYD4, SMYD5, SUV39H1, SUV39H2, SUV420H1, or SUV420H2; histone demethylases (e.g., the KDM1, KDM2, KDM3, KDM4, KDM5, or KDM6 family); histone acetyltransferases (e.g., the GNAT or HAT family acetyltransferases); or histone deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC4, HDAC5, HDAC6, HDAC7, HDAC8, HDAC9, HDAC10, HDAC11, SIRT1, SIRT2, SIRT3, SIRT4, SIRT5, SIRT6, or SIRT7). Histone editors may include yeast, eukaryotic, mammalian, or human histone editors.

[0143] On one hand, this disclosure provides a vector containing the nucleic acid described herein. In some cases, the vector further contains nucleic acid encoding an engineered guide ribonucleic acid (BRNA) structure. The engineered guide ribonucleic acid structure can be configured to form a complex with a nuclease. In some cases, the engineered guide ribonucleic acid structure contains a guide ribonucleic acid sequence. In some cases, the guide ribonucleic acid sequence is configured to hybridize with a target deoxyribonucleic acid (DNA) sequence. In some cases, the engineered guide ribonucleic acid structure contains a tracr ribonucleic acid (TRACR) sequence. In some cases, the TRACR ribonucleic acid (TRACR) sequence is configured to bind to a nuclease. In some cases, the vector is a plasmid, microcircle, CELiD, adeno-associated virus (AAV)-derived viral particle, or lentivirus.

[0144] On the one hand, this disclosure provides a cell that includes any of the vectors described herein.

[0145] On one hand, this disclosure provides a method for producing endonucleases. The method may include culturing any of the cells described herein.

[0146] On one hand, this disclosure provides a method for binding, cleaving, labeling, or modifying double-stranded deoxyribonucleic acid (DDNA) polynucleotides. The method may include contacting the DDNA polynucleotide with a nuclease. In some cases, the nuclease is a Cas nuclease. In some cases, the nuclease is a type 2 nuclease. In some cases, the nuclease is a type 2 type II Cas nuclease. The nuclease may be complexed with an engineered guide ribonucleic acid structure. In some cases, the engineered guide ribonucleic acid structure is configured to bind to both the nuclease and the DDNA polynucleotide. In some cases, the DDNA polynucleotide contains a protospacer adjacent motif (PAM). In some cases, endonucleases have the following molecular weights: about 120 kDa or less, about 110 kDa or less, about 100 kDa or less, about 90 kDa or less, about 80 kDa or less, about 70 kDa or less, about 60 kDa or less, about 50 kDa or less, about 40 kDa or less, about 30 kDa or less, about 20 kDa or less, or about 10 kDa or less. In some cases, the endonuclease comprises a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668.

[0147] On one hand, this disclosure provides a method for binding, cleaving, labeling, or modifying double-stranded deoxyribonucleic acid (DDNA) polynucleotides. The method may include contacting the DDNA polynucleotide with a nuclease. In some cases, the nuclease is a Cas nuclease. In some cases, the nuclease is a type 2 nuclease. In some cases, the nuclease is a type 2 type II Cas nuclease. The nuclease may be complexed with an engineered guide ribonucleic acid structure. In some cases, the engineered guide ribonucleic acid structure may be configured to bind to both the nuclease and the DDNA polynucleotide. In some cases, the DDNA polynucleotide contains a protospacer adjacent motif (PAM). In some cases, the PAM is NGG. In some cases, the endonuclease comprises a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any one of SEQ ID NO:1-198, 221-459, 463-612, or 617-668.

[0148] In some cases, the endonuclease is not Cas9, Cas14, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, or Cas13d. In some cases, the endonuclease originates from uncultured microorganisms. In some cases, the double-stranded deoxyribonucleic acid (DDNA) is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human DDNA. In some cases, the DDNA is a prokaryotic, archaeal, or bacterial DDNA from a species other than the species from which the endonuclease originates.

[0149] On one hand, this disclosure provides a method for modifying a target nucleic acid locus. The method may include delivering the engineered nuclease system described herein to the target nucleic acid locus. In some cases, the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure. In some cases, the complex is configured such that when the complex binds to the target nucleic acid locus, the complex modifies the target nucleic acid locus. In some cases, modifying the target nucleic acid locus includes binding, cleaving, splitting, or labeling the target nucleic acid locus.

[0150] In some cases, the target nucleic acid locus includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, the target nucleic acid includes genomic eukaryotic DNA, viral DNA, or bacterial DNA. In some cases, the target nucleic acid includes bacterial DNA. Bacterial DNA can originate from a bacterial species different from the species from which the endonuclease originates. In some cases, the target nucleic acid locus is in vitro. In some cases, the nucleic acid locus is intracellular. In some cases, the endonuclease and engineered guide nucleic acid structure are encoded by separate nucleic acid molecules. In some cases, the cell is a prokaryotic cell, bacterial cell, eukaryotic cell, fungal cell, plant cell, animal cell, mammalian cell, rodent cell, primate cell, or human cell. In some cases, the cell originates from a species different from the species from which the endonuclease originates.

[0151] In some cases, delivering an engineered nuclease system to a target nucleic acid locus includes delivering the nucleic acid or the vector described herein. In some cases, delivering an engineered nuclease system to a target nucleic acid locus includes delivering a nucleic acid containing an open reading frame encoding a nuclease. In some cases, the nucleic acid contains a promoter operatively linked to an open reading frame encoding a nuclease. In some cases, delivering an engineered nuclease system to a target nucleic acid locus includes delivering a capped mRNA containing an open reading frame encoding the nuclease. In some cases, delivering an engineered nuclease system to the target nucleic acid locus includes delivering a translated polypeptide.

[0152] In some cases, engineered nuclease delivery systems to target nucleic acid loci include the delivery of deoxyribonucleic acid (DNA) encoding an engineered guide ribonucleic acid structure operatively linked to the ribonucleic acid (RNA) polIII promoter. In some cases, the endonuclease induces single-strand or double-strand breaks at or near the target locus.

[0153] The systems disclosed herein can be used for a variety of applications, such as nucleic acid editing (e.g., gene editing) and binding to nucleic acid molecules (e.g., sequence-specific binding). Such systems can be used, for example, to address (e.g., remove or replace) genetic mutations that may cause a target disease; to inactivate genes to determine their function in cells; as diagnostic tools for detecting pathogenic gene components (e.g., via the cleavage of retroviral RNA or amplified DNA sequences encoding pathogenic mutations); as inactivating enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria); to render viruses inactive or unable to infect host cells by targeting viral genomes; to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites; to establish gene drivers for evolutionary selection; and as biosensors to detect interference from foreign small molecules and nucleotides in cells.

[0154] Example

[0155] Example 1 - Discovery of novel Cas effectors through metagenomics

[0156] Metagenomics mining

[0157] Metagenomic samples were collected from sediments, soil, and animals. Deoxyribonucleic acid (DNA) was extracted using the Zymobiomics DNA mini-prep kit and analyzed in Illumina. Sequencing was performed at 2500 bp. Samples were collected with the owner's consent. DNA was extracted from samples using the Qiagen DNeasy PowerSoil kit or the Zymo BIOMICS DNA Miniprep kit. The DNA was sent to the Vincent J. Coates Genome Sequencing Laboratory at UC Berkeley for sequencing library preparation (Illumina TruSeq) and sequencing on an Illumina HiSeq 4000 or Novaseq (150-base-pair (bp) reads, target insert size 400–800 bp). Additionally, publicly available high-temperature, soil, and marine metagenomic sequencing data were downloaded from NCBI SRA. Sequencing reads were trimmed using BBMap (Bushnell B., sourceforge.net / projects / bbmap / ) and assembled using Megahit (https: / / paperpile.com / c / QSZG6K / clMrh). Protein sequences were predicted using Prodigal (https: / / paperpile.com / c / QSZG6K / BJ6oW). HMM maps of known type II CRISPR nucleases were constructed using HMMER3 (hmmer.org), and a search was performed for all predicted proteins. Minced ( https: / / github.com / ctSkennerton / minced or (https: / / paperpile.com / c / QSZG6K / OPC44) Predicts CRISPR arrays on assembled contigs. Proteins are classified using Kaiju (https: / / paperpile.com / c / QSZG6K / nMi6k), and contig classification is determined by finding common sequences among all encoded proteins.

[0158] MAFFT (https: / / paperpile.com / c / QSZG6K / sVHNH) was used to align predicted and reference (e.g., SpCas9, SaCas9, and AsCas9) type II effector proteins, and FastTree2 (https: / / paperpile.com / c / QSZG6K / osZNM) was used to infer phylogenetic trees. Novel families were identified from the evolutionary branches formed by the sequences obtained in this study. From the families, candidates were selected if they contained all the components necessary for laboratory analysis (i.e., they were present on well-assembled annotated contigs with CRISPR arrays). MUSCLE (https: / / paperpile.com / c / QSZG6K / ITOla) was used to align the selected representative and reference sequences to identify catalytic and PAM interacting residues.

[0159] This metagenomics workflow produces a profile of the SMART (SMall ARchaeal-associated) endonuclease system described in this paper.

[0160] Discovery of SMART endonucleases containing active residues

[0161] Mining of tens of thousands of high-quality CRISPR Cas systems assembled from metagenomic data revealed novel effectors containing both RuvC and HNH domains, but with unusually small sizes (<900 aa). These effector nucleases showed only low sequence similarity (amino acid identity <20%) to the archaea Cas9 endonuclease used as a reference. Phylogenetic analysis of the effector protein sequences indicated that the SMART system is a divergent group relative to well-studied type II systems from subtypes A, B, or C. Figure 1 A).

[0162] These compact "SMART" effectors (approximately 400-1000 amino acids, Figure 2 These loci appear in the genome adjacent to the CRISPR array. Some of these adjacent SMART loci also include sequences predicting the encoding of tracrRNA and CRISPR adaptation genes (e.g., genes involved in spacer acquisition) cas1, cas2, and / or cas4 within the same operon. Figure 3 Despite its compact size, the SMART effector contains six putative HNH and RuvC catalytic residues when aligned with the reference SaCas9 sequence. Figure 4 Furthermore, 3D structure prediction identified residues involved in guidance and target binding as well as PAM recognition, indicating that the SMART effector is an active dsDNA endonuclease. Multiple groups of SMART endonucleases...

[0163] Based on the location of important catalytic and binding residues, SMART nucleases contain three RuvC domains, an arginine-rich region typically containing an RRxRR motif (e.g., a domain homologous to PF14239), an HNH endonuclease domain, and a putative recognition domain. Figure 5 and Figure 6 These domains have low sequence similarity to the reference sequence. Figure 7 Furthermore, the frequencies of RRxRR and zinc-binding band motifs (CX[2-4]C or CX[2-4]H) in SMART effectors and reference archaea sequences were significantly higher than those in Cas9 nucleases. Figure 8Furthermore, unlike Cas9 effector sequences, most SMART effectors show significant hits to the Pfam domain PF14239, which is generally associated with different endonucleases. Based on differences in SMART effector size, phylogenetic relationships, and both operon and domain architecture, we classify these systems into two main groups: SMART I and SMART II. Table 3 below summarizes the key characteristics of these groups and also illustrates the differences compared to the two types of type II A / B / C Cas enzymes.

[0164] Table 3: Properties of SMART Group I and II enzymes described in this paper

[0165]

[0166] SMART I endonuclease

[0167] SMARTI effectors range in size from approximately 700 to 1,050 amino acids. Their common genomic background features adaptive module genes (e.g., genes involved in spacer acquisition) and predicted tracrRNAs near the CRISPR array, organized similarly to type II and type V CRISPR systems. Figure 3 A, 3B, and 3C). The region containing the RRXRR motif in the SMART I effector is unique but may function similarly to the arginine-rich bridging helix in the Cas9 nuclease. When modeling the SaCas9 crystal structure, the predicted 3D structure of the SMART I effector shows unaligned regions within the recognition leaf (typically containing the Pfam domain PF14239) and the RuvC-II domain. Figure 5 The results indicate that these domains have a different origin relative to other type II effectors. This is further supported by their divergent positions in the type II effector phylogenetic tree and their low sequence similarity to known type II effectors. Figure 1 A) These results indicate that the SMART I endonuclease belongs to a new group of type II CRISPR systems. According to the accepted classification of CRISPR systems, these SMART I systems are divided into types II-D.

[0168] Environmental RNA expression data from the SMART I MG34-1 system were used to engineer putative single-guide RNAs (sgRNAs). Furthermore, multiple sgRNAs predicted and designed from SMART I repetitive sequences and tracrRNA were tested in vitro in a PAM enrichment assay. In the case of SMART I enzymes, end repair and blunt-end ligation were used in this step for optimal identification of PAM sequences, demonstrating that these enzymes can generate staggered double-strand DNA breaks. The assay confirmed dsDNA cleavage of MG34-1 (SEQ ID NO:2), MG34-9 (SEQ ID NO:9), and MG34-16 (SEQ ID NO:17) designed with various sgRNAs. Figure 7 The use of SEQ ID NO:612-615 is described. MG34-1 exhibits a preference for NGGN PAM in target recognition and fragmentation. Figure 8 A). Analysis of the cleavage site indicates preferential cleavage at position 7 ( Figure 8 B). These results indicate that this is a novel biochemical mechanism compared to the cleavage mechanisms of other type II enzymes, which preferentially cleave from position 2-3 of PAM, thus supporting a new classification of the SMARTICRISPR system.

[0169] Some SMART I system environmental expression data confirmed in situ transcription of the CRISPR array and the intergenic region encoding the predicted tracrRNA. Figure 3 B and 3C). Additionally, active CRISPR targeting was assessed by searching for spacer sequences that matched other genomic sequences assembled from the same or related metagenomics. Along these routes, phage genomes targeted by one of the spacers encoded in the SMARTICRISPR array were identified. Figure 3 C and 3D). Analysis of the regions adjacent to the target sequence showed that the 3'PAM sequence contains the GG motif (C and 3D). Figure 3 D). These results indicate that the SMARTICRISPR system is active in its natural environment as an RNA-directed effector involved in phage defense, and may act as a nuclease, i.e., cleaving or degrading targeted DNA or RNA.

[0170] SMART I effectors are active RNA-directed dsDNA CRISPR endonucleases.

[0171] Environmental RNA expression data from the SMART I MG34-1 and MG34-16 systems were used to engineer putative guide RNAs (sgRNAs). Figure 3 B and 3C, and Figure 9In addition, multiple sgRNAs predicted and designed from SMART I repeat sequences and tracrRNA were tested in vitro in a PAM enrichment assay. Figure 10 The assay confirmed programmable dsDNA cleavage of MG34-1, MG34-9, and MG34-16 designed with multiple sgRNAs. Figure 10 ). MG34-1 and MG34-9 require NGGN PAM for target identification and fragmentation. Figure 11 A and 11C). Analysis of the cleavage site indicated preferential cleavage at position 7 (A and 11C). Figure 11 These results indicate a novel biochemical cleavage mechanism compared to the Cas9 enzyme, which preferentially cleaves from position 3 of PAM, and provide further support for a new classification of the SMARTICRISPR system.

[0172] PAM enrichment assays without an end-repair step showed no activity against SMARTI nucleases. In the PAM enrichment protocol, end-repair is required before ligation to create blunt-ended fragments, instructing these enzymes to produce staggered double-strand DNA breaks.

[0173] Experiments in *E. coli* showed that the system possesses the activity necessary to act as a nuclease in the cell. *E. coli* strains expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid containing the sgRNA target. Successful targeting and cleavage of the antibiotic resistance plasmid in the presence of antibiotics resulted in growth defects. The assay showed approximately 2-fold growth inhibition compared to a control experiment using a kanamycin resistance plasmid without the sgRNA target. Figure 12 ).

[0174] SMART II endonuclease

[0175] Compared to SMARTII effectors, SMARTII effectors have a smaller size distribution (approximately 400-600 amino acids). Their genomic background indicates unusual repetitive regions or CRISPR arrays. Non-CRISPR repetitive regions contain directly repeating sequences, ranging in size from approximately 10 to over 30 bp. In some cases, these include multiple distinct repeat units. Sometimes, standard CRISPR identification algorithms will label these regions as CRISPR systems; however, closer inspection reveals that regions identified as spacer sequences recur repeatedly in the arrays. These arrays are not directly adjacent to effectors, but they reside within the same genomic region. Figure 3 A, MG35-236 and Figure 13 A, for example, effector genes >20kb). SMART II system operons typically lack adaptation module genes (e.g., genes involved in interval acquisition).

[0176] In addition to all six RuvC and HNH nuclease catalytic residues commonly found in type II Cas effectors ( Figure 6 In addition, structural prediction identified characteristic residues of the Cas enzyme involved in guide RNA binding, target cleavage, and PAM recognition and interaction. Furthermore, the SMART II effector contains multiple RRXRR and zinc-binding band motifs (CX). [2-4] C or CX [2-4] H), which may be involved in the recognition and binding of target nucleic acid motifs. Based on the location of important residues, the predicted domain architecture of the SMART II nuclease consists of three RuvC subdomains, an arginine-rich region containing the RRxRR motif (e.g., a domain homologous to PF14239), an HNH endonuclease domain, an unknown domain, and a recognition domain (REC). Figure 6 The domain architecture of the SMARTII effector differs from the known domain architecture of type II Cas9 nucleases. Figure 6 and Figure 14 ).

[0177] Environmental transcriptome data from some SMART II systems have confirmed in situ expression of CRISPR arrays and other repetitive regions in the natural environment. Figure 13 A). Transcription of the 5' untranslated region (UTR) of some SMARTII effectors was also observed in environmental expression data. Figure 13 B), thus indicating that this region may be important for the regulation of nuclease activity or the SMART system.

[0178] Preliminary in vitro experiments using SMART II effector proteins, repetitive regions, and related intergenic regions suggest that these enzymes may have the ability to cleave dsDNA in a programmable manner. Figure 15 The results indicate that SMARTII nuclease activity can be directed by RNA and / or DNA, which may require the use of repetitive regions such as CRISPR arrays, or by recognizing features encoded within a locus, such as TIR or 5'UTR.

[0179] Some SMART II effectors were observed to be located next to the putative insertion sequences (IS) encoding transposases TnpA and TnpB. Figure 3 A). The ends of the IS were identified as containing terminal inverted repeat sequences (TIRs) with predicted hairpin structures, and the target site repeats into which the IS is most likely to integrate were also identified. Furthermore, some SMARTII loci encode putative TIRs located flanking SMARTII effectors (e.g., Figure 3 ).

[0180] Example 2 - Identification / confirmation of the PAM sequence of the endonuclease described herein

[0181] The proposed SMART endonuclease was expressed in an *E. coli* lysate-based expression system (PURExpress, New England Biolabs). In this system, the *E. coli* endonuclease was codon-optimized and cloned into a vector containing a T7 promoter and a C-terminal His tag. The gene was amplified by PCR, with primer binding sites located 150 bp upstream of the T7 promoter and downstream of the terminator sequence, respectively. This PCR product was added to NEBPURExpress at a final concentration of 5 nM and expressed at 37°C for 2 hours to produce the endonuclease for PAM assays.

[0182] Proposed sgRNAs compatible with each SMART Cas enzyme described herein were identified from RNAseq reads assembled from sequencing data to contiguous CRISPR loci: the secondary structure of the tracr region was determined from RNAseq data and repetitive sequences of CRISPR arrays in the Geneious software package (https: / / www.geneious.com), and the resulting helices were trimmed and cascaded with GAAA quadruplexes. Multiple lengths of repeat-antirepetitive helical trimming, as well as different spacer lengths and different tracr termination points, were tested. Figure 12 (Showing SEQ ID NO: 612-615). Each sgRNA was then assembled by assembly-type PCR, purified with SPRI beads, and in vitro transcribed (IVT) was performed according to the manufacturer's recommended short RNA transcript protocol (HiScribe T7 kit, NEB). RNA transcription reactions were washed with a Monarch RNA kit, and purity was checked using Tapestation (Agilent).

[0183] The PAM sequence is determined by sequencing plasmids containing randomly generated potential PAM sequences, which can be cleaved by the putative nuclease. In this system, under the control of the T7 promoter, an E. codon-optimized nucleotide sequence encoding the putative nuclease is transcribed and translated from a PCR fragment in vitro. A second PCR fragment with a minimal CRISPR array, consisting of the T7 promoter and a repeat-spacer-repeat sequence, is transcribed in the same reaction. Successful expression of the endonuclease and repeat-spacer-repeat sequence in the TXTL system, followed by CRISPR array processing, provides an active in vitro CRISPR nuclease complex.

[0184] The target plasmid library contains a spacer sequence (potential PAM sequence) that matches the minimal array of 8N mixed degenerate bases preceding it. This target plasmid library is incubated with the output of a TXTL reaction (10 mM Tris pH 7.5, 100 mM NaCl, and 10 mM MgCl2, containing a 5-fold dilution of the translated Cas enzyme, 5 nM of the 8N PAM plasmid library, and 50 nM of sgRNA targeting the PAM library). After 1–3 hours, the reaction is stopped, and the DNA is recovered using a DNA cleanup kit. The aptamer sequence is blunt-ended and ligated to DNA containing the active PAM sequence that has been cleaved by a nuclease, while uncleavage is not ligated. The DNA segment containing the active PAM sequence is then amplified by PCR using primers specific to the library and aptamer sequences. The PCR amplification products are cleaved on a gel to identify the amplicons corresponding to the cleavage events. The amplified segments from the cleavage reaction are also used as templates for preparing NGS libraries or as substrates for Sanger sequencing. Sequencing of the resulting library (a subset of the starting 8N library) revealed PAM-active sequences compatible with the CRISPR complex. For PAM assays using the processed RNA construct, the same procedure was repeated, except that in vitro transcribed RNA was added along with the plasmid library, and the minimal CRISPR array / tracr template was omitted. The following spacer sequence was used as a target in these assays (5'-CGUGAGCCACCACGUCGCAAGCCUCGAC-3').

[0185] After obtaining raw sequencing reads from PAM assays, reads were filtered by a Phred quality score >20. A 24 bp sequence representing a known DNA sequence from the backbone adjacent to the PAM was used as a reference for finding proximal regions of the PAM, and the adjacent 8 bp were identified as putative PAMs. The distance between the PAM and the ligated aptamer was also measured for each read. Reads that did not precisely match the reference or aptamer sequence were excluded. PAM sequences were filtered by cleavage site frequency so that only PAMs with the most frequent cleavage sites ±2 bp were included in the analysis. The filtered list of PAMs was used to generate sequence logos using Logomaker (Tareen A, Kinney JB. Logomaker: beautiful sequence logos in Python. Bioinformatics. 2020; 36(7):2272-2274, which is incorporated herein by reference).

[0186] Example 3 - The predicted RNA folding scheme

[0187] The method from Andronescu 2007 was used to calculate the predicted RNA folding of an active single RNA sequence at 37°C. The color of a base corresponds to the probability of that base pairing, with red representing a high probability and blue representing a low probability.

[0188] Example 4 - In vitro lysis efficiency

[0189] In protease-deficient *E. coli* strain B, a nuclease was expressed from an inducible T7 promoter as a His-tagged fusion protein. The nuclease was fused with two nuclear localization signals (N-terminal NLS nucleoplasmic protein bityping and C-terminal simian virus 40T-antigen NLS PPKKKRK), a maltose-binding protein (MBP) tag, a tobacco etch virus (TEV) protease cleavage site, and a 6XHis tag, in the following order from N-terminus to C-terminus: 6XHis-MBP-TEV-NLS-gene-NLS-terminus. This protein was expressed in *E. coli* NEB Iq under a pTac promoter via self-inducible medium (MagicMedia ThermoFisher), grown at 30°C, and induced at 16°C.

[0190] Cells expressing His-tagged proteins were dissolved by sonication and purified by Ni-NTA affinity chromatography on a HisTrap FF column (GE Lifescience) on an AKTA Avant FPLC (GE Lifescience). The eluent was separated by SDS-PAGE on an acrylamide gel (Bio-Rad) and stained with Instant Blue Ultrafast Coomassie (Sigma-Aldrich). Protein band density was determined using ImageLab software (Bio-Rad) to confirm purity. The purified endonuclease was dialyzed into a storage buffer (pH 7.5) consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, and 5% glycerol and stored at -80°C.

[0191] Target DNA containing a spacer sequence and a PAM sequence (e.g., as determined in Example 2) was constructed via DNA synthesis. A single representative PAM was selected for testing when it had degenerate bases. The target DNA consisted of 2200 bp of linear DNA obtained from a plasmid via PCR amplification, with the PAM and spacer located 700 bp from one end. Successful lysis yielded 700 and 1500 bp fragments. The target DNA, in vitro transcribed single RNA, and purified recombinant protein were combined with excess protein and RNA in lysis buffer (10 mM Tris, 100 mM NaCl, 10 mM MgCl2) and incubated for 5 minutes to 3 hours, typically 1 hour. The reaction was stopped by adding RNase A and incubating for 60 minutes. The reaction was then cleaved on a 1.2% TAE agarose gel, and the fraction of lysed target DNA was quantified in ImageLab software.

[0192] Example 5. - Activity in Escherichia coli

[0193] Escherichia coli lacks the ability to effectively repair double-strand DNA breaks. Therefore, cleavage of genomic DNA can be a lethal event. Utilizing this phenomenon, endonuclease activity in E. coli was tested by recombinantly expressing endonucleases and guide RNA with spacers / targets and PAM sequences integrated into its genomic DNA in target strains.

[0194] To test nuclease activity in bacterial cells, BL21(DE3) strain (NEB) was transformed with plasmids containing a T7-driven effector and sgRNA (10 ng per plasmid), plated, and grown overnight. The resulting colonies were cultured in triplicate overnight, then subcultured in SOB and grown to OD 0.4–0.6. Cell cultures were chemically activated to 0.5 OD equivalents according to a standard kit protocol (Zymo Mix and Go kit) and transformed with 130 ng kanamycin plasmid (with or without spacers and PAM in the backbone). After heat shock, transformation was restored in SOC at 37°C for 1 h, and nuclease efficiency was determined by a 5-fold series of dilutions grown on induction medium (LB agar plates containing antibiotics and 0.05 mM IPTG). Colonies were quantified from the dilution series to measure overall inhibition due to nuclease-driven plasmid cleavage.

[0195] The results of this type of measurement are shown in Figure 12 In. Figure 12Figure (A) shows the plating of *E. coli* strains, demonstrating plasmid cleavage; *E. coli* expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid containing the sgRNA target (+sp). The plate quadrants showing growth defects (+sp) compared to the negative control (no target and PAM(-sp)) indicate successful targeting and cleavage via the enzyme. Experiments were repeated twice and performed in triplicate. Figure 12 In Figure B, a measurement of colony-forming units (CFU) from the replication plating experiment in A is shown, revealing growth inhibition under target conditions (+sp) compared to the non-target control (-sp), indicating that the plasmid was cleaved.

[0196] Engineered strains with PAM sequences (e.g., as determined in Example 2) integrated into their genomic DNA were transformed with DNA encoding a nuclease. The transformants were then chemically activated and transformed with 50 ng of guide RNA (e.g., crRNA) that was specific to the target sequence (“targeted”) or non-specific to the target (“non-target”). After heat shock, transformation was recovered at 37°C in SOC for 2 hours. Nuclease efficiency was then determined by a 5-fold dilution series grown on induction medium. Colonies were quantified in triplicate from the dilution series.

[0197] Example 6. - Testing the genome cleavage activity of the MG CRISPR complex in mammalian cells.

[0198] To demonstrate targeting and cleavage activity in mammalian cells, MG Cas effector protein sequences were tested in two mammalian expression vectors: (a) a vector with a C-terminal SV40 NLS and a 2A-GFP tag, and (b) a vector without a GFP tag but with two SV40 NLS sequences, one at the N-terminus and one at the C-terminus. NLS sequences include any NLS sequences described herein. In some cases, the nucleotide sequences encoding endonucleases are codon-optimized for expression in mammalian cells.

[0199] The corresponding crRNA sequence with the attached target sequence was cloned into a second mammalian expression vector. Both plasmids were co-transfected into HEK293T cells. 72 hours after co-transfection of the expression plasmid and gRNA targeting plasmid into HEK293T cells, DNA was extracted and used to prepare an NGS library. The percentage of NHEJ was measured by insertions and deletions in target site sequencing to demonstrate the enzyme's targeting efficiency in mammalian cells. At least 10 different target sites were selected to test the activity of each protein.

[0200] Example 7 - Predictive activity of the MG family described herein

[0201] In situ expression and protein sequence analysis indicated that these enzymes are active nucleases. They contain the predicted endonuclease-associated domains (matching the RRXRR and HNH_ endonuclease Pfam domains); Figure 2 , 3 A and 3B), and contain the predicted HNH and RuvC catalytic residues (e.g., Figure 2 , 3 A and 3B, rectangles). Furthermore, the presence of the RRXRR motif found in the H-like ribonuclease family indicates potential RNA targeting and / or nuclease activity (see [link]). Figure 2 ).

[0202] Expression data confirmed the in situ native activity of the candidate MG34-1 nuclease, tracrRNA, and CRISPR array. Figure 4 ).

[0203] Example 8. - Activity in mammalian cells with mRNA delivery

[0204] For genome editing using cell transfection / mRNA conversion, mouse or human codon optimization of the coding sequence was performed using algorithms from Twist Bioscience or Thermo Fisher Scientific (GeneArt). A cassette was constructed using two nuclear localization signals attached to the coding endonuclease sequence: SV40 at the N-terminus and a nucleoplasmic protein at the C-terminus. Additionally, the untranslated region from human complement 3 (C3) was attached to the 5' and 3' of the coding sequence within the cassette.

[0205] This box was then cloned into an mRNA production vector upstream of a long poly A segment. The mRNA construct was organized as follows: 5' UTR from C3 - SV40 NLS - codon-optimized SMART gene - nucleoplasmic protein NLS - 3' UTR from C3 - 107 poly A tail. The mRNA was then transcribed via the T7 promoter using engineered T7 RNA polymerase (Hi-T7: New England Biolabs). 5' capping of the mRNA was performed co-transcribed using CleanCap AG (Trilink Biolabs). The mRNA was then purified using the MEGAclear Transcription Clean-Up kit (Thermo Fisher Scientific).

[0206] Mammalian cells were co-transfected with transcribed mRNA and a set of at least 10 guides targeting specific genomic regions using a Lipofectamine Messenger Max (Thermo Fisher Scientific). Cells were incubated for a period of time (e.g., 48 hours), and then genomic DNA was isolated using the Purelink Genomic DNA Extraction Kit (Fisher Scientific). Target regions were amplified using specific primers. Edits were then evaluated using Sanger sequencing via inference of CRISPR Edits and NGS for a comprehensive analysis of the results.

[0207] While preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited to the specific embodiments provided herein. Although the invention has been described with reference to the foregoing detailed description, the description and illustrative examples of embodiments herein are not intended to be construed as limiting. Various variations, modifications, and alternatives will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein can be used to implement the invention. Therefore, it is contemplated that the invention should also cover any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the invention and are thus intended to cover methods and structures within the scope of these claims and their equivalents.

Claims

1. An engineered nuclease system, comprising: (a) The endonuclease of SEQ ID NO: 2; and (b) An engineered guide RNA (RNA) structure configured to form a complex with the said endonuclease, the engineered guide RNA structure comprising: (i) a guide RNA sequence configured to hybridize with a target deoxyribonucleic acid (DNA) sequence; and (ii) An RNA sequence configured to bind to the endonuclease, wherein the RNA sequence comprises a sequence of non-variable nucleotides of any one of SEQ ID NO:203 or 613, wherein the sequence of non-variable nucleotides is a scaffold sequence of any one of SEQ ID NO:203 or 613.

2. The engineered nuclease system of claim 1, wherein the engineered nuclease system further comprises one or more nuclear localization sequences (NLS) located near the N-terminus of the endonuclease.

3. The engineered nuclease system of claim 1 further comprises a single-stranded or double-stranded DNA repair template, wherein the single-stranded or double-stranded DNA repair template comprises, from 5' to 3': a first homologous arm comprising a sequence having at least 20 nucleotides at 5' of the target DNA sequence; a synthetic DNA sequence having at least 10 nucleotides; and a second homologous arm comprising a sequence having at least 20 nucleotides at 3' of the target DNA sequence.

4. The engineered nuclease system of claim 3, wherein the first homologous arm comprises a sequence having at least 40 nucleotides.

5. The engineered nuclease system of claim 1, wherein the engineered guide RNA structure comprises a single RNA polynucleotide, the RNA polynucleotide comprising the guide RNA sequence and the RNA sequence configured to bind to the nuclease.

6. The engineered nuclease system of claim 1, wherein the guide RNA sequence is complementary to a eukaryotic genome sequence.

7. The engineered nuclease system of claim 1, wherein the guide RNA sequence is complementary to the fungal genome sequence.

8. The engineered nuclease system of claim 1, wherein the guide RNA sequence is complementary to the plant genome sequence.

9. The engineered nuclease system of claim 1, wherein the guide RNA sequence is complementary to the mammalian genome sequence.

10. The engineered nuclease system of claim 1, wherein the guide RNA sequence is complementary to the human genome sequence.

11. The engineered nuclease system of claim 1, wherein the guide RNA sequence is 15-24 nucleotides in length.

12. The engineered nuclease system of claim 1, wherein the engineered nuclease system further comprises one or more NLSs, the NLSs being located near the C-terminus of the endonuclease.

13. The engineered nuclease system of claim 3, wherein the second homologous arm comprises a sequence having at least 40 nucleotides.

14. The engineered nuclease system of claim 1, wherein the engineered guide RNA structure comprises an RNA sequence containing a hairpin, the hairpin comprising a stem and a loop, wherein the stem comprises at least 12 pairs of ribonucleotides.

15. The engineered nuclease system of claim 14, wherein the engineered guide RNA structure further comprises a second stem and a second loop, wherein the second stem comprises at least 5 pairs of ribonucleotides.

16. The engineered nuclease system of claim 14, wherein the engineered guide RNA structure further comprises an RNA structure containing at least two hairpins.

17. The engineered nuclease system of claim 1, wherein the RNA sequence configured to bind to the nuclease comprises nucleotides 23-157 of SEQ ID NO: 203 or nucleotides 23-145 of SEQ ID NO: 613.

Citation Information

Patent Citations

  • Poly-substituted-phenyl-oligoribo nucleotides having enhanced stability and membrane permeability and methods of use

    US5858988A

  • Antiviral anticancer poly-substituted phenyl derivatized oligoribonucleotides and methods for their use

    US6291438B1

  • Base editing enzyme

    CN116867897A