Novel CRISPR DNA Targeting Enzymes and Systems

The engineered CRISPR-Cas system CLUST.091979 addresses the limitations of current CRISPR-Cas systems by providing programmable effectors for nucleic acid modification in diverse organisms, enhancing editing capabilities and delivery strategies.

JP7705382B2Active Publication Date: 2025-07-09ARBOR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022515511
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-09
Filing Date
2020-09-09
Publication Date
2025-07-09
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

Current CRISPR-Cas systems lack additional programmable effectors with unique PAM sequence requirements for nucleic acid modification, limiting their versatility and applicability in non-native environments such as bacteria and eukaryotic cells.

Method used

Development of non-naturally engineered CRISPR-Cas systems, specifically CLUST.091979, comprising engineered CRISPR-associated proteins with specific amino acid sequences and RNA guides, capable of binding and modifying target nucleic acids in various organisms, including eukaryotic cells.

Benefits of technology

Enables efficient and programmable nucleic acid editing and regulation in diverse cellular environments, offering novel editing properties, smaller size for versatile delivery, and programmed perturbations, expanding the toolbox of genome and epigenome manipulation techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705382000045
    Figure 0007705382000045
  • Figure 0007705382000046
    Figure 0007705382000046
  • Figure 0007705382000047
    Figure 0007705382000047
Patent Text Reader

Abstract

This disclosure describes novel systems, methods, and compositions for manipulating nucleic acids in a targeted manner. This disclosure describes non-naturally occurring engineered CRISPR systems, components, and methods for targeting and modifying nucleic acids. Each system includes one or more protein components and one or more nucleic acid components that work together to target nucleic acids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 897,859, filed on September 9, 2019, the entire content of which is incorporated herein by reference.

[0002] Sequence Listing This application is electronically filed in ASCII format and includes a sequence listing which is incorporated herein by reference in its entirety. The ASCII copy created on September 9, 2020, has the file name A2186-7028WO_SL.txt and a size of 475,511 bytes.

[0003] The present disclosure relates to systems and methods for genome editing and regulation of gene expression using novel, clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes.

Background Art

[0004] In recent years, advances in genome sequencing technology and analysis have provided important insights into the genetic basis of biological activity in a wide variety of natural areas, from prokaryotic biosynthetic pathways to human pathology. To fully understand and evaluate the vast amount of information obtained, there is a need to improve the scale, effectiveness, and ease of sequence technologies for corresponding genomic and epigenomic manipulation. Such new technologies will accelerate the development of new applications in numerous areas, including biotechnology, agriculture, and human therapeutics.

[0005] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes are collectively known as the CRISPR-Cas or CRISPR / Cas system, which is an adaptive immune system in archaea and bacteria that defends specific species from foreign genetic elements. The CRISPR-Cas system includes a highly diverse group of protein effectors, non-coding elements, and locus architectures, and some of its examples have been engineered and adapted to generate important biotechnology advancements.

[0006] Components of this system involved in host defense include one or more effector proteins with the ability to modify nucleic acids and an RNA guide element responsible for targeting the effector protein to specific sequences on phage nucleic acids. The RNA guide is composed of CRISPR RNA (crRNA) and may require an additional trans-activating RNA (tracrRNA) to enable the manipulation of target nucleic acids by one or more effector proteins. The crRNA consists of a direct repeat that is responsible for protein binding to the crRNA and a spacer sequence complementary to the desired nucleic acid target sequence. The CRISPR system can be reprogrammed to target another DNA or RNA target by modifying the spacer sequence of the crRNA.

[0007] The CRISPR-Cas system can be broadly divided into two classes: Class 1 systems are composed of multiple effector proteins that together form a complex around the crRNA, and Class 2 systems consist of a single effector protein that complexes with an RNA guide to target a nucleic acid substrate. The single-subunit effector composition of Class 2 systems provides a more convenient set of components for engineering and application transfer, and has thus far been an important source of programmable effectors. Nevertheless, there is still a need for additional programmable effectors and systems that go beyond current CRISPR-Cas systems, such as smaller effectors and / or effectors with unique PAM sequence requirements that enable new applications by virtue of their unique properties for modifying nucleic acids and polynucleotides (i.e., DNA, RNA, or any hybrid, derivative, or modified form thereof). SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0008] This disclosure provides non-naturally engineered systems and compositions for novel single-effector Class 2 CRISPR-Cas systems that were first computationally identified from genomic databases and then engineered and experimentally validated. In particular, the identification of the components of these CRISPR-Cas systems enables their use in non-native environments, such as bacteria other than those in which the system was first discovered or eukaryotic cells such as mammalian cells. These novel effectors differ in sequence and function compared to orthologs and homologs of existing Class 2 CRISPR effectors.

[0009] In one aspect, the present disclosure provides an engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-Cas system of CLUST.091979, comprising a CRISPR-associated protein comprising an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid, wherein the CRISPR-associated protein can bind to the RNA guide and modify a target nucleic acid sequence complementary to the spacer sequence. In one aspect, the present disclosure provides a CRISPR-associated protein or a nucleic acid encoding the CRISPR-associated protein, comprising an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid, or a nucleic acid encoding the RNA guide, of an engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-Cas system of CLUST.091979, wherein the CRISPR-associated protein can bind to the RNA guide and modify a target nucleic acid sequence complementary to the spacer sequence.

[0010] In some embodiments, the present disclosure provides an engineered, non-naturally occurring, clustered regularly interspaced short palindromic repeat (CRISPR)-Cas system of CLUST.091979, comprising a CRISPR-associated protein or a nucleic acid encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises the amino acid sequence of SEQ ID NO: 241; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid, wherein the CRISPR-associated protein can bind to the RNA guide and modify a target nucleic acid sequence complementary to the spacer sequence. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4, SEQ ID NO: 10, SEQ ID NO: 12, or SEQ ID NO: 14.

[0011] In some embodiments of any part of the systems described herein, the CRISPR-associated protein comprises at least one (e.g., one, two, or three) RuvC domain or at least one split RuvC domain.

[0012] In some embodiments of any part of the systems described herein, the CRISPR-related protein comprises one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216), where X1 is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M; (b) RX1X2X3L (SEQ ID NO: 217), where X1 is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K; (c) NX1YX2 (SEQ ID NO: 218), where X1 is I or L or F and X2 is K or R or V or E; (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219), where X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N; (e) LX1NX2 (SEQ ID NO: 220), where X1 is G or S or C or T and X2 is N or Y or K or S; (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221), where X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I; (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222), where X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P's Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q; and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223), where X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L).In some embodiments of any part of the system described in this specification, the sequence of SEQ ID NO: 216 is an N-terminal sequence. In some embodiments of any part of the system described in this specification, the sequence of SEQ ID NO: 219 is a C-terminal sequence. In some embodiments of any part of the system described in this specification, the sequence of SEQ ID NO: 220 is a C-terminal sequence. In some embodiments of any part of the system described in this specification, the sequence of SEQ ID NO: 221 is a C-terminal sequence. In some embodiments of any part of the system described in this specification, the sequence of SEQ ID NO: 222 is a C-terminal sequence. In some embodiments of any part of the system described in this specification, the sequence of SEQ ID NO: 223 is a C-terminal sequence.

[0013] In some embodiments of any part of the systems described herein, the CRISPR-related protein comprises one or more of the following sequences: (a) ECPITKDVINEYK (SEQ ID NO: 290); (b) NLTSITIG (SEQ ID NO: 231); (c) NYRTKIRTLN (SEQ ID NO: 232); (d) ISYIENVEN (SEQ ID NO: 233); (e) ELLSVEQLK (SEQ ID NO: 234); (f) HINSMTINIQDFKIE (SEQ ID NO: 235); (g) KENSLGFIL (SEQ ID NO: 236); (h) GNRQIKKG (SEQ ID NO: 237); (i) DVNFKHA (SEQ ID NO: 238); (j) GYINLYKYLLEH (SEQ ID NO: 239); (k) KEQVLSKLLY (SEQ ID NO: 240); (l) EYIYVSCVNKLRAKYVSYFILKEKYYEKQKEYDIEMGF (SEQ ID NO: 241); (m) DDSTESKESMDKRR (SEQ ID NO: 242); (n) NVQQDINGCLKNIINY (SEQ ID NO: 243); (o) ALENLENSNFEK (SEQ ID NO: 244); (p) QVLPTIKSLL (SEQ ID NO: 245); (q) YHKLENQN (SEQ ID NO: 246); (r) ASDKVKEYIE (SEQ ID NO: 247); (s) TNENNEIVDAKYT (SEQ ID NO: 248); (t) ANFFNLMMKSLHFAS (SEQ ID NO: 249); (u) LLSNNGKTQIALVPSE (SEQ ID NO: 250); (v) HINGLNADFNAANNIKYI (SEQ ID NO: 251), or a sequence having one, two, or three or fewer sequence differences (e.g., substitutions) relative to any of the foregoing. In some embodiments, the CRISPR-related protein has a sequence that is at least 70% identical to SEQ ID NO: 4. In some embodiments, the CRISPR-related protein has a sequence that is at least 70% identical to SEQ ID NO: 10.

[0014] In some embodiments of any part of the systems described herein, the directory repeat array comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, SEQ ID NOs: 118 to 151, or SEQ ID NO: 213. In some embodiments of any part of the systems described herein, the directory repeat array comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, SEQ ID NOs: 118 to 151, or SEQ ID NO: 213.

[0015] In some embodiments of any part of the systems described herein, the directory repeat array is one or more of the following arrays: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (where X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (where X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 227)); and (c) X1X2X3AC (SEQ ID NO: 228) (where X1 is A or C or G, X2 is C or A, X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any part of the systems described herein, SEQ ID NO: 224 is proximal to the 5' end of the directory repeat. In some embodiments of any part of the systems described herein, SEQ ID NO: 228 is proximal to the 3' end of the directory repeat.

[0016] In some embodiments of any part of the systems described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM), where the PAM sequence comprises a nucleic acid sequence, which is described as a nucleic acid sequence of 5'-NTTN-3', 5'-NTTR-3', 5'-RTTR-3', 5'-TNNT-3', 5'-TNRT-3', 5'-TSRT-3', 5'-TGRT-3', 5'-TNRY-3', 5'-TTNR-3', 5'-TTYR-3', 5'-TTTR-3', 5'-TTCV-3', 5'-DTYR-3', 5'-WTTR-3', 5'-NNR-3', 5'-NYR-3', 5'-YYR-3', 5'-TYR-3', 5'-TTN-3', 5'-TTR-3', 5'-CNT-3', 5'-NGG-3', 5'-BGG-3', or 5'-R-3', where "N" is any nucleotide, "B" is C or G or T, "D" is A or G or T, "R" is A or G, "S" is G or C, "V" is A or C or G, "W" is A or T, and "Y" is C or T.

[0017] In some embodiments of any part of the systems described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any part of the systems described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any part of the systems described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', "N" is any nucleotide, and "R" is A or G.

[0018] In some embodiments of any part of the systems described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any part of the systems described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any part of the systems described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G.

[0019] In some embodiments of any part of the systems described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any part of the systems described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any part of the systems described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), "N" is any nucleotide, and "R" is A or G.

[0020] In some embodiments of any part of the systems described herein, the spacer sequence of the RNA guide comprises from about 15 nucleotides to about 55 nucleotides. In some embodiments of any part of the systems described herein, the spacer sequence of the RNA guide comprises 20 to 45 nucleotides.

[0021] In some embodiments of any part of the systems described herein, the CRISPR-related protein comprises catalytic residues (e.g., aspartic acid or glutamic acid). In some embodiments of any part of the systems described herein, the CRISPR-related protein cleaves a target nucleic acid. In some embodiments of any part of the systems described herein, the CRISPR-related protein further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcription modification factor, a light-gated control factor, a chemically inducible factor, or a chromatin visualization factor.

[0022] In some embodiments of any part of the systems described herein, the nucleic acid encoding the CRISPR-related protein is codon-optimized for expression in a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell). In some embodiments of any part of the systems described herein, the nucleic acid encoding the CRISPR-related protein is operably linked to a promoter. In some embodiments of any part of the systems described herein, the nucleic acid encoding the CRISPR-related protein is within a vector. In some embodiments, the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.

[0023] In some embodiments of any part of the systems described herein, the target nucleic acid is a DNA molecule. In some embodiments of any part of the systems described herein, the target nucleic acid comprises a PAM sequence.

[0024] In some embodiments of any part of the systems described herein, the CRISPR-related protein has non-specific nuclease activity.

[0025] In some embodiments of any part of the systems described herein, modification of a target nucleic acid occurs by recognition of the target nucleic acid by a CRISPR-related protein and an RNA guide. In some embodiments of any part of the systems described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any part of the systems described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any part of the systems described herein, an insertion event occurs due to the modification of the target nucleic acid. In some embodiments of any part of the systems described herein, a deletion event occurs due to the modification of the target nucleic acid. In some embodiments of any part of the systems described herein, cytotoxicity or cell death occurs due to the modification of the target nucleic acid.

[0026] In some embodiments of any part of the systems described herein, the system further comprises a donor template nucleic acid. In some embodiments of any part of the systems described herein, the donor template nucleic acid is a DNA molecule. In some embodiments of any part of the systems described herein, the donor template nucleic acid is an RNA molecule.

[0027] In some embodiments of any part of the systems described herein, the RNA guide optionally comprises a tracrRNA and / or a modulator RNA. In some embodiments of any part of the systems described herein, the system further comprises a tracrRNA. In some embodiments of any part of the systems described herein, the system does not comprise a tracrRNA. In some embodiments of any part of the systems described herein, the CRISPR-related protein is self-processing. In some embodiments of any part of the systems described herein, the system further comprises a modulator RNA.

[0028] In some embodiments of any part of the systems described herein, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 1, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 152, SEQ ID NO: 153, or SEQ ID NO: 154.

[0029] In some embodiments of any part of the systems described herein, the system is present in a delivery composition comprising nanoparticles, liposomes, exosomes, microvesicles, or a gene gun.

[0030] In some embodiments of any part of the systems described herein, the system is intracellular. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a prokaryotic cell.

[0031] In another aspect, the present disclosure provides a cell, wherein the cell comprises a CRISPR-related protein having an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid. In another aspect, the present disclosure provides a cell, wherein the cell comprises a CRISPR-related protein or a nucleic acid encoding a CRISPR-related protein, wherein the CRISPR-related protein has an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid, or a nucleic acid encoding the RNA guide.

[0032] In some embodiments of any of the cells described herein, the CRISPR-related protein comprises at least one (e.g., one, two, or three) RuvC domain or at least one split RuvC domain.

[0033] In some embodiments of any of the cells described herein, the CRISPR-related protein is one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216) (where X1 is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M); (b) RX1X2X3L (SEQ ID NO: 217) (where X1 is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (c) NX1YX2 (SEQ ID NO: 218) (where X1 is I or L or F, and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (where X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (e) LX1NX2 (SEQ ID NO: 220) (where X1 is G or S or C or T, and X2 is N or Y or K or S); (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (where X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (where X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P's Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q); and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (where X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L).In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 216 is the N-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 219 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 220 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 221 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 222 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 223 is the C-terminal sequence.

[0034] In some embodiments of any of the cells described herein, the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. In some embodiments of any of the cells described herein, the direct repeat sequence comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213.

[0035] In some embodiments of any of the cells described herein, the direct repeat sequence is one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (where X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, X8 is A or G or T) (for example, ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (where X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, X9 is G or A or C) (for example, TTTTWTARG (SEQ ID NO: 227)); and (c) X1X2X3AC (SEQ ID NO: 228) (where X1 is A or C or G, X2 is C or A, X3 is A or C) (for example, ACAAC (SEQ ID NO: 229)). In some embodiments of any of the cells described herein, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments of any of the cells described herein, SEQ ID NO: 228 is proximal to the 3' end of the direct repeat.

[0036] In some embodiments of any of the cells described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the cells described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the cells described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', "N" is any nucleotide, and "R" is A or G.

[0037] In some embodiments of any of the cells described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the cells described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the cells described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', "N" is any nucleotide, and "R" is A or G.

[0038] In some embodiments of any of the cells described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the cells described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the cells described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), "N" is any nucleotide, and "R" is A or G.

[0039] In some embodiments of any of the cells described herein, the spacer sequence comprises from about 15 nucleotides to about 55 nucleotides. In some embodiments of any of the cells described herein, the spacer sequence comprises 20 to 45 nucleotides.

[0040] In some embodiments of any of the cells described herein, the CRISPR-associated protein comprises catalytic residues (e.g., aspartic acid or glutamic acid). In some embodiments of any of the cells described herein, the CRISPR-associated protein cleaves a target nucleic acid. In some embodiments of any of the cells described herein, the CRISPR-associated protein further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcriptional modification factor, a light-gated control factor, a chemically inducible factor, or a chromatin visualization factor.

[0041] In some embodiments of any of the cells described herein, the nucleic acid encoding the CRISPR-associated protein is codon-optimized for expression in a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell). In some embodiments of any of the cells described herein, the nucleic acid encoding the CRISPR-associated protein is operably linked to a promoter. In some embodiments of any of the cells described herein, the nucleic acid encoding the CRISPR-associated protein is within a vector. In some embodiments, the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.

[0042] In some embodiments of any of the cells described herein, the RNA guide optionally comprises a tracrRNA and / or a modulator RNA. In some embodiments of any of the cells described herein, the cell further comprises a tracrRNA. In some embodiments of any of the cells described herein, the cell does not comprise a tracrRNA. In some embodiments of any of the cells described herein, the CRISPR-associated protein is self-processing. In some embodiments of any of the cells described herein, the cell further comprises a modulator RNA.

[0043] In some embodiments of any of the cells described herein, the cell is a eukaryotic cell. In some embodiments of any of the cells described herein, the cell is a mammalian cell. In some embodiments of any of the cells described herein, the cell is a human cell. In some embodiments of any of the cells described herein, the cell is a prokaryotic cell.

[0044] In some embodiments of any of the cells described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the cells described herein, the target nucleic acid comprises a PAM sequence.

[0045] In some embodiments of any of the cells described herein, the CRISPR-related protein has non-specific nuclease activity.

[0046] In some embodiments of any of the cells described herein, recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide results in modification of the target nucleic acid. In some embodiments of any of the cells described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the cells described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the cells described herein, an insertion event occurs due to the modification of the target nucleic acid. In some embodiments of any of the cells described herein, a deletion event occurs due to the modification of the target nucleic acid. In some embodiments of any of the cells described herein, cytotoxicity or cell death occurs due to the modification of the target nucleic acid.

[0047] In another aspect, the disclosure provides a method of binding a system described herein to a target nucleic acid in a cell, comprising: (a) providing the system; and (b) delivering the system to the cell, wherein the cell comprises the target nucleic acid, the CRISPR-related protein binds to the RNA guide, and the spacer sequence binds to the target nucleic acid. In some embodiments, the cell is a eukaryotic cell, such as a mammalian cell, such as a human cell.

[0048] In another aspect, the present disclosure provides a method for modifying a target nucleic acid, comprising delivering an engineered non - naturally occurring CRISPR - Cas system to the target nucleic acid, the system comprising a CRISPR - associated protein having an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1 - 56; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to the target nucleic acid, wherein the CRISPR - associated protein has the ability to bind to the RNA guide, and recognition of the target nucleic acid by the CRISPR - associated protein and the RNA guide results in modification of the target nucleic acid. In another aspect, the present disclosure provides a method for modifying a target nucleic acid, comprising delivering an engineered non - naturally occurring CRISPR - Cas system to the target nucleic acid, the system comprising a CRISPR - associated protein or a nucleic acid encoding the CRISPR - associated protein, wherein the CRISPR - associated protein has an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1 - 56; and an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to the target nucleic acid, wherein the CRISPR - associated protein has the ability to bind to the RNA guide, and recognition of the target nucleic acid by the CRISPR - associated protein and the RNA guide results in modification of the target nucleic acid.

[0049] In some embodiments of any of the methods described herein, the CRISPR-related protein comprises one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216), where X1 is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M; (b) RX1X2X3L (SEQ ID NO: 217), where X1 is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K; (c) NX1YX2 (SEQ ID NO: 218), where X1 is I or L or F and X2 is K or R or V or E; (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219), where X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N; (e) LX1NX2 (SEQ ID NO: 220), where X1 is G or S or C or T and X2 is N or Y or K or S; (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221), where X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I; (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222), where X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P's Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q; and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223), where X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L.In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 216 is an N-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 219 is a C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 220 is a C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 221 is a C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 222 is a C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 223 is a C-terminal sequence.

[0050] In some embodiments of any of the methods described herein, the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. In some embodiments of any of the methods described herein, the direct repeat sequence comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213.

[0051] In some embodiments of any of the methods described herein, the directory repeat array is one or more of the following arrays: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (where X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (where X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 227)); and (c) X1X2X3AC (SEQ ID NO: 228) (where X1 is A or C or G, X2 is C or A, X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any of the methods described herein, SEQ ID NO: 224 is proximal to the 5' end of the directory repeat. In some embodiments of any of the methods described herein, SEQ ID NO: 228 is proximal to the 3' end of the directory repeat.

[0052] In some embodiments of any of the methods described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the methods described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the methods described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', "N" is any nucleotide, and "R" is A or G.

[0053] In some embodiments of any of the methods described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G.

[0054] In some embodiments of any of the methods described herein, the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods described herein, the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods described herein, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), "N" is any nucleotide, and "R" is A or G.

[0055] In some embodiments of any of the methods described herein, the spacer sequence comprises from about 15 nucleotides to about 55 nucleotides. In some embodiments of any of the methods described herein, the spacer sequence comprises 20 to 45 nucleotides.

[0056] In some embodiments of any of the methods described herein, the RNA guide optionally comprises a tracrRNA and / or a modulator RNA. In some embodiments of any of the methods described herein, the system further comprises a tracrRNA. In some embodiments of any of the methods described herein, the system does not comprise a tracrRNA. In some embodiments of any of the methods described herein, the CRISPR-associated protein is self-processing. In some embodiments of any of the methods described herein, the system further comprises a modulator RNA.

[0057] In some embodiments of any of the methods described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the methods described herein, the target nucleic acid comprises a PAM sequence.

[0058] In some embodiments of any of the methods described herein, the CRISPR-associated protein has non-specific nuclease activity.

[0059] In some embodiments of any of the methods described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid results in an insertion event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid results in a deletion event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid results in cytotoxicity or cell death.

[0060] In another aspect, the present disclosure provides a method for editing a target nucleic acid, the method comprising contacting the system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for modifying the expression of a target nucleic acid, the method comprising contacting the system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for targeting the insertion of a payload nucleic acid at a site of a target nucleic acid, the method comprising contacting the system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for targeting the excision of a payload nucleic acid from a site of a target nucleic acid, the method comprising contacting the system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for non-specifically degrading single-stranded DNA upon recognition of a DNA target nucleic acid, the method comprising contacting the system described herein with the target nucleic acid.

[0061] In some embodiments of any of the systems or methods provided herein, the contacting includes direct contact or indirect contact. In some embodiments of any of the systems or methods provided herein, contacting indirectly includes administering one or more nucleic acids encoding the RNA guide or CRISPR-related protein described herein under conditions that allow for the production of the RNA guide and / or CRISPR-related protein. In some embodiments of any of the systems or methods provided herein, the contacting includes in vivo contact or in vitro contact. In some embodiments of any of the systems or methods provided herein, contacting the target nucleic acid with the system includes contacting a cell containing the nucleic acid with the system under conditions that allow the CRISPR-related protein and the guide RNA to reach the target nucleic acid. In some embodiments of any of the systems or methods provided herein, contacting a cell with the system in vivo includes administering the system to a subject containing the cell under conditions that allow the CRISPR-related protein and the guide RNA to reach or be produced within the cell.

[0062] In another aspect, the present disclosure provides a system for use in an in vitro or ex vivo method, which is (a) a method for targeting and editing a target nucleic acid; (b) a method for non-specific degradation of a single-stranded nucleic acid in response to recognition of a nucleic acid; (c) a method for targeting and nicking a non-spacer complementary strand of a double-stranded target in response to recognition of a spacer complementary strand of the double-stranded target; (d) a method for targeting and cleaving a double-stranded target nucleic acid; (e) a method for detecting a target nucleic acid in a sample; (f) a method for specifically editing a double-stranded nucleic acid; (g) a method for base editing of a double-stranded nucleic acid; (h) a method for inducing genotype-specific or transcription state-specific cell death or dormancy in a cell; (i) a method for creating indels in a double-stranded nucleic acid target; (j) a method for inserting a sequence into a double-stranded nucleic acid target; or (k) a method for deleting or inverting a sequence in a double-stranded nucleic acid target.

[0063] In another aspect, the present disclosure provides a method for introducing an insertion or deletion into a target nucleic acid in a mammalian cell, which includes (a) a nucleic acid sequence encoding a CRISPR-related protein, wherein the CRISPR-related protein includes an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1-56; and (b) transfection of an RNA guide (or a nucleic acid encoding the RNA guide) including a direct repeat sequence and a spacer sequence having the ability to hybridize to the target nucleic acid, wherein the CRISPR-related protein has the ability to bind to the RNA guide, and recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide results in modification of the target nucleic acid.

[0064] In some embodiments of any of the methods provided herein, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments of any of the methods provided herein, the CRISPR-related protein comprises an amino acid sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods provided herein, the target nucleic acid is adjacent to a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5’-NTTN-3’, 5’-NTTR-3’ (e.g., 5’-TTTG-3’), or 5’-NNR-3’, where "N" is any nucleotide and "R" is A or G.

[0065] In some embodiments of any of the methods provided herein, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. In some embodiments of any of the methods provided herein, the CRISPR-related protein comprises an amino acid sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods provided herein, the target nucleic acid is adjacent to a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G.

[0066] In some embodiments of any of the methods provided herein, the transfection is transient transfection. In some embodiments of any of the methods provided herein, the cell is a human cell.

[0067] In another aspect, the present disclosure provides a composition comprising (a) a CRISPR-related protein or a nucleic acid encoding a CRISPR-related protein and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence; wherein the CRISPR-related protein is one or more of the following amino acid sequences: (i) PX1X2X3X4F (SEQ ID NO: 216) (wherein X1 is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M); (ii) RX1X2X3L (SEQ ID NO: 217) (wherein X1 is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (iii) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F, and X2 is K or R or V or E); (iv) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (v) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T, and X2 is N or Y or K or S); (vi) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I); (vii) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P's Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q);and (viii) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L), and the CRISPR-related protein can bind to the RNA guide and modify a target nucleic acid sequence complementary to the spacer sequence.;

[0068] In some embodiments of any of the compositions described herein, the direct repeat sequence is one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (wherein X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 227)); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any of the compositions described herein, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments of any of the compositions described herein, SEQ ID NO: 228 is proximal to the 3' end of the direct repeat.

[0069] In some embodiments of any of the compositions described herein, the CRISPR-associated protein comprises at least one (e.g., one, two, or three) RuvC domain or at least one split RuvC domain.

[0070] In some embodiments of any of the compositions described herein, the spacer sequence of the RNA guide comprises from about 15 nucleotides to about 55 nucleotides. In some embodiments of any of the compositions described herein, the spacer sequence of the RNA guide comprises 20 to 45 nucleotides.

[0071] In some embodiments of any of the compositions described herein, the CRISPR-associated protein comprises catalytic residues (e.g., aspartic acid or glutamic acid). In some embodiments of any of the compositions described herein, the CRISPR-associated protein cleaves a target nucleic acid. In some embodiments of any of the compositions described herein, the CRISPR-associated protein further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcriptional modification factor, a light gate control factor, a chemical inducibility factor, or a chromatin visualization factor.

[0072] In some embodiments of any of the compositions described herein, the nucleic acid encoding the CRISPR-associated protein is codon-optimized for expression in a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell). In some embodiments of any of the compositions described herein, the nucleic acid encoding the CRISPR-associated protein is operably linked to a promoter. In some embodiments of any of the compositions described herein, the nucleic acid encoding the CRISPR-associated protein is within a vector. In some embodiments, the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.

[0073] In some embodiments of any of the compositions described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the compositions described herein, the target nucleic acid comprises a PAM sequence.

[0074] In some embodiments of any of the compositions described herein, the CRISPR-related protein has non-specific nuclease activity.

[0075] In some embodiments of any of the compositions described herein, recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide results in modification of the target nucleic acid. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid results in an insertion event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid results in a deletion event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid results in cytotoxicity or cell death.

[0076] In some embodiments of any of the compositions described herein, the system further comprises a donor template nucleic acid. In some embodiments of any of the compositions described herein, the donor template nucleic acid is a DNA molecule. In some embodiments of any of the compositions described herein, the donor template nucleic acid is an RNA molecule.

[0077] In some embodiments of any of the compositions described herein, the RNA guide optionally includes a tracrRNA. In some embodiments of any of the compositions described herein, the system further includes a tracrRNA. In some embodiments of any of the compositions described herein, the system does not include a tracrRNA. In some embodiments of any of the compositions described herein, the CRISPR-associated protein is self-processing.

[0078] In some embodiments of any of the compositions described herein, the system is present in a delivery composition comprising nanoparticles, liposomes, exosomes, microvesicles, or a gene gun.

[0079] In some embodiments of any of the compositions described herein, the composition is intracellular. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a prokaryotic cell.

[0080] The effectors described herein provide additional features including, but not limited to: 1) novel nucleic acid editing properties and control mechanisms; 2) smaller size compared to higher versatility in delivery strategies; 3) cellular processes such as cell death induced by genotype; and 4) programmable RNA-guided DNA insertion, excision, and mobilization; and 5) a differentiated profile of existing immunity through non-human co-source. See, for example, Examples 1, 4, and 5 and Figures 1-3 and 5-11D. The addition of the novel DNA targeting system described herein to the toolbox of genome and epigenome manipulation techniques enables a wide range of applications to specific and programmed perturbations.

[0081] Other features and advantages of the invention will be apparent from the following detailed description and from the claims.

[0082] The drawings are a series of schematic representations showing the analysis results of a protein cluster called CLUST.091979.

Brief Description of the Drawings

[0083]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 1F

Figure 1G

Figure 1H

Figure 1I

Figure 1J

Figure 1K

Figure 1L

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11A

Figure 11B

Figure 11C

Figure 11D

BEST MODE FOR CARRYING OUT THE INVENTION

[0084] The CRISPR-Cas system is naturally diverse and contains various active mechanisms and functional elements that can be harnessed for programmable biotechnology. In nature, these systems provide efficient defense against foreign DNA and viruses while providing self-nonself discrimination to avoid self-targeting. In engineered settings, these systems provide a diverse toolbox of molecular technologies and define the boundaries of the targeting space. Using the methods described herein, additional mechanisms and parameters have been discovered that expand the ability of RNA-programmable nucleic acid manipulation within single subunit class 2 effector systems.

[0085] Unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In the practice or testing of the present invention, methods and materials similar or equivalent to those described herein may be used, but suitable methods and materials are described below. Publications, patent applications, patents, and other references mentioned herein are hereby incorporated by reference in their entirety. In case of conflict, this specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. Applicants reserve the right to alternatively claim any disclosed invention using transitional phrases such as "comprising", "consisting essentially of", or "consisting of" in accordance with standard patent law practice.

[0086] As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. For example, reference to "a nucleic acid" means one or more nucleic acids.

[0087] Note that in this specification, terms such as "preferably", "suitably", "generally", and "typically" are not used to limit the scope of the claimed invention or to imply that a particular feature is critical, essential, or more important to the structure or function of the claimed invention. Rather, these terms are merely intended to highlight alternative or additional features that may or may not be utilized in particular embodiments of the present invention.

[0088] Note that for the purposes of describing and defining the present invention, the term "substantially" is used herein to represent the degree of inherent uncertainty that can be attributed to any quantitative comparison, value, measurement, or other representation. The term "substantially" is also used herein to represent the degree to which a quantitative representation may vary from the reference being described without causing a change in the basic function of the subject matter in question.

[0089] As used herein, the term "CRISPR-Cas system" refers to nucleic acids and / or proteins that are involved in the expression of a CRISPR effector, or direct its activity, including an array encoding a CRISPR effector, an RNA guide, and other sequences and transcripts from a CRISPR locus.

[0090] As used herein interchangeably, the terms "CRISPR-associated protein", "CRISPR-Cas effector", "CRISPR effector", "effector", "effector protein", "CRISPR enzyme", etc. refer to a protein that binds to a target site on a nucleic acid specified by an enzyme activity-performing protein or an RNA guide. In some embodiments, the CRISPR effector has endonuclease activity, nickase activity, and / or exonuclease activity.

[0091] As used herein, the terms "RNA guide", "guide RNA", "gRNA", and "guide sequence" refer to any RNA molecule that facilitates targeting of an effector described herein to a target nucleic acid such as DNA and / or RNA. Exemplary "RNA guides" include, but are not limited to, crRNA, and crRNA hybridized or fused to either tracrRNA and / or a modulator RNA. In some embodiments, the RNA guide comprises both crRNA and tracrRNA fused to a single RNA molecule or as separate RNA molecules. In some embodiments, the RNA guide comprises crRNA and a modulator RNA fused to a single RNA molecule or as separate RNA molecules. In some embodiments, the RNA guide comprises crRNA, tracrRNA, and a modulator RNA fused to a single RNA molecule or as separate RNA molecules.

[0092] The terms "CRISPR effector complex", "effector complex", or "monitoring complex", as used herein, refer to a complex comprising a CRISPR effector and an RNA guide. The CRISPR effector complex may further comprise one or more accessory proteins. The one or more accessory proteins may be non-catalytic and / or non-target-binding. The crRNA may comprise a sequence that hybridizes to the tracrRNA. The crRNA:tracrRNA duplex may then bind to the CRISPR effector. As used herein, the term "pre-crRNA" refers to an unprocessed RNA molecule comprising a DR-spacer-DR sequence. As used herein, the term "mature crRNA" refers to the processed form of the pre-crRNA. The mature crRNA may comprise a DR spacer sequence, where the DR is a shortened form of the DR of the pre-crRNA and / or the spacer is a shortened form of the spacer of the pre-crRNA.

[0093] The terms "CRISPR RNA" and "crRNA", as used herein, refer to an RNA molecule comprising a guide sequence used by a CRISPR effector to specifically recognize a nucleic acid sequence. The crRNA "spacer" sequence is complementary to a nucleic acid target sequence and has the ability to bind partially or entirely to the nucleic acid target sequence.

[0094] The term "trans-activating crRNA" or "tracrRNA", as used herein, refers to an RNA molecule comprising a sequence that forms the structural and / or sequence motifs necessary for a CRISPR effector to bind to a specific target nucleic acid.

[0095] As used herein, the term "CRISPR array" refers to a nucleic acid (e.g., DNA) segment that includes CRISPR repeats and spacers, beginning with the first nucleotide of the first CRISPR repeat and ending with the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeats. As used herein, the terms "CRISPR repeat", "CRISPR direct repeat", and "direct repeat" refer to a plurality of short, direct, repeating sequences that show little or no sequence variation within a CRISPR array.

[0096] As used herein, the term "modulator RNA" refers to any RNA molecule that modulates (e.g., increases or decreases) the activity of a CRISPR effector or a ribonucleoprotein complex comprising a CRISPR effector. In some embodiments, the modulator RNA modulates the nuclease activity of a CRISPR effector or a ribonucleoprotein complex comprising a CRISPR effector.

[0097] As used herein, the term "target nucleic acid" refers to a nucleic acid that includes a nucleotide sequence complementary to all or part of a spacer in an RNA guide. In some embodiments, the target nucleic acid includes a gene. In some embodiments, the target nucleic acid includes a non-coding region (e.g., a promoter). In some embodiments, the target nucleic acid is single-stranded. In some embodiments, the target nucleic acid is double-stranded. As used herein, a "transcriptionally active site" refers to a site within a nucleic acid sequence that is actively being transcribed.

[0098] As used herein, the term "protospacer adjacent motif" or "PAM" refers to a DNA sequence adjacent to a target sequence to which a complex comprising an effector and an RNA guide binds. In some embodiments, a PAM is required for enzymatic activity. As used herein, the term "adjacent" includes cases where the RNA guide of the complex specifically binds, interacts, or associates with a target sequence that is directly adjacent to the PAM. In such cases, no nucleotides are present between the target sequence and the PAM. The term "adjacent" also includes cases where a small number (e.g., 1, 2, 3, 4, or 5) of nucleotides are present between the target sequence to which the targeting moiety binds and the PAM. As used herein, the term "recognizing a PAM sequence" refers to the binding of a complex comprising a CRISPR-associated protein and a crRNA to a target nucleic acid, where the target nucleic acid is adjacent to the PAM sequence.

[0099] The terms "activated CRISPR effector complex", "activated CRISPR complex", and "activated complex", as used herein, refer to a CRISPR effector complex capable of modifying a target nucleic acid. In some embodiments, the activated CRISPR complex is capable of modifying a target nucleic acid after the activated CRISPR complex binds to the target nucleic acid. In some embodiments, binding of the activated CRISPR complex to the target nucleic acid results in additional cleavage events such as collateral cleavage.

[0100] The term "cleavage event", as used herein, refers to the cleavage of a nucleic acid such as DNA and / or RNA. In some embodiments, the cleavage event, as used herein, refers to the cleavage in a target nucleic acid created by a nuclease of the CRISPR system described herein. In some embodiments, the cleavage event is a double-stranded DNA cleavage. In some embodiments, the cleavage event is a single-stranded DNA cleavage. In some embodiments, the cleavage event refers to the cleavage of collateral nucleic acids.

[0101] As used herein, the term "collateral nucleic acid" refers to a nucleic acid substrate that is non-specifically cleaved by an activated CRISPR complex. As used herein when referring to a CRISPR effector, the term "collateral DNase activity" refers to the non-specific DNase activity of an activated CRISPR complex. As used herein when referring to a CRISPR effector, the term "collateral RNase activity" refers to the non-specific RNase activity of an activated CRISPR complex.

[0102] As used herein, the term "donor template nucleic acid" refers to a nucleic acid molecule that can be used to effect templated changes to a target sequence or target-proximal sequence after a CRISPR effector as described herein has modified a target nucleic acid. In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome).

[0103] As used herein, the terms "polynucleotide", "nucleotide", "oligonucleotide", and "nucleic acid" may be used interchangeably to refer to nucleic acids including DNA, RNA, their derivatives, or combinations thereof. One of ordinary skill in the art can construct gene expression constructs and recombinant cells according to the present invention using methods well known in the art. These methods include in vitro recombinant DNA techniques, synthetic techniques, in vivo recombinant techniques, and polymerase chain reaction (PCR) techniques. See, for example, Maniatis et al., 1989, MOLECULAR CLONING: A LABORATORY MANUAL, Cold Spring Harbor Laboratory, New York; Ausubel et al., 1989, CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, Greene Publishing Associates and Wiley Interscience, New York, and the techniques described in PCR Protocols: A Guide to Methods and Applications (Innis et al., 1990, Academic Press, San Diego, Calif.).

[0104] The term "gene modification" or "genetic engineering" broadly refers to the manipulation of the genome or nucleic acids of a cell. Similarly, the terms "genetically engineered" and "engineered" refer to cells containing a manipulated genome or nucleic acid. Methods of gene modification include, for example, heterologous gene expression, insertion or deletion of genes or promoters, nucleic acid mutations, changes in gene expression or inactivation, enzyme engineering, directed evolution, knowledge-based design, random mutagenesis, gene shuffling, and codon optimization.

[0105] The term "recombinant" indicates that a nucleic acid, protein, or cell is a product of genetic modification, engineering, or recombination. Generally, the term "recombinant" refers to a nucleic acid, protein, or cell that contains or is encoded by genetic material from multiple sources. As used herein, the term "recombinant" can also be used to describe a cell that contains a mutant nucleic acid or protein, including a mutant form of an endogenous nucleic acid or protein. The terms "recombinant cell" and "recombinant host" can be used interchangeably. In some embodiments, the recombinant cell comprises a CRISPR effector disclosed herein. The CRISPR effector can be codon-optimized for expression in the recombinant cell. In some embodiments, the recombinant cell disclosed herein further comprises an RNA guide. In some embodiments, the RNA guide of the recombinant cell disclosed herein comprises a tracrRNA. In some embodiments, the recombinant cell disclosed herein comprises a modulator RNA. In some embodiments, the recombinant cell is a prokaryotic cell such as an E. coli cell. In some embodiments, the recombinant cell is a eukaryotic cell such as a mammalian cell including a human cell.

[0106] Identification of CLUST.091979 This application relates to the identification, engineering, and use of a novel protein family referred to herein as "CLUST.091979". As shown in Figure 2, the proteins of CLUST.091979 contain RuvC domains (designated as RuvC I, RuvC II, and RuvC III). As shown in Table 5, the size of the effector of CLUST.091979 ranges from about 700 amino acids to about 800 amino acids. Thus, as shown below, the effector of CLUST.091979 is smaller than effectors known in the art. See, for example, Table 1.

[0107] [Table 1]

[0108] The effectors of CLUST.091979 were identified using computational methods and algorithms to search for and identify proteins that exhibit strong co-occurrence patterns with other specific functions. In certain embodiments, these computational methods related to identifying proteins that co-occur in close proximity to the CRISPR array. The methods disclosed herein are also useful in identifying proteins that naturally occur within a range in close proximity to other features, both non-coding and protein-coding (e.g., fragments of phage sequences in the non-coding region of a bacterial locus; or CRISPR Cas1 proteins). It is understood that the methods and calculations described herein may be implemented on one or more computing devices.

[0109] A set of genomic sequences was obtained from a genomic or metagenomic database. The database included short reads, or contig-level data, or assembled scaffolds, or complete genomic sequences of organisms. Similarly, the database may include genomic sequence data from prokaryotes or eukaryotes, or data from metagenomic environmental samples. Examples of database repositories include RefSeq of the National Center for Biotechnology Information (NCBI), GenBank of the NCBI, the Whole Genome Shotgun (WGS) of the NCBI, and the Integrated Microbial Genomes (IMG) of the Joint Genome Institute (JGI).

[0110] In some embodiments, a minimum size requirement is imposed on the selection of genomic sequence data of a specified minimum length. In certain exemplary embodiments, the minimum contig length may be 100 nucleotides, 500 nt, 1 kb, 1.5 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 40 kb, or 50 kb.

[0111] In some embodiments, known or predicted proteins are extracted from complete or a selected set of genomic sequence data. In some embodiments, known or predicted proteins are taken from extracting the coding sequence (CDS) annotations provided by a source database. In some embodiments, predicted proteins are determined by applying computational methods to identify proteins from nucleotide sequences. In some embodiments, the GeneMark suite is used to predict proteins from genomic sequences. In some embodiments, Prodigal is used to predict proteins from genomic sequences. In some embodiments, multiple protein prediction algorithms are used for the same set of sequence data, and duplicates may be removed from the resulting set of proteins.

[0112] In some embodiments, CRISPR arrays are identified from genomic sequence data. In some embodiments, PILER-CR is used to identify CRISPR arrays. In some embodiments, the CRISPR Recognition Tool (CRT) is used to identify CRISPR arrays. In some embodiments, CRISPR arrays are identified by a discovery approach that identifies nucleotide motifs that are repeated a minimum number (e.g., 2, 3, or 4 times), where the interval between successive occurrences of the repeated motif does not exceed a specified length (e.g., 50, 100, or 150 nucleotides). In some embodiments, multiple CRISPR array identification tools are used for the same set of sequence data, and duplicates may be removed from the resulting set of CRISPR arrays.

[0113] In some embodiments, proteins are identified that are in very close proximity to a CRISPR array (referred to herein as a "CRISPR proximal protein cluster"). In some embodiments, proximity is defined as a nucleotide distance and may be within 20 kb, 15 kb, or 5 kb. In some embodiments, proximity is defined as the number of open reading frames (ORFs) between a protein and a CRISPR array, and certain exemplary distances can be 10, 5, 4, 3, 2, 1, or 0 ORFs. Proteins identified as being within a very close range to a CRISPR array are then grouped into homologous protein clusters. In some embodiments, blastclust is used to form CRISPR proximal protein clusters. In certain other embodiments, mmseqs2 is used to form CRISPR proximal protein clusters.

[0114] To establish a strong co-occurrence pattern among members of a CRISPR proximal protein cluster, a BLAST search of each member of the protein cluster may be performed against a pre-compiled complete set of known and predicted proteins. In some embodiments, UBLAST or mmseqs2 may be used to search for similar proteins. In some embodiments, the search may be performed only for a representative subset of proteins within a family.

[0115] In some embodiments, co-occurrence is determined by ranking or filtering CRISPR proximal protein clusters by a metric. One exemplary metric is the ratio of the number of elements within a protein cluster to the number of BLAST matches to a particular E-value threshold. In some embodiments, a fixed E-value threshold may be used. In other embodiments, the E-value threshold may be determined by the most distant member of the protein cluster. In some embodiments, a global set of proteins is clustered and the co-occurrence metric is the ratio of the number of elements of CRISPR proximal proteins to the number of elements of one or more global clusters included.

[0116] In some embodiments, by using a manual review process, the potential functionality and minimal set of components of a system engineered based on the naturally occurring locus structure of proteins in a cluster are evaluated. In some embodiments, a graphical representation of the protein cluster can be useful for the manual review, which can include information including pairwise sequence similarity, phylogenetic tree, source organism / environment, predicted functional domains, and graphical depiction of the locus structure. In some embodiments, the graphical depiction of the locus structure can filter highly representative neighboring protein families. In some embodiments, the representativeness may be calculated by the ratio of the number of relevant neighboring proteins to the size of one or more of the included global clusters. In certain exemplary embodiments, the graphical representation of the protein cluster can include a depiction of the CRISPR array structure of the naturally occurring locus. In some embodiments, the graphical representation of the protein cluster can include a depiction of the number of conserved direct repeats relative to the length of the putative CRISPR array, or the number of unique spacer sequences relative to the length of the putative CRISPR array. In some embodiments, the graphical representation of the protein cluster includes a depiction of various co-occurrence metrics of putative effectors with the CRISPR array, which can predict novel CRISPR-Cas systems and identify their components.

[0117] Pooled screening of CLUST.091979 To efficiently verify the activity, mechanism, and functional parameters of the engineered CLUST.091979 CRISPR-Cas system identified herein, a pooled screening approach in Escherichia coli (E. coli) was used as described in Example 4. First, from the computational identification of the conserved proteins and non-coding elements of the CRIST.091979 CRISPR-Cas system, individual components were assembled into a single artificial expression vector (based on the pET-28a+ backbone in one embodiment) using DNA synthesis and molecular cloning. In a second embodiment, the effector and non-coding elements are transcribed into mRNA transcripts and the individual effectors are translated using different ribosome binding sites.

[0118] Second, native crRNAs and targeting spacers are replaced with a library of unprocessed crRNAs containing non-native spacers targeting the second plasmid pACYC184. This crRNA library is cloned into a vector backbone (e.g., pET-28a+) containing the effector and non-coding elements, and subsequently this library is transformed into Escherichia coli (E. coli) together with the pACYC184 plasmid target. As a result, each resulting Escherichia coli (E. coli) cell contains only one targeting array. In an alternative embodiment, a library of unprocessed crRNAs containing non-native spacers further targets essential Escherichia coli (E. coli) genes such as those cited from sources such as Baba et al. (2006) Mol. Syst. Biol. 2:2006.0008; and Gerdes et al. (2003) J. Bacteriol. 185(19):5673-84 (the entire contents of each of which are incorporated herein by reference). In this embodiment, the positive targeted activity of the novel CRISPR-Cas system that disrupts essential gene function results in cell death or growth arrest. In some embodiments, essential gene targeting spacers can be combined with the pACYC184 target.

[0119] Thirdly, grow Escherichia coli (E. coli) under antibiotic selection. In one embodiment, triple antibiotic selection: kanamycin is used to confirm the success of transformation of the pET-28a+ vector containing the engineered CRISPR effector system, and chloramphenicol and tetracycline are used to confirm the success of co-transformation of the pACYC184 target vector. pACYC184 usually confers resistance to chloramphenicol and tetracycline. Under antibiotic selection, due to the positive activity of the novel CRISPR-Cas system targeting this plasmid, cells that actively express the effector, non-coding elements, and specific active elements of the crRNA library will be eliminated. Typically, the population of surviving cells is analyzed 12 - 14 hours after transformation. In some embodiments, the analysis of surviving cells is performed 6 - 8 hours after transformation, 8 - 12 hours after transformation, up to 24 hours after transformation, or more than 24 hours after transformation. Examining the population of surviving cells at a later time point compared to an earlier time point results in a depletion of the signal compared to the inactive crRNA.

[0120] In some embodiments, dual antibiotic selection is used. Removing either chloramphenicol or tetracycline to remove the selection pressure can provide new information regarding the targeting substrate, sequence specificity, and efficacy. For example, negative selection can occur in Escherichia coli (E. coli) where depletion of both the selected and unselected genes is observed due to cleavage of dsDNA in the selected or unselected genes. If the CRISPR-Cas system interferes with transcription or translation (e.g., by binding or cleavage of transcripts), the selection is observed only for the target of the selected resistance gene rather than the unselected resistance gene.

[0121] In some embodiments, the successful transformation of a pET-28a+ vector containing an engineered CRISPR-Cas system is confirmed using only kanamycin. This embodiment is suitable for libraries containing spacers targeting essential E. coli genes, as no additional selection other than kanamycin is required to observe growth changes. In this embodiment, chloramphenicol and tetracycline dependencies are removed, and their targets (if present) in the library provide additional sources of negative or positive information regarding the targeting substrate, sequence specificity, and efficacy.

[0122] Since the pACYC184 plasmid contains a diverse set of features and sequences that can affect the activity of the CRISPR-Cas system, mapping the active crRNAs from a pooled screen to pACYC184 provides activity patterns that may suggest various mechanisms of action and functional parameters. In this way, the features required for the reconstitution of novel CRISPR-Cas systems in heterologous prokaryotic species can be more comprehensively tested and studied.

[0123] Important advantages of the in vivo pooled screen described herein include the following: (1) Versatility - Plasmid design allows for the expression of multiple effectors and / or non-coding elements; library cloning strategies enable the expression of crRNAs in both transcriptional directions predicted computationally. (2) Comprehensive testing of mechanisms of action and functional parameters allows for the evaluation of diverse interference mechanisms, including nucleic acid cleavage; co-occurrence of features such as transcription and plasmid DNA replication; and examination of flanking sequences for the crRNA library to reliably determine PAMs with 4N complexity equivalence. (3) Sensitivity - pACYC184 is a low-copy plasmid, and even with a low interference rate, it can remove the antibiotic resistance encoded by the plasmid, thus achieving high sensitivity for CRISPR-Cas activity; and (4) Optimized molecular biology steps that achieve higher speed and throughput for efficiency-RNA sequencing enable the protein expression samples to be directly collected from viable cells in the screen.

[0124] The novel CRIST.091979 CRISPR-Cas family described herein was evaluated by using this in vivo pool-type screen to assess its operable elements, mechanisms and parameters, and its ability to be active and reprogrammed in engineered systems outside of its endogenous cellular environment.

[0125] Activity and Modification of CRISPR Effectors In some embodiments, the CRISPR effector and RNA guide of CRIST.091979 form a binary complex that may include other components. The binary complex is activated upon binding to a nucleic acid substrate (i.e., a sequence-specific substrate or target nucleic acid) complementary to the spacer sequence in the RNA guide. In some embodiments, the sequence-specific substrate is double-stranded DNA. In some embodiments, the sequence-specific substrate is single-stranded DNA. In some embodiments, the sequence-specific substrate is single-stranded RNA. In some embodiments, the sequence-specific substrate is double-stranded RNA. In some embodiments, sequence specificity requires a perfect match between the spacer sequence in the RNA guide (e.g., crRNA) and the target substrate. In other embodiments, sequence specificity requires a partial (continuous or discontinuous) match between the spacer sequence in the RNA guide (e.g., crRNA) and the target substrate.

[0126] In some embodiments, the CRISPR effector of the present invention has enzymatic activity, for example, nuclease activity, over a wide range of pH conditions. In some embodiments, the nuclease has enzymatic activity, for example, nuclease activity, at a pH of about 3.0 to about 12.0. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 4.0 to about 10.5. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 5.5 to about 8.5. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 6.0 to about 8.0. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 7.0.

[0127] In some embodiments, the CRISPR effector of the present invention has enzymatic activity, for example, nuclease activity, in a temperature range of about 10°C to about 100°C. In some embodiments, the CRISPR effector of the present invention has enzymatic activity in a temperature range of about 20°C to about 90°C. In some embodiments, the CRISPR effector of the present invention has enzymatic activity at a temperature of about 20°C to about 25°C or at a temperature of about 37°C.

[0128] In some embodiments, the binary complex becomes activated upon binding to the target substrate. In some embodiments, the activated complex exhibits "multiple metabolic turnovers" activity, and thus, upon acting on the target substrate (e.g., upon cleavage thereof), the activated complex remains in the activated state. In some embodiments, the activated binary complex exhibits "single metabolic turnover" activity, and thus, upon acting on the target substrate, the binary complex returns to the inactive state. In some embodiments, the activated binary complex exhibits non-specific (i.e., "collateral") cleavage activity, and thus the complex cleaves non-target nucleic acids. In some embodiments, the non-target nucleic acid is a DNA molecule (e.g., single-stranded or double-stranded DNA). In some embodiments, the non-target nucleic acid is an RNA molecule (e.g., single-stranded or double-stranded RNA).

[0129] In some embodiments where the CRISPR effector of the present invention induces double-strand cleavage or single-strand cleavage in a target nucleic acid (e.g., genomic DNA), the double-strand cleavage can stimulate an endogenous DNA repair pathway in the cell, including homologous recombination (HDR), non-homologous end joining (NHEJ), or alternative non-homologous end joining (A-NHEJ). NHEJ can repair the cleaved target nucleic acid without the need for a homology template, resulting in the deletion or insertion of one or more nucleotides at the target locus. HDR can occur using a homology template such as donor DNA. The homology template can include a sequence homologous to the sequence flanking the target nucleic acid cleavage site. In some examples, HDR can insert an exogenous polynucleotide sequence into the cleavage target locus. Modifications of the target DNA resulting from NHEJ and / or HDR can result in, for example, mutations, deletions, alterations, integrations, gene corrections, gene replacements, gene tagging, transgene knock-ins, gene disruptions, and / or gene knockouts.

[0130] In some embodiments, the CRISPR effectors described herein can be fused to one or more peptide tags, including a His tag, GST tag, FLAG tag, or myc tag. In some embodiments, the CRISPR effectors described herein can be fused to a detectable moiety, such as a fluorescent protein (e.g., green fluorescent protein or yellow fluorescent protein). In some embodiments, the CRISPR effector and / or accessory protein of the present disclosure is fused to a peptide or non-peptide moiety that allows the protein to penetrate or localize to a tissue, cell, or region of the cell. For example, the CRISPR effector of the present disclosure may include a nuclear localization sequence (NLS), such as the SV40 (simian virus 40) NLS, c-Myc NLS, or other suitable monopartite NLS. The NLS may be fused to the N-terminus and / or C-terminus of the CRISPR effector and may be fused alone (i.e., a single NLS) or concatenated (e.g., a chain of 2, 3, 4, etc. NLSs).

[0131] In some embodiments, at least one nuclear export signal (NES) is attached to the nucleic acid sequence encoding the CRISPR effector. In some embodiments, C-terminal and / or N-terminal NLS or NES are attached for optimal expression and nuclear targeting in eukaryotic cells, such as human cells.

[0132] In embodiments where a tag is fused to the CRISPR effector, such tag can facilitate affinity-based or charge-based purification of the CRISPR effector, for example, by liquid chromatography or bead separation using immobilized affinity or ion exchange reagents. As a non-limiting example, the recombinant CRISPR effector of the present disclosure includes a polyhistidine (His) tag and is loaded onto a chromatography column containing immobilized metal ions for purification (e.g., Zn 2+ , Ni 2+ , Cu 2+ ions, and this resin may be a resin prepared individually or a commercially available resin or a ready-made column such as the HisTrap FF column commercialized by GE Healthcare Life Sciences, Marlborough, Massachusetts. After the loading step, the column is optionally rinsed using, for example, one or more suitable buffer solutions, and then the protein with the His tag added is eluted using a suitable elution buffer. Alternatively or in addition, if the recombinant CRISPR effector of the present disclosure utilizes a FLAG tag, such protein may be purified using immunoprecipitation methods known in the art. Other suitable purification methods for the tagged CRISPR effector or accessory protein of the present disclosure will be apparent to those skilled in the art.

[0133] The proteins described herein (e.g., CRISPR effectors or accessory proteins) can be delivered or used either as nucleic acid molecules or polypeptides. When using nucleic acid molecules, the nucleic acid molecules encoding CRISPR effectors can be codon-optimized. The nucleic acids can be codon-optimized for use in any organism of interest, specifically in human cells or bacteria. For example, the nucleic acids can be codon-optimized for any non-human eukaryote, including mice, rats, rabbits, dogs, livestock, or non-human primates. Codon usage tables are readily available in the “Codon Usage Database” available, for example, at www.kazusa.orjp / codon / , and these tables can be adapted in several ways. See Nakamura et al. Nucl. Acids Res. 28:292 (2000) (incorporated herein by reference in its entirety). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, such as Gene Forge (Aptagen; Jacobus, PA).

[0134] In some examples, the nucleic acids of the disclosure encoding a CRISPR effector for expression in eukaryotic cells (e.g., human, or other mammalian cells) include one or more introns, i.e., one or more non-coding sequences that include a splice donor sequence at a first end (e.g., the 5′ end) and a splice acceptor sequence at a second end (e.g., the 3′ end). In various embodiments of the disclosure, any suitable splice donor / splice acceptor can be used, including, without limitation, simian virus 40 (SV40) introns, β-globin introns, and synthetic introns. Alternatively or in addition, the nucleic acids of the disclosure encoding a CRISPR effector or accessory protein can include a transcription termination signal, such as a polyadenylation (polyA) signal, at the 3′ end of the DNA coding sequence. In some examples, the polyA signal is located very close to or adjacent to an intron, such as an SV40 intron.

[0135] Inactivated CRISPR effector The CRISPR effectors described herein can be modified such that they have reduced nuclease activity, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation when compared to a wild-type CRISPR effector. Nuclease activity can be reduced by several methods well known in the art, e.g., by introducing mutations into the nuclease domain of the protein. In some embodiments, the catalytic residues of the nuclease activity are identified and the nuclease activity may be reduced by substituting those amino acid residues with different amino acid residues (e.g., glycine or alanine).

[0136] The inactivated CRISPR effector can contain or be associated with one or more functional domains (e.g., via a fusion protein, a linker peptide, a "GS" linker, etc.). Such functional domains can have various activities, e.g., methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and switch activity (e.g., light-inducible). In some embodiments, the functional domains are Kruppel-associated box (KRAB), VP64, VP16, Fok1, P65, HSF1, MyoD1, and biotin-APEX.

[0137] Positioning one or more functional domains on an inactivated CRISPR effector enables the functional domains to be in the correct spatial orientation to target the effect of the function attributed to them. For example, if the functional domain is a transcriptional activator (e.g., VP16, VP64, or p65), the transcriptional activator is placed in a spatial orientation that enables it to affect the transcription of the target. Similarly, a transcriptional repressor is positioned to affect the transcription of the target, and a nuclease (e.g., Fok1) is positioned to cleave or partially cleave the target. In some embodiments, the functional domain is located at the N-terminus of the CRISPR effector. In some embodiments, the functional domain is located at the C-terminus of the CRISPR effector. In some embodiments, the inactivated CRISPR effector is modified to include a first functional domain at the N-terminus and a second functional domain at the C-terminus.

[0138] Split enzyme The present disclosure also provides split versions of the CRISPR effectors described herein. Split versions of the CRISPR effector may be advantageous for delivery. In some embodiments, the CRISPR effector is split into two parts of the enzyme, which together comprise a substantially functional CRISPR effector.

[0139] The splitting can be done in such a way that one or more catalytic domains are not affected. The CRISPR effector may function as a nuclease or may be an inactivated enzyme that is an RNA-binding protein with very little or no catalytic activity (e.g., due to one or more mutations in its catalytic domain).

[0140] In some embodiments, the nuclease lobe and the α-helix lobe are expressed as separate polypeptides. These lobes do not interact with each other per se, but an RNA guide recruits them into a complex that recapitulates the activity of the full-length CRISPR effector and catalyzes site-specific DNA cleavage. Using a modified RNA guide, dimerization is prevented, rendering split enzyme activity ineffective and enabling the development of an inducible dimerization system. Split enzymes are described, for example, in Wright et al. “Rational design of a split-Cas9 enzyme complex,” Proc. Nat’l. Acad. Sci., 112.10 (2015):2984-2989 (incorporated herein by reference in its entirety).

[0141] In some embodiments, the split enzyme may be fused to a dimerization partner, for example, by utilizing a rapamycin-sensitive dimerization domain. This enables the creation of a chemically inducible CRISPR effector for temporally controlling CRISPR effector activity. By being split into two fragments in this way, the CRISPR effector can be made chemically inducible, and a rapamycin-sensitive dimerization domain can be used for the controlled reassembly of the CRISPR effector.

[0142] The split point is typically designed in silico and cloned into a construct. Mutations may be introduced into the split enzyme during this process, and non-functional domains may be removed. In some embodiments, the two parts or fragments of the split CRISPR effector (i.e., the N-terminal and C-terminal fragments) can form a complete CRISPR effector that includes, for example, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% of the sequence of the wild-type CRISPR effector.

[0143] Self-activating or inactivating enzymes The CRISPR effectors described herein may be designed to be self-activating or self-inactivating. In some embodiments, the CRISPR effector is self-inactivating. For example, a target sequence can be introduced into the construct encoding the CRISPR effector. Thus, when the CRISPR effector cleaves the target sequence, the construct encoding the enzyme can thereby self-inactivate its expression. Methods for constructing self-inactivating CRISPR systems are described, for example, in Epstein et al., “Engineering a Self-Inactivating CRISPR System for AAV Vectors,” Mol. Ther., 24(2016):S50 (incorporated herein by reference in its entirety).

[0144] In some other embodiments, an additional RNA guide that is expressed under the control of a weak promoter (e.g., the 7SK promoter) can target the nucleic acid sequence encoding the CRISPR effector to inhibit and / or prevent its expression (e.g., by interfering with nucleic acid transcription and / or translation). Transfecting a cell with a vector expressing a CRISPR effector, an RNA guide, and an RNA guide that targets the nucleic acid encoding the CRISPR effector can lead to efficient destruction of the nucleic acid encoding the CRISPR effector, reduce the CRISPR effector level, and thus limit genome editing activity.

[0145] In some embodiments, the genome editing activity of a CRISPR effector can be regulated through endogenous RNA signatures (e.g., miRNA) in mammalian cells. By using miRNA complementary sequences in the 5'-UTR of the mRNA encoding the CRISPR effector, a CRISPR effector switch can be created. This switch selectively and efficiently responds to miRNA in target cells. Thus, this switch can differentially control genome editing by sensing endogenous miRNA activity within a heterogeneous cell population. Therefore, this switch system can provide a framework for cell type-selective genome editing and cell engineering based on intracellular miRNA information (Hirosawa et al. “Cell-type-specific genome editing with a microRNA-responsive CRISPR-Cas9 switch,” Nucl. Acids Res., 2017 Jul 27;45(13):e118).

[0146] Inducible CRISPR effector The CRISPR effector may be inducible, for example, photoinducible or chemically inducible. This mechanism enables activation of the functional domain in the CRISPR enzyme. The photoinducible ability can be achieved by various methods known in the art, for example, by designing a fusion complex in which the CRY2PHR / CIBN pair is used in a split CRISPR effector (see, for example, Konermann et al. “Optical control of mammalian endogenous transcription and epigenetic states,” Nature, 500.7463 (2013): 472). The chemically inducible ability can be achieved, for example, by designing a fusion complex in which the FKBP / FRB (FK506 binding protein / FKBP rapamycin binding domain) pair is used in a split CRISPR effector. Rapamycin is required for the formation of the fusion complex and thus activates the CRISPR effector (see, for example, Zetsche et al. “A split-Cas9 architecture for inducible genome editing and transcription modulation,” Nature Biotech., 33.2 (2015): 139-142).

[0147] Furthermore, the expression of the CRISPR effector can be regulated by inducible promoters, such as transcriptional activation under the control of tetracycline or doxycycline (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. When delivered as RNA, the expression of the RNA-targeting effector protein may be regulated by riboswitches that can sense small molecule-like tetracyclines (see, e.g., Goldfless et al., “Direct and specific chemical control of eukaryotic translation with a synthetic RNA-protein interaction,” Nucl. Acids Res., 40.9 (2012): e64-e64).

[0148] Various embodiments of inducible CRISPR effectors and inducible CRISPR systems are described, for example, in U.S. Patent No. 8,871,445, U.S. Patent Application Publication No. 2016 / 0208243, and International Publication No. 2016 / 205764, each of which is incorporated herein by reference in its entirety.

[0149] Functional mutation Various mutations or modifications can be introduced into the CRISPR effectors as described herein to improve specificity and / or robustness. In some embodiments, amino acid residues that recognize the protospacer adjacent motif (PAM) are identified. The CRISPR effectors described herein can be further modified to recognize different PAMs, for example, by substituting the amino acid residues that recognize the PAM with other amino acid residues. In some embodiments, the CRISPR effector can recognize, for example, 5'-NTTN-3', 5'-NTTR-3', 5'-RTTR-3', 5'-TNNT-3', 5'-TNRT-3', 5'-TSRT-3', 5'-TGRT-3', 5'-TNRY-3', 5'-TTNR-3', 5'-TTYR-3', 5'-TTTR-3', 5'-TTCV-3', 5'-DTYR-3', 5'-WTTR-3', 5'-NNR-3', 5'-NYR-3', 5'-YYR-3', 5'-TYR-3', 5'-TTN-3', 5'-TTR-3', 5'-CNT-3', 5'-NGG-3', 5'-BGG-3', or 5'-R-3', where "N" is any nucleotide, "B" is C or G or T, "D" is A or G or T, "R" is A or G, "S" is G or C, "V" is A or C or G, "W" is A or T, and "Y" is C or T.

[0150] In some embodiments, the CRISPR effectors described herein can have one or more functional activities modified by mutating one or more amino acid residues. For example, in some embodiments, the CRISPR effector has its helicase activity modified by mutating one or more amino acid residues. In some embodiments, the CRISPR effector has its nuclease activity (e.g., endonuclease activity or exonuclease activity) modified by mutating one or more amino acid residues. In some embodiments, the CRISPR effector has its ability functionally related to the RNA guide modified by mutating one or more amino acid residues. In some embodiments, the CRISPR effector has its ability functionally related to the target nucleic acid modified by mutating one or more amino acid residues.

[0151] In some embodiments, the CRISPR effectors described herein have the ability to cleave a target nucleic acid molecule. In some embodiments, the CRISPR effector cleaves both strands of the target nucleic acid molecule. However, in some embodiments, the CRISPR effector has its cleavage activity modified by mutating one or more amino acid residues. For example, in some embodiments, the CRISPR effector may include one or more mutations that increase the ability of the CRISPR effector to cleave the target nucleic acid. In another example, in some embodiments, the CRISPR effector may include one or more mutations such that the enzyme no longer has the ability to cleave the target nucleic acid. In other embodiments, the CRISPR effector may include one or more mutations such that the enzyme has the ability to cleave only one strand of the target nucleic acid (i.e., nickase activity). In some embodiments, the CRISPR effector has the ability to cleave the strand of the target nucleic acid that is complementary to the strand to which the RNA guide hybridizes. In some embodiments, the CRISPR effector has the ability to cleave the strand of the target nucleic acid to which the RNA guide hybridizes.

[0152] In some embodiments, one or more residues of the CRISPR effectors disclosed herein are mutated in the arginine moiety. In some embodiments, one or more residues of the CRISPR effectors disclosed herein are mutated in the glycine moiety. In some embodiments, one or more residues of the CRISPR effectors disclosed herein are mutated based on the consensus residues of the phylogenetic alignment of the CRISPR effectors disclosed herein.

[0153] In some embodiments, the CRISPR effectors described herein may be engineered to include deletions in one or more amino acid residues in order to reduce the size of the enzyme while retaining one or more desired functional activities (e.g., nuclease activity and the ability to interact functionally with an RNA guide). This truncated CRISPR effector may advantageously be used in combination with a delivery system with limited payload.

[0154] In one aspect, the present disclosure provides a nucleic acid sequence that is at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the nucleic acid sequences described herein while maintaining the domain configuration shown in FIG. 2. In another aspect, the present disclosure also provides an amino acid sequence that is at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequences described herein while maintaining the domain configuration shown in FIG. 2.

[0155] In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., consecutive or non-consecutive nucleotides) that is the same as the sequence described herein. In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., consecutive or non-consecutive nucleotides) that is different from the sequence described herein.

[0156] In some embodiments, the amino acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., consecutive or non-consecutive amino acid residues) that is the same as the sequence described herein. In some embodiments, the amino acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., consecutive or non-consecutive amino acid residues) that is different from the sequence described herein.

[0157] To determine the percent identity between two amino acid sequences, or two nucleic acid sequences, the sequences are aligned for optimal comparison (e.g., gaps may be introduced in one or both of the first and second amino acid or nucleic acid sequences for optimal alignment and non-homologous sequences may be disregarded for comparison purposes). Generally, the length of the reference sequence aligned for comparison purposes should be at least 80% of the length of the reference sequence, and in some embodiments, at least 90%, 95%, or 100% of the length of the reference sequence. Next, the amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by those sequences, taking into account the number of gaps that need to be introduced to optimally align the two sequences and the length of each gap. For the purposes of the present disclosure, sequence comparison and determination of percent identity between two sequences can be accomplished using the Blossum 62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5.

[0158] In some embodiments, the nuclease comprises a sequence described as PX1X2X3X4F (SEQ ID NO: 216), where X1 is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M. In some embodiments, the sequence set forth in SEQ ID NO: 216 is an N-terminal sequence. In some embodiments, the nuclease comprises a sequence described as RX1X2X3L (SEQ ID NO: 217), where X1 is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K. In some embodiments, the nuclease comprises a sequence described as NX1YX2 (SEQ ID NO: 218), where X1 is I or L or F and X2 is K or R or V or E. In some embodiments, the nuclease comprises a sequence described as KX1X2X3FAX4X5KD (SEQ ID NO: 219), where X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 219 is a C-terminal sequence. In some embodiments, the nuclease comprises a sequence described as LX1NX2 (SEQ ID NO: 220), where X1 is G or S or C or T and X2 is N or Y or K or S. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 220 is a C-terminal sequence. In some embodiments, the nuclease comprises a sequence described as PX1X2X3X4SQX5DS (SEQ ID NO: 221), where X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 221 is a C-terminal sequence.In some embodiments, the nuclease comprises a sequence described as KX1X2VRX3X4QEX5H (SEQ ID NO: 222), where X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P's Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q. In some embodiments of any part of the systems described herein, the sequence of SEQ ID NO: 222 is a C-terminal sequence. In some embodiments, the nuclease comprises a sequence described as X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223), where X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L. In some embodiments of any part of the systems described herein, the sequence of SEQ ID NO: 223 is a C-terminal sequence.

[0159] RNA Guide and RNA Guide Modification In some embodiments, the RNA guide described herein comprises uracil (U). In some embodiments, the RNA guide described herein comprises thymine (T). In some embodiments, the direct repeat sequence of the RNA guide described herein comprises uracil (U). In some embodiments, the direct repeat sequence of the RNA guide described herein comprises thymine (T). In some embodiments, the direct repeat sequence according to Table 2 or Table 8 comprises a sequence containing uracil at one or more locations shown as thymine in the corresponding sequences of Table 2 or Table 8.

[0160] In some embodiments, the direct repeat comprises only one copy of the sequence repeated in the endogenous CRISPR array. In some embodiments, the direct repeat is a full-length sequence adjacent to (e.g., flanking) one or more spacer sequences found in the endogenous CRISPR array. In some embodiments, the direct repeat is a part (e.g., processed part) of a full-length sequence adjacent to (e.g., flanking) one or more spacer sequences found in the endogenous CRISPR array.

[0161] Spacers and direct repeats The spacer length of the RNA guide may range from about 15 to 55 nucleotides. The spacer length of the RNA guide may range from about 20 to 45 nucleotides. In some embodiments, the spacer length of the RNA guide is at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, or at least 22 nucleotides. In some embodiments, the spacer length is 15 - 17 nucleotides, 15 - 23 nucleotides, 16 - 22 nucleotides, 17 - 20 nucleotides, 20 - 24 nucleotides (e.g., 20, 21, 22, 23, or 24 nucleotides), 23 - 25 nucleotides (e.g., 23, 24, or 25 nucleotides), 24 - 27 nucleotides, 27 - 30 nucleotides, 30 - 45 nucleotides (e.g., 30, 31, 32, 33, 34, 35, 40, or 45 nucleotides), 30 or 35 - 40 nucleotides, 41 - 45 nucleotides, 45 - 50 nucleotides, or more.

[0162] In some embodiments, the direct repeat length of the RNA guide is at least 16 nucleotides, or 16 - 20 nucleotides (e.g., 16, 17, 18, 19, or 20 nucleotides). In some embodiments, the direct repeat length of the RNA guide is about 19 - about 40 nucleotides.

[0163] Exemplary directory repeat arrays (e.g., the directory repeat array of a pre-crRNA (e.g., unprocessed crRNA) or the directory repeat array of a mature crRNA (e.g., processed crRNA)) are shown in Table 2. See also Table 8.

[0164] [Table 2]

[0165] In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 57. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 2, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 58. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 3, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 59.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 60. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 14, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 128.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 15, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 63. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 17, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 130. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 18, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 70.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 21, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 72. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 22, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 73. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 23, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 74.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 24, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 63. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 27, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 76. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 28, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 77.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 29, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 139. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 31, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 58. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 32, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 80. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 32, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 80. It comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 35. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 77, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 77. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 36, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 139. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 38, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 80.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 39, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 58. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 41, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 83. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 42, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 84.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 44, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 86. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 45, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 130. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 46, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 84.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 47, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 87. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 48, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 88. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 51, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 84.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 53, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 84. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 55, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 88. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 56, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 90.

[0166] In some embodiments, the RNA guide comprises the directory repeat sequence described in FIG. 3. For example, in some embodiments, the RNA guide comprises the directory repeat of the consensus sequence shown in FIG. 3 or a portion of the consensus sequence shown in FIG. 3. In some embodiments, the RNA guide comprises a directory repeat having a sequence described as X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224), where X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T. For example, in some embodiments, the RNA guide comprises a directory repeat having a sequence described as ATTGTTGDA (SEQ ID NO: 225). In some embodiments, SEQ ID NO: 224 is proximate to the 5' end of the directory repeat. In some embodiments, SEQ ID NO: 225 is proximate to the 5' end of the directory repeat. In some embodiments, the RNA guide comprises a directory repeat having a sequence described as X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226), where X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C. For example, in some embodiments, the RNA guide comprises a directory repeat having a sequence described as TTTTWTARG (SEQ ID NO: 227). In some embodiments, the RNA guide comprises a directory repeat having a sequence described as X1X2X3AC (SEQ ID NO: 228), where X1 is A or C or G, X2 is C or A, and X3 is A or C. For example, in some embodiments, the RNA guide comprises a directory repeat having a sequence described as ACAAC (SEQ ID NO: 229). In some embodiments, SEQ ID NO: 228 is proximate to the 3' end of the directory repeat. In some embodiments, SEQ ID NO: 229 is proximate to the 3' end of the directory repeat.

[0167] In some embodiments, the RNA-guided spacer binds to a target nucleic acid adjacent to the PAM sequence of Table 3. For example, in some embodiments, the complex of the effector and the RNA guide disclosed herein binds to a target nucleic acid adjacent to the PAM sequence shown in Table 3.

[0168] [Table 3]

[0169] [Table 4]

[0170] In some embodiments, the RNA guide further comprises a tracrRNA. In some embodiments, the tracrRNA is not required (e.g., the tracrRNA is optional). In some embodiments, the tracrRNA is part of the non-coding sequence shown in Table 9. For example, in some embodiments, the tracrRNA is the sequence of Table 4.

[0171] [Table 5]

[0172] [Table 6]

[0173] [Table 7]

[0174] [Table 8]

[0175] [Table 9]

[0176]

Table 10

[0177] In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 1, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 152, SEQ ID NO: 153, or SEQ ID NO: 154. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 2, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 155, SEQ ID NO: 156, SEQ ID NO: 157, or SEQ ID NO: 158. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 3, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 159, SEQ ID NO: 160, or SEQ ID NO: 161.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 14, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 162. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 17, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, or SEQ ID NO: 166. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 18, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 167 or SEQ ID NO: 168.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 21, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 169, SEQ ID NO: 170, or SEQ ID NO: 171. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 22, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, or SEQ ID NO: 175. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 23, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 176, SEQ ID NO: 177, SEQ ID NO: 178, or SEQ ID NO: 179.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 27, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 180 or SEQ ID NO: 181. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 29, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 182, SEQ ID NO: 183, or SEQ ID NO: 184. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 31, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 185, SEQ ID NO: 186, SEQ ID NO: 187, or SEQ ID NO: 188.In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 32, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 189 or SEQ ID NO: 190. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 36, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 182, SEQ ID NO: 183, or SEQ ID NO: 184. In some embodiments, the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 38, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 189 or SEQ ID NO: 190.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 39, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 185, SEQ ID NO: 186, SEQ ID NO: 187, or SEQ ID NO: 188. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 41, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%,... It comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 43, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 197, SEQ ID NO: 198, or SEQ ID NO: 199. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 44, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 195 or SEQ ID NO: 196. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 45, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, or SEQ ID NO: 166.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 48, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 200, SEQ ID NO: 201, or SEQ ID NO: 202. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 52, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 197, SEQ ID NO: 198, or SEQ ID NO: 199. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 55, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 200, SEQ ID NO: 201, or SEQ ID NO: 202.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 56, and the tracrRNA sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 203 or SEQ ID NO: 204.

[0178] The RNA guide sequence may be modified in such a way that it allows for the formation of the CRISPR complex and successful binding to the target, but at the same time does not allow for successful nuclease activity (i.e., no nuclease activity / does not cause indels). Such a modified guide sequence is referred to as a "dead guide" or "dead guide sequence". Such a dead guide or dead guide sequence may be catalytically inactive or conformationally inactive in terms of nuclease activity. The dead guide sequence is typically shorter than each guide sequence that gives rise to active RNA cleavage. In some embodiments, the dead guide is 5%, 10%, 20%, 30%, 40%, or 50% shorter than each guide RNA having nuclease activity. The dead guide sequence of the RNA guide may be 13-15 nucleotides in length (e.g., 13, 14, or 15 nucleotides in length), 15-19 nucleotides in length, or 17-18 nucleotides in length (e.g., 17 nucleotides in length).

[0179] Accordingly, in one aspect, the present disclosure provides a non-naturally occurring or engineered CRISPR system comprising a functional CLUST.091979CRISPR effector as described herein and an RNA guide, where the RNA guide comprises a dead guide sequence, and thus the RNA guide has the ability to hybridize to a target sequence such that the CRISPR system is directed to a target genomic locus in a cell without detectable cleavage activity. A detailed description of dead guides is provided, for example, in International Publication No. WO 2016 / 094872, which is incorporated herein by reference in its entirety.

[0180] Inducible RNA guide The RNA guide can be engineered as a component of an inducible system. The inducible nature of this system allows for spatiotemporal control of gene editing or gene expression. In some embodiments, stimuli for the inducible system include, for example, electromagnetic radiation, acoustic energy, chemical energy, and / or thermal energy.

[0181] In some embodiments, transcription of the RNA guide can be regulated by inducible promoters such as transcriptional activation under tetracycline or doxycycline control (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. Other examples of inducible systems include, for example, small molecule 2-hybrid transcriptional activation systems (FKBP, ABA, etc.), light-inducible systems (phytochrome, LOV domain, or cryptochrome), or light-inducible transcriptional effectors (LITE). These inducible systems are described, for example, in International Publication No. WO 2016 / 205764 and U.S. Patent No. 8,795,965, which are incorporated herein by reference in their entireties, respectively.

[0182] Chemical modifications Chemical modifications can be applied to the phosphate backbone, sugar, and / or bases of guide RNAs. Backbone modifications such as phosphorothioates modify the charge on the phosphate backbone and are useful for delivery of oligonucleotides and nuclease resistance (see, e.g., Eckstein, “Phosphorothioates, essential components of therapeutic oligonucleotides,” Nucl. Acid Ther., 24 (2014), pp. 374-387); sugar modifications such as 2'-O-methyl (2'-OMe), 2'-F, and locked nucleic acid (LNA) enhance both base pairing and nuclease resistance (see, e.g., Allerson et al. “Fully 2‘-modified oligonucleotide duplexes with improved in vitro potency and stability compared to unmodified small interfering RNA,” J. Med. Chem., 48.4 (2005): 901-904). Chemically modified bases, particularly such as 2-thiouridine or N6-methyladenosine, can enable either stronger or weaker base pairing (see, e.g., Bramsen et al., “Development of therapeutic-grade small interfering RNAs by chemical engineering” Front. Genet., 2012 Aug 20; 3:154). In addition, RNA is suitable for conjugation at both the 5' and 3' ends with various functional moieties including fluorescent dyes, polyethylene glycol, or proteins.

[0183] A wide variety of modifications can be applied to chemically synthesized RNA guide molecules. For example, modifying oligonucleotides with 2'-OMe to improve nuclease resistance can change the binding energy of Watson-Crick base pairing. Furthermore, 2'-OMe modifications can affect how oligonucleotides interact with transfection reagents, proteins, or any other molecule in the cell. The effects of these modifications can be determined by empirical testing.

[0184] In some embodiments, the RNA guide comprises one or more phosphorothioate modifications. In some embodiments, the RNA guide comprises one or more locked nucleic acids for the purpose of enhancing base pairing and / or increasing nuclease resistance.

[0185] For an overview of these chemical modifications, see, for example, Kelley et al., “Versatility of chemically synthesized guide RNAs for CRISPR-Cas9 genome editing,” J. Biotechnol. 2016 Sep 10;233:74-83; International Publication No. WO 2016 / 205764; and U.S. Patent No. 8,795,965, each of which is incorporated by reference in its entirety.

[0186] Sequence Modifications The sequences and lengths of the RNA guides, tracrRNAs, and crRNAs described herein can be optimized. In some embodiments, the optimized length of the RNA guide may be determined by identifying processed forms of tracrRNA and / or crRNA or by empirical length studies for RNA guides of crRNAs.

[0187] The RNA guide can also include one or more aptamer sequences. An aptamer is an oligonucleotide or peptide molecule capable of binding to a specific target molecule. The aptamer may be specific for a gene effector, gene activator, or gene repressor. In some embodiments, the aptamer is specific for a protein, which in turn may be specific for and recruit / bind to a specific gene effector, gene activator, or gene repressor. The effector, activator, or repressor can be present in the form of a fusion protein. In some embodiments, the RNA guide has two or more aptamer sequences specific for the same adapter protein. In some embodiments, the two or more aptamer sequences are specific for different adapter proteins. Examples of adapter proteins include, for example, MS2, PP7, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. Thus, in some embodiments, the aptamer is selected from binding proteins that specifically bind to any one of the adapter proteins as described herein. In some embodiments, the aptamer sequence is an MS2 loop. For a detailed description of aptamers, reference can be made, for example, to Nowak et al., “Guide RNA engineering for versatile Cas9 functionality,” Nucl. Acid. Res., 2016 Nov 16;44(20):9555-9564; and International Publication No. WO 2016 / 205764 (each of which is incorporated herein by reference in its entirety).

[0188] Guide: Target Sequence Matching Requirement In the CRISPR system, the degree of complementarity between the guide sequence and its corresponding target sequence may be about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100%. To reduce off-target interactions, for example, to reduce guides that interact with target sequences with low complementarity, mutations may be introduced into the CRISPR system so that the CRISPR system can distinguish between the target sequence and off-target sequences having a complementarity higher than 80%, 85%, 90%, or 95%. In some embodiments, the degree of complementarity is between 80% and 95%, for example, about 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, or 95% (e.g., to distinguish between a target having 18 nucleotides and an off-target of 18 nucleotides having 1, 2, or 3 mismatches). Thus, in some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is higher than 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 99.9%. In some embodiments, the degree of complementarity is 100%.

[0189] In the art, it is known that perfect complementarity is not a requirement if there is sufficient complementarity to be functional. By introducing mismatches between the spacer sequence and the target sequence, such as one or more mismatches, such as 1 or 2 mismatches, including the position of the mismatches along the spacer / target, the cleavage efficiency can be exploited. The closer the mismatch, such as a double mismatch, is located towards the center (i.e., not at the 3' or 5' end), the greater the effect on the cleavage efficiency. Thus, the cleavage efficiency can be regulated by the choice of the mismatch position along the spacer sequence. For example, if less than 100% cleavage of the target is desired (e.g., in a cell population), one or two mismatches between the spacer and the target sequence may be introduced into the spacer sequence.

[0190] Methods of using the CRISPR system The CRISPR systems described herein have a wide variety of utilities, including modification (e.g., deletion, insertion, translocation, inactivation, or activation) of target polynucleotides in a very large number of cell types. This CRISPR system has a wide range of applications, for example, in DNA / RNA detection (e.g., specific high sensitivity enzymatic reporter unlocking (SHERLOCK)), nucleic acid tracking and labeling, enrichment assays (extraction of a desired sequence from a background), detection of circulating tumor DNA, preparation of next-generation libraries, drug screening, disease diagnosis and prognosis determination, and treatment of various genetic disorders.

[0191] DNA / RNA detection In one aspect, the CRISPR systems described herein can be used in DNA / RNA detection. By reprogramming a single effector RNA-guided DNase with a CRISPR RNA (crRNA), a platform for specific single-stranded DNA (ssDNA) sensing can be provided. Upon recognition of its DNA target, the activated type V single effector DNA-guided DNase is involved in "collateral" cleavage of neighboring non-target ssDNA. This collateral cleavage activity programmed by the crRNA enables the CRISPR system to detect the presence of specific DNA by non-specific degradation of labeled ssDNA.

[0192] In DNA detection applications, collateral ssDNA activity can be combined with a reporter, such as in a method called DNA Endonuclease-Targeted CRISPR trans reporter (DETECTR), which achieves attomolar concentration DNA detection sensitivity (see, e.g., Chen et al., Science, 360(6387):436-439, 2018, which is incorporated herein by reference in its entirety). One application of the enzymes described herein is the degradation of non-specific ssDNA in an in vitro environment. “Reporter” ssDNA molecules linked to fluorophores and quenchers can also be added to this in vitro system along with an unknown DNA sample (either single-stranded or double-stranded). Recognition of the target sequence in the unknown DNA fragment causes the effector complex to cleave the reporter ssDNA, resulting in a fluorescent readout.

[0193] In other embodiments, the SHERLOCK method (Specific High Sensitivity Enzymatic Reporter UnLOCKing) also provides an in vitro nucleic acid detection platform with attomolar concentration (or single molecule) sensitivity based on nucleic acid amplification and collateral cleavage of reporter ssDNA, enabling real-time detection of targets. The use of CRISPR in SHERLOCK is described in detail, for example, in Gootenberg, et al. “Nucleic acid detection with CRISPR-Cas13a / C2c2,” Science, 356(6336):438-442 (2017), which is incorporated herein by reference in its entirety.

[0194] In some embodiments, the CRISPR systems described herein can be used in multiplexed error-robust fluorescence in situ hybridization (MERFISH). Such methods are described, for example, in Chen et al., “Spatially resolved, highly multiplexed RNA profiling in single cells,” Science, 2015 Apr 24;348(6233):aaa6090, which is incorporated herein by reference in its entirety.

[0195] Tracking and Labeling of Nucleic Acids Cellular processes rely on networks of molecular interactions among proteins, RNAs, and DNAs. Accurate detection of protein-DNA and protein-RNA interactions is key to understanding such processes. In vitro proximity labeling techniques use affinity tags combined with reporter groups, such as photoactivatable groups, to label polypeptides and RNAs near a protein or RNA of interest in vitro. After ultraviolet irradiation, the photoactivatable group reacts with proteins and other molecules in close proximity to the tagged molecule, thereby labeling them. The labeled interacting molecules can then be recovered and identified. This RNA targeting effector protein can be used, for example, to target probes to selected RNA sequences. Such applications can also be applied to in vivo imaging of diseases or cell types that are difficult to culture in animal models. Methods for tracking and labeling nucleic acids are described, for example, in U.S. Patent No. 8,795,965; International Publication No. WO 2016 / 205764; and International Publication No. WO 2017 / 070605, each of which is incorporated herein by reference in its entirety.

[0196] High-Throughput Screening The CRISPR systems described herein can be used for the preparation of next-generation sequencing (NGS) libraries. For example, to create a cost-effective NGS library, the CRISPR system can be used to disrupt the coding sequence of a target gene, and at the same time, clones transfected with the CRISPR effector can be screened by next-generation sequencing (e.g., on an Ion Torrent PGM system). For a detailed description of methods for preparing NGS libraries, reference can be made, for example, to Bell et al., “A high-throughput screening strategy for detecting CRISPR-Cas9 induced mutations using next-generation sequencing,” BMC Genomics, 15.1 (2014): 1002 (which is incorporated herein by reference in its entirety).

[0197] Engineered cells Microorganisms (e.g., Escherichia coli (E. coli), yeast, and microalgae) are widely used in synthetic biology. The development of synthetic biology has a wide range of utilities, including various clinical applications. For example, a programmable CRISPR system can be used to split a protein of a toxic domain for targeted cell death using, for example, cancer-related RNA as a target transcript. Further, a fusion complex with an appropriate effector such as, for example, a kinase or an enzyme can affect a pathway involving protein-protein interactions in a synthetic biological system.

[0198] In some embodiments, an RNA guide sequence targeting a phage sequence can be introduced into a microorganism. Accordingly, the present disclosure also provides a method of “vaccinating” a microorganism (e.g., a production strain) against phage infection.

[0199] In some embodiments, by engineering a microorganism using the CRISPR system provided herein, for example, the yield can be improved or the fermentation efficiency can be improved. For example, by engineering a microorganism such as yeast using the CRISPR system described herein, biofuels or biopolymers can be produced from fermentable sugars, or plant-derived lignocellulose derived from agricultural waste as a fermentable sugar source can be decomposed. More specifically, the methods described herein can be used to modify the expression of endogenous genes required for biofuel production and / or to modify endogenous genes that may interfere with biofuel synthesis. These methods of engineering microorganisms are described, for example, in Verwaal et al., “CRISPR / Cpf1 enables fast and simple genome editing of Saccharomyces cerevisiae,” Yeast, 2017 Sep 8. doi:10.1002 / yea.3278; and Hlavova et al., “Improving microalgae for biotechnology-from genetics to synthetic biology,” Biotechnol. Adv., 2015 Nov 1;33:1194-203 (each of which is incorporated herein by reference in its entirety).

[0200] In some embodiments, the CRISPR system provided herein can be used to engineer eukaryotic cells or eukaryotes. For example, the CRISPR system described herein can be used to engineer eukaryotic cells not limited to plant cells, fungal cells, mammalian cells, reptilian cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, or human cells. In some embodiments, the eukaryotic cells are in an in vitro culture. In some embodiments, the eukaryotic cells are in vivo. In some embodiments, the eukaryotic cells are ex vivo.

[0201] In some embodiments, the cells are derived from a cell line. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, 293T, MF7, K562, HeLa, and their transgenic variants. Cell lines are available from a variety of sources known to those of skill in the art (see, e.g., American Type Culture Collection (ATCC) (Manassas, Va.)). In some embodiments, cells transfected with one or more nucleic acids (e.g., a nuclease polypeptide coding vector and an RNA guide) are used to establish a new cell line comprising one or more vector-derived sequences, thereby establishing a new cell line comprising a modification of a target nucleic acid or target locus. In some embodiments, the cells are immortal or immortalized cells.

[0202] In some embodiments, the cells are primary cells. In some embodiments, the cells are stem cells such as totipotent stem cells (e.g., omnipotent), pluripotent stem cells, multipotent stem cells, oligopotent stem cells, or unipotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs) or are derived from iPSCs. In some embodiments, the cells are differentiated cells. For example, in some embodiments, the differentiated cells are muscle cells (e.g., myocytes), adipocytes (e.g., adipocytes), bone cells (e.g., osteoblasts, osteocytes, osteoclasts), blood cells (e.g., monocytes, lymphocytes, neutrophils, eosinophils, basophils, macrophages, erythrocytes, or platelets), nerve cells (e.g., neurons), epithelial cells, immune cells (e.g., lymphocytes, neutrophils, monocytes, or macrophages), hepatocytes (e.g., hepatocytes), fibroblasts, or germ cells. In some embodiments, the cells are terminally differentiated cells. For example, in some embodiments, the terminally differentiated cells are nerve cells, adipocytes, cardiomyocytes, skeletal muscle cells, epidermal cells, or intestinal cells. In some embodiments, the cells are mammalian cells, e.g., human cells or murine cells. In some embodiments, the murine cells are derived from wild-type mice, immunosuppressed mice, or disease-specific mouse models.

[0203] Gene drive A gene drive is a phenomenon in which there is a favorable bias in the genetic traits of a specific gene or group of genes. A gene drive can be constructed using the CRISPR system described herein. For example, the CRISPR system can be designed to target and disrupt a specific allele of a gene, causing the cell to copy a second allele and fix the sequence. Due to this copying, the first allele will be converted to the second allele, increasing the likelihood that the second allele will be inherited by the offspring. For detailed methods on how to construct a gene drive using the CRISPR system described herein, see, for example, Hammond et al., “A CRISPR-Cas9 gene drive system targeting female reproduction in the malaria mosquito vector Anopheles gambiae,” Nat. Biotechnol., 2016 Jan; 34(1): 78-83 (which is hereby incorporated by reference in its entirety).

[0204] Pooled screening As described herein, pooled CRISPR screening is a powerful tool for identifying genes involved in biological mechanisms such as cell proliferation, drug resistance, and viral infection. Cells are transduced en masse with a library of RNA-guided coding vectors as described herein, and the distribution of gRNAs is measured before and after the application of a selective challenge. Pooled CRISPR screens function well for mechanisms that affect cell survival and proliferation and can be extended to the measurement of the activity of individual genes (e.g., by using engineered reporter cell lines). For arrayed CRISPR screens where only one gene is targeted at a time, it becomes possible to use RNA-seq as a readout. In some embodiments, the CRISPR system as described herein can be used for single-cell CRISPR screens. For a detailed description of pooled CRISPR screening, see, e.g., Datlinger et al., “Pooled CRISPR screening with single-cell transcriptome read-out,” Nat. Methods., 2017 Mar;14(3):297-301, which is incorporated herein by reference in its entirety.

[0205] Saturation mutagenesis (“bashing”) The CRISPR systems described herein can be used for in situ saturating mutagenesis. In some embodiments, a pooled RNA guide library can be used to perform in situ saturating mutagenesis on a particular gene or regulatory element. Such methods can reveal the definitive minimal features and the individual vulnerabilities of those genes or regulatory elements (e.g., enhancers). These methods are described, for example, in Canver et al., “BCL11A enhancer dissection by Cas9-mediated in situ saturating mutagenesis,” Nature, 2015 Nov 12;527(7577):192-7 (which is incorporated herein by reference in its entirety).

[0206] Therapeutic applications In some embodiments, the CRISPR systems described herein can be used to edit a target nucleic acid to modify the target nucleic acid (e.g., by inserting, deleting, or mutating one or more amino acid residues). For example, in some embodiments, the CRISPR systems described herein include an exogenous donor template nucleic acid (e.g., a DNA molecule or an RNA molecule) that contains a desired nucleic acid sequence. Upon resolution of a cleavage event induced by the CRISPR systems described herein, the cellular machinery can utilize the exogenous donor template nucleic acid when repairing and / or resolving the cleavage event. Alternatively, the cellular machinery can utilize an endogenous template when repairing and / or resolving the cleavage event. In some embodiments, the CRISPR systems described herein can be used to modify a target nucleic acid to result in an insertion, deletion, and / or point mutation. In some embodiments, the insertion is a scarless insertion (i.e., insertion of the intended nucleic acid sequence into the target nucleic acid does not result in additional unintended nucleic acid sequences upon resolution of the cleavage event). The donor template nucleic acid can be a double-stranded or single-stranded nucleic acid molecule (e.g., DNA or RNA). Methods for designing the exogenous donor template nucleic acid are described, for example, in WO 2016 / 094874 (the entire content of which is hereby expressly incorporated by reference).

[0207] In another aspect, the disclosure provides for the use of the systems described herein in a method selected from the group consisting of RNA sequence-specific interference; RNA sequence-specific gene regulation; screening of RNA, RNA products, lncRNA, non-coding RNA, nuclear RNA, or mRNA; mutagenesis; inhibition of RNA splicing; fluorescence in situ hybridization; breeding; induction of cell dormancy; induction of cell cycle arrest; decrease in cell growth and / or cell proliferation; induction of cell anergy; induction of cell apoptosis; induction of cell necrosis; induction of cell death; or induction of programmed cell death.

[0208] The CRISPR systems described herein can have various therapeutic applications. In some embodiments, the novel CRISPR systems can be used to treat various diseases and disorders, such as genetic disorders (e.g., single gene diseases) or diseases that can be treated by nuclease activity (e.g., Pcsk9 targeting or BCL11a targeting). In some embodiments, the methods described herein are used for the treatment of a subject, such as a mammalian subject, e.g., a human patient. The mammalian subject can also be a domesticated mammal such as a dog, cat, horse, monkey, rabbit, rat, mouse, female cow, goat, or sheep.

[0209] The method can include the condition or disease being infectious, where the infectious agent is selected from the group consisting of human immunodeficiency virus (HIV), herpes simplex virus type 1 (HSV1), and herpes simplex virus type 2 (HSV2).

[0210] In one aspect, the CRISPR systems described herein can be used to treat diseases caused by the overexpression of RNA, toxic RNA, and / or mutant RNA (e.g., splicing defects or truncations). For example, the expression of toxic RNA can be associated with the formation of nuclear inclusions and late-onset degenerative changes in the brain, heart, or skeletal muscle. In some embodiments, the disorder is myotonic dystrophy. In myotonic dystrophy, the main pathogenic effect of toxic RNA is to sequester binding proteins and impair the regulation of alternative splicing (see, e.g., Osborne et al., “RNA-dominant diseases,” Hum. Mol. Genet., 2009 Apr 15;18(8):1471-81). Myotonic dystrophy (dystrophia myotonica (DM)) is of particular interest to geneticists because it gives rise to a very broad range of clinical features. The classical form of DM, now called DM type 1 (DM1), is caused by an expansion of CTG repeats in the 3' untranslated region (UTR) of the gene DMPK, which encodes a cytosolic protein kinase. The CRISPR systems described herein can target either the overexpressed RNA or toxic RNA, such as the DMPK gene, or alternative splicing misregulated in DM1 skeletal muscle, heart, or brain.

[0211] The CRISPR systems described herein can also target trans-acting mutations that affect RNA-dependent functions that cause various diseases such as Prader-Willi syndrome, spinal muscular atrophy (SMA), dyskeratosis congenita, etc. The list of diseases that can be treated using the CRISPR systems described herein is summarized in Cooper et al., “RNA and disease,” Cell, 136.4 (2009):777-793 and WO 2016 / 205764 (each of which is incorporated herein by reference in its entirety).

[0212] The CRISPR systems described herein can also be used for the treatment of various tauopathies, including, for example, primary age-related tauopathy (PART) / neurofibrillary tangle (NFT)-predominant senile dementia (with NFTs similar to those seen in Alzheimer's disease (AD), but without plaques), boxing dementia (chronic traumatic encephalopathy), and progressive supranuclear palsy, as well as primary and secondary tauopathies. A useful list of tauopathies and methods of treating these diseases is described, for example, in International Publication No. WO 2016 / 205764 (incorporated herein by reference in its entirety).

[0213] The CRISPR systems described herein can also be used to target splicing defects and mutations that disrupt cis-acting splicing codes that can cause diseases. These diseases include, for example, motor neuron degenerative diseases (such as spinal muscular atrophy) caused by deletions in the SMN1 gene, Duchenne muscular dystrophy (DMD), frontotemporal dementia, and parkinsonism associated with chromosome 17 (FTDP-17), and cystic fibrosis.

[0214] The CRISPR systems described herein can further be used, particularly for antiviral activity against RNA viruses. The effector protein can target viral RNA using an appropriate RNA guide selected to target the viral RNA sequence.

[0215] Furthermore, in vitro RNA sensing assays can be used to detect specific RNA substrates. RNA-targeting effector proteins can be used for RNA-based sensing in living cells. An example of an application is, for example, diagnosis by sensing disease-specific RNA.

[0216] A detailed description of the therapeutic uses of the CRISPR systems described herein can be found, for example, in U.S. Patent No. 8,795,965, European Patent No. 3,009,511, International Publication No. WO 2016 / 205764, and International Publication No. WO 2017 / 070605, each of which is incorporated herein by reference in its entirety.

[0217] Application in Plants The CRISPR systems described herein have a wide variety of utilities in plants. In some embodiments, the CRISPR systems can be used to engineer the genomes of plants (e.g., to improve production, to produce products with desired post-translational modifications, or to introduce genes for the production of industrial products). In some embodiments, the CRISPR systems can be used to introduce desired traits (e.g., with or without heritable modifications to the genome) into plants or to regulate the expression of endogenous genes in plant cells or whole plants.

[0218] In some embodiments, the present CRISPR systems can be used to identify, edit, and / or silence genes encoding specific proteins, such as allergen proteins (e.g., allergen proteins in peanuts, soybeans, lentils, peas, green beans, and fava beans). For a detailed description of methods for identifying, editing, and / or silencing genes encoding proteins, see, for example, Nicolaou et al., “Molecular diagnosis of peanut and legume allergy,” Curr. Opin. Allergy Clin. Immunol., 11(3):222-8 (2011), and International Publication No. WO 2016 / 205764, each of which is incorporated herein by reference in its entirety.

[0219] Delivery of the CRISPR System Through the present disclosure and the knowledge in the art, the CRISPR systems described herein, or components thereof, the nucleic acid molecules thereof, or the nucleic acid molecules encoding or providing such components can be delivered by various delivery systems such as vectors, e.g., plasmids, or viral delivery vectors. The CRISPR effectors and / or any RNAs (e.g., RNA guides) disclosed herein can be delivered using suitable vectors, e.g., plasmids, or viral vectors such as adeno-associated virus (AAV), lentivirus, adenovirus, and other viral vectors, or combinations thereof. The effector and one or more RNA guides can be packaged into one or more vectors, e.g., plasmids or viral vectors.

[0220] In some embodiments, the vector, e.g., plasmid or viral vector, is delivered to the target tissue by, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such delivery can be by either a single dose or multiple doses. Those skilled in the art will understand that the actual dosage delivered herein can vary widely depending on various factors including, but not limited to, the choice of vector, target cells, organism, tissue, general condition of the subject being treated, degree of transformation / modification sought, route of administration, mode of administration, and type of transformation / modification sought.

[0221] In certain embodiments, the delivery is by adenovirus, which can be in a single dose containing at least 1×10 5 particles (also referred to as particle units, pu) of adenovirus. In some embodiments, preferably the dose is at least about 1×10 6 particles, at least about 1×10 7 particles, at least about 1×10 8 particles, and at least about 1×10 9The particle is an adenovirus. Regarding the delivery method and dosage, for example, they are described in International Publication No. WO2016205764 pamphlet and U.S. Patent No. 8454972 specification (both of which are incorporated herein by reference in their entireties).

[0222] In some embodiments, the delivery is by plasmid. The dosage can be a number of plasmids sufficient to elicit a response. In some cases, a suitable amount of plasmid DNA in the plasmid composition may be about 0.1 to about 2 mg. The plasmid generally comprises: (i) a promoter; (ii) a sequence encoding a nucleic acid targeting CRISPR effector operably linked to the promoter; (iii) a selectable marker; (iv) an origin of replication; and (v) a transcription terminator downstream of and operably linked to (ii). The plasmid can also encode the RNA components of the CRISPR complex, but alternatively, one or more of these can be encoded by different vectors. The dosing frequency is within the purview of a medical or veterinary practitioner (e.g., a physician, veterinarian), or one of ordinary skill in the art.

[0223] In another embodiment, the delivery is by liposomes or lipofectin formulations, etc., which can be prepared by methods known to those of ordinary skill in the art. Such methods are described, for example, in International Publication No. WO2016205764 pamphlet and U.S. Patent No. 5593972 specification; No. 5589466 specification; and No. 5580859 specification (each of which is incorporated herein by reference in its entirety).

[0224] In some embodiments, the delivery is by nanoparticles or exosomes. For example, exosomes have been shown to be particularly useful for RNA delivery.

[0225] A further means of introducing one or more components of the CRISPR system described herein into cells is by use of a cell-penetrating peptide (CPP). In some embodiments, the cell-penetrating peptide is linked to a CRISPR effector. In some embodiments, the CRISPR effector and / or RNA guide is coupled to one or more CPPs and transported into the cell (e.g., a plant protoplast). In some embodiments, the CRISPR effector and / or one or more RNA guides are encoded by one or more circular or linear DNA molecules that are coupled to one or more CPPs for cell delivery.

[0226] A CPP is a short-chain peptide of less than 35 amino acids derived from either a protein or chimeric sequence having the ability to transport biomolecules across the cell membrane in a receptor-independent manner. The CPP may be a cationic peptide, a peptide having a hydrophobic sequence, an amphipathic peptide, a peptide having a proline-rich antimicrobial sequence, and a chimeric or bipartite peptide. Examples of CPPs include, for example, Tat (which is the transcriptional activator protein required for viral replication by human immunodeficiency virus type 1), penetratin, the Kaposi fibroblast growth factor (FGF) signal peptide sequence, the integrin β3 signal peptide sequence, the polyarginine peptide Arg sequence, the guanine-rich molecular transporter, and the sweet arrow peptide. CPPs and methods of using them are described, for example, in Haellbrink et al., “Prediction of cell-penetrating peptides,” Methods Mol. Biol., 2015;1324:39-58; Ramakrishna et al., “Gene disruption by cell-penetrating peptide-mediated delivery of Cas9 protein and guide RNA,” Genome Res., 2014 Jun;24(6):1020-7; and International Publication No. WO 2016 / 205764, each of which is incorporated herein by reference in its entirety.

[0227] The various delivery methods for the CRISPR systems described herein are also described, for example, in U.S. Patent No. 8,795,965, European Patent No. 3,009,511, International Publication No. 2016 / 205764, and International Publication No. 2017 / 070605, each of which is incorporated herein by reference in its entirety.

Examples

[0228] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.

[0229] Example 1 - Identification of Components of the CLUST.091979 CRISPR-Cas System This protein family was identified using the computational methods described above. The CLUST.091979 system includes single effectors associated with CRISPR systems found in uncultured metagenomic sequences collected from environments not limited to the gut, bovine gut, human gut, sheep gut, terrestrial, feces, and mammalian digestive system environments (Table 5). Exemplary CLUST.091979 effectors include those shown in Tables 5 and 6 below. As shown in FIGS. 1A - 1L, the effector sequences set forth in SEQ ID NOs: 1 - 4, 14, 15, 17 - 19, 21 - 25, 27 - 33, 35 - 49, 51 - 56 were aligned to identify regions of sequence similarity. The bar graph shows sequence similarity, with the tallest bar indicating the residues with the greatest sequence similarity. Non-limiting regions of sequence similarity are shown in Table 7. The regions of sequence similarity indicate that the effectors disclosed herein are a family that has a representative conserved C-terminal RuvC domain of a nuclease.

[0230]

Table 11

[0231]

Table 12

[0232]

Table 13

[0233]

Table 14

[0234]

Table 15

[0235]

Table 16

[0236]

Table 17

[0237]

Table 18

[0238]

Table 19

[0239]

Table 20

[0240]

Table 21

[0241]

Table 22

[0242]

Table 23

[0243]

Table 24

[0244]

Table 25

[0245]

Table 26

[0246]

Table 27

[0247]

Table 28

[0248]

Table 29

[0249] Examples of directory repeat arrays and spacer lengths for these systems are shown in Table 8.

[0250]

Table 30

[0251]

Table 31

[0252]

Table 32

[0253]

Table 33

[0254]

Table 34

[0255]

Table 35

[0256] Example 2 - Identification of Trans-activating RNA Elements In addition to effector proteins and crRNAs, some CRISPR systems described herein may also include additional small RNAs that activate a robust enzymatic activity called trans-activating RNA (tracrRNA). Such tracrRNAs typically contain a complementary region that hybridizes to the crRNA. The crRNA-tracrRNA hybrid forms a complex with the effector, resulting in the activation of programmable enzymatic activity. ·The tracrRNA sequence can be identified by searching for genomic sequences flanking the CRISPR array for short sequence motifs homologous to the direct repeat portion of the crRNA. Search methods include exact or degenerate sequence matching for the full direct repeat (DR) or DR subsequences. For example, a DR of length n nucleotides can be broken down into a set of overlapping 6-10 nt kmers. These kmers can be aligned with the sequences flanking the CRISPR locus, and regions homologous to one or more kmer alignments can be identified as DR homology regions for experimental validation as tracrRNA. Alternatively, the RNA simultaneous fold free energy can be calculated for short kmer sequences from the full DR or DR subsequence and genomic sequences flanking elements of the CRISPR system. Flanking sequence elements with a low minimum free energy structure can be identified as DR homology regions for experimental validation as tracrRNA. ·The tracrRNA element occurs frequently in close proximity to the CRISPR-associated gene or CRISPR array. As an alternative to searching for DR homology regions to identify the tracrRNA element, non-coding sequences flanking the CRISPR effector or CRISPR array can be isolated by cloning or gene synthesis for direct experimental validation of the tracrRNA. ·Experimental validation of the tracrRNA element can be performed using small RNA sequencing of the host organism for a CRISPR system or synthetic sequence heterologously expressed in a non-native species. Alignment of small RNA sequences from the native genomic locus can be used to identify the expression RNA products containing the DR homology region and the canonical processing typical of the full tracrRNA element. ·The fully identified tracrRNA candidates by RNA sequencing can be verified in vitro or in vivo by expressing the crRNA and effector with or without the tracrRNA candidates and monitoring the activation of effector enzyme activity. ·In engineered constructs, the expression of tracrRNA can be driven by promoters including, but not limited to, the U6, U1, and H1 promoters for expression in mammalian cells or the J23119 promoter for expression in bacteria. ·In some examples, tracrRNA can be fused with crRNA and expressed as a single RNA guide. ·The system can include tracrRNA contained within the non-coding sequences listed in Table 9. For example, in some embodiments, the system includes tracrRNA described in any one of SEQ ID NOs: 152 to 204.

[0257]

Table 36

[0258]

Table 37

[0259]

Table 38

[0260]

Table 39

[0261]

Table 40

[0262]

Table 41

[0263]

Table 42

[0264] Example 3 - Identification of a Novel RNA Modulator of Enzyme Activity In addition to effector proteins and crRNAs, some CRISPR systems described herein may also include additional small RNAs, referred to herein as RNA modulators, for activating or regulating effector activity. · RNA modulators are predicted to occur very close to CRISPR - associated genes or CRISPR arrays. To identify and validate RNA modulators, non - coding sequences flanking the CRISPR effector or CRISPR array can be isolated by cloning or gene synthesis for direct experimental validation. · Experimental validation of RNA modulators can be performed using small RNA sequencing of host organisms for heterologously expressed CRISPR systems or synthetic sequences in non - native species. Alignment of small RNA sequences to the native genomic locus can be used to identify expressed RNA products and the processing of maturation that contain DR homology regions. · Candidate RNA modulators identified by RNA sequencing can be verified in vitro or in vivo by expressing the crRNA and effector, either alone or in combination with the candidate RNA modulator, and monitoring changes in effector enzyme activity. · In engineered constructs, RNA modulators can be driven by promoters including, but not limited to, the U6, U1, and H1 promoters for expression in mammalian cells, or the J23119 promoter for expression in bacteria. ·In some examples, the RNA modulator can be artificially fused with either crRNA, tracrRNA, or both, and expressed as a single RNA element.

[0265] Example 4 - Functional verification of the engineered CLUST.091979 CRISPR-Cas system After identifying the components of the CLUST.091979 CRISPR-Cas system, loci from the metagenomic source designated AUXO013988882 (SEQ ID NO: 1) and the metagenomic source designated SRR3181151 (SEQ ID NO: 4) were selected for functional verification.

[0266] DNA synthesis and effector library cloning To test the activity of the exemplary CLUST.091979 CRISPR-Cas system, the system was designed and synthesized using the pET28a(+) vector. Briefly, an E. coli codon-optimized nucleic acid sequence encoding the CLUST.091979 AUXO013988882 effector (SEQ ID NO: 1 shown in Table 6) and an E. coli codon-optimized nucleic acid sequence encoding the CLUST.091979 SRR3181151 effector (SEQ ID NO: 4 shown in Table 6) were synthesized (Genscript) and individually cloned into a custom expression system derived from pET-28a(+) (EMD-Millipore). The vector contained a nucleic acid encoding the CLUST.091979 effector under the control of the lac promoter and an E. coli ribosome binding sequence. The vector also contained an acceptor site for the CRISPR array library driven by the J23119 promoter following the open reading frame of the CLUST.091979 effector. As shown in Table 9, the non-coding sequence used for the CLUST.091979 AUXO013988882 effector (SEQ ID NO: 1) is described in SEQ ID NO: 98, and the non-coding sequence used for the CLUST.091979 SRR3181151 effector (SEQ ID NO: 4) is described in SEQ ID NO: 99. Additional conditions were tested where the CLUST.091979 effector was individually cloned into pET28a(+) without the non-coding sequence. See Figure 4A.

[0267] An oligonucleotide library synthesis (OLS) pool containing a "repeat-spacer-repeat" array was computationally designed, where "repeat" corresponds to the consensus direct repeat array found in the CRISPR array associated with the effector, and "spacer" corresponds to the sequence tiling the pACYC184 plasmid or the essential genes of Escherichia coli (E. coli). In particular, as shown in Table 8, the repeat sequence used for the CLUST.091979 AUXO013988882 effector (SEQ ID NO: 1) is described in SEQ ID NO: 57, and the repeat sequence used for the 091979 SRR3181151 effector (SEQ ID NO: 4) is described in SEQ ID NO: 60. The spacer length was determined by the mode of the spacer lengths found in the endogenous CRISPR arrays. Restriction sites were added to the repeat-spacer-repeat sequence to enable bidirectional cloning of the fragments into the aforementioned CRISPR array library acceptor site and unique PCR priming sites for the specific amplification of the repeat-spacer-repeat library from a larger pool.

[0268] Next, the repeat-spacer-repeat library was cloned into a plasmid using the Golden Gate assembly method. Briefly, the inventors first amplified each repeat-spacer-repeat from the OLS pool (Agilent Genomics) using unique PCR primers and pre-linearized the plasmid backbone using BsaI to reduce potential background. Both DNA fragments were purified with Ampure XP (Beckman Coulter) before being added to the Golden Gate assembly master mix (New England Biolabs) and incubated according to the manufacturer's instructions. The Golden Gate reaction was further purified and concentrated to achieve maximum transformation efficiency in subsequent steps of bacterial screening.

[0269] A plasmid library containing different repeat-spacer-repeat elements and CRISPR effectors was electroporated into E. Cloni electrocompetent E. coli (Lucigen) using a Gene Pulser Xcell® (Bio-rad) according to the protocol recommended by Lucigen. The library was either co-transformed with a purified pACYC184 plasmid or directly transformed into E. Cloni electrocompetent E. coli (Lucigen) containing pACYC184, seeded onto an agar medium containing chloramphenicol (Fisher), tetracycline (Alfa Aesar), and kanamycin (Alfa Aesar) in a BioAssay® dish (Thermo Fisher), and incubated at 37 °C for 10 - 12 hours. After estimating the approximate colony number to ensure sufficient library representation on the bacterial plate, the bacteria were harvested, and plasmid DNA was extracted using a QIAprep Spin Miniprep® kit (Qiagen) to create an "output library". Barcoded next-generation sequencing libraries were generated from both the "input library" before transformation and the "output library" after recovery by performing PCR using custom primers containing barcodes and sites compatible with Illumina sequencing chemistry, pooled, and loaded onto a Nextseq550 (Illumina) to evaluate the effectors. To ensure consistency, at least two independent biological replicates were performed for each screen. See Figure 4B.

[0270] Bacterial screen sequencing analysis Next-generation sequencing data of the screen input and output libraries was demultiplexed using Illumina bcl2fastq. Reads in the fastq files obtained for each sample contained CRISPR array elements for the screening plasmid library. The orientation of the array was determined using the direct repeat sequences of the CRISPR array, and the corresponding targets were determined by mapping the spacer sequences to the source (pACYC184 or E.Cloni) or negative control sequences (GFP). For each sample, the total number of reads (r a ) for each unique array element in a given plasmid library was counted and normalized as follows: (r a + 1) / total number of reads for all library array elements. The depletion score was calculated by dividing the normalized output read count for a given array element by the normalized input read count.

[0271] To identify specific parameters that result in enzyme activity and bacterial cell death, the inventors used next-generation sequencing (NGS) to quantify and compare the representation of individual CRISPR arrays (i.e., repeat-spacer-repeat) in the PCR products of the input and output plasmid libraries. The depletion rate of an array was defined as the normalized output read count divided by the normalized input read count. If the depletion rate was less than 0.3 (more than 3-fold depletion), the array was considered "strongly depleted" and indicated by a dashed line in Figures 5 and 8. When calculating the array depletion rate across biological replicates, the value of the maximum depletion rate for a given CRISPR array across all experiments was taken (i.e., strongly depleted arrays must be strongly depleted in all biological replicates). For each spacer target, a matrix was created that included the array depletion rate and the following characteristics: target strand, transcript targeting, ORI targeting, target sequence motif, flanking sequence motif, and target secondary structure. The extent to which different features in this matrix explained target depletion for the CLUST.091979 system was investigated.

[0272] Figures 5 and 8 show the degree of interference activity of the CLUST.091979 composition engineered with non-coding sequences by plotting the normalized ratio of sequencing reads in screen output versus screen input for a given target. The results are plotted for each DR transcription direction. In the functional screening of the composition, the active effector that formed a complex with the active RNA guide interfered with the ability of pACYC184, which confers E. coli resistance to chloramphenicol and tetracycline, resulting in cell death and depletion of spacer elements within the pool. Comparing the results of the deep sequencing of the initial DNA library (screen input) and viable transformed E. coli (screen output) suggests specific target sequences and DR transcription directions that enable an active and programmable CRISPR system. The screen also shows that the effector complex is active in only one direction of the DR. Thus, the screen showed that the CLUST.091979 AUXO013988882 effector was active in the "forward" direction of the DR (5'-ACTA…AACT-[spacer]-3') (Figure 5), and that the CLUST.091979 SRR3181151 effector was active in the "reverse" direction of the DR (5'-CCTG…CAAC-[spacer]-3') (Figure 8).

[0273] Figures 6A and 6B respectively show the positions of strongly depleted targets for the CLUST.091979 AUXO013988882 effector (+ non-coding sequence) targeting pACYC184 and the essential genes of Escherichia coli (E. coli) E.Cloni. Similarly, Figures 9A and 9B respectively show the positions of strongly depleted targets for the CLUST.091979 SRR3181151 effector targeting pACYC184 and the essential genes of Escherichia coli (E. coli) E.Cloni. The adjacent sequences of the depleted targets were analyzed to determine the PAM sequences of CLUST.091979 AUXO013988882 and CLUST.091979 SRR3181151. The WebLogo representations (Crooks et al., Genome Research 14:1188-90, 2004) of the PAM sequences of CLUST.091979 AUXO013988882 and CLUST.091979 SRR3181151 are shown in Figures 7 and 10 respectively, and position "20" corresponds to the nucleotide adjacent to the 5' end of the target.

[0274] Therefore, multiple effectors of CLUST.091979 CRISPR-Cas show activity in vivo.

[0275] Example 5 - Targeting of Mammalian Genes by CLUST.091979 This example describes the indel evaluation for multiple targets using nucleases from CLUST.091979 introduced into mammalian cells by transient transfection.

[0276] The effectors of SEQ ID NO: 4, SEQ ID NO: 8, and SEQ ID NO: 10 were cloned into the pcda3.1 backbone (Invitrogen). Next, the plasmid was maxiprepped and diluted to 1 μg / μL. For the preparation of RNA guides, the dsDNA fragment encoding crRNA was derived by an ultramer containing a scaffold of the target sequence and the U6 promoter. The ultramer was resuspended in 10 mM Tris·HCl at pH 7.5 to a final stock concentration of 100 μM. Subsequently, the working stock was diluted to 10 μM and used as a template for the PCR reaction again using 10 mM Tris·HCl. The amplification of crRNA was performed in a 50 μL reaction using the following components: 0.02 μl of the aforementioned template, 2.5 μl of the forward primer, 2.5 μl of the reverse primer, 25 μL of NEB HiFi polymerase, and 20 μl of water. The cycling conditions were 1× (30 seconds at 98°C), 30× (10 seconds at 98°C, 15 seconds at 67°C), 1× (2 minutes at 72°C). The PCR product was cleaned up by 1.8X SPRI treatment and normalized to 25 ng / μL. The prepared crRNA sequences and their corresponding target sequences are shown in Table 10. The direct repeat sequences of the mature crRNAs of SEQ ID NO: 205, SEQ ID NO: 207, SEQ ID NO: 252, SEQ ID NO: 254, SEQ ID NO: 256, SEQ ID NO: 258, SEQ ID NO: 260, SEQ ID NO: 262, SEQ ID NO: 264, SEQ ID NO: 266, SEQ ID NO: 268, SEQ ID NO: 270, SEQ ID NO: 272, SEQ ID NO: 274, and SEQ ID NO: 276 are described in SEQ ID NO: 60. The direct repeats of the mature crRNAs of SEQ ID NO: 209 and SEQ ID NO: 214 are described in SEQ ID NO: 62. The direct repeats of the mature crRNAs of SEQ ID NO: 211, SEQ ID NO: 278, SEQ ID NO: 280, SEQ ID NO: 282, SEQ ID NO: 284, SEQ ID NO: 286, and SEQ ID NO: 288 are described in SEQ ID NO: 213.

[0277]

Table 43

[0278]

Table 44

[0279] Approximately 16 hours before transfection, 100 μl of 25,000 HEK293T cells in DMEM / 10% FBS + Pen / Strep were plated into each well of a 96-well plate. On the day of transfection, the cells were 70 - 90% confluent. For each well to be transfected, a mixture of 0.5 μl of Lipofectamine 2000 and 9.5 μl of Opti-MEM was prepared and then incubated at room temperature for 5 - 20 minutes (Solution 1). After incubation, the lipofectamine:OptiMEM mixture was added to another mixture containing 182 ng of effector plasmid, 14 ng of crRNA, and up to 10 μL of water (Solution 2). For the negative control, crRNA was not included in Solution 2. The mixtures of Solution 1 and Solution 2 were mixed up and down by pipetting and then incubated at room temperature for 25 minutes. After incubation, 20 μL of the mixture of Solution 1 and Solution 2 was dropped into each well of the 96-well plate containing the cells. 72 hours after transfection, 10 μL of TrypLE was added to the center of each well to trypsinize the cells and incubated for approximately 5 minutes. Next, 100 μL of D10 medium was added to each well and mixed to resuspend the cells. The cells were then spun down at 500 g for 10 minutes and the supernatant was discarded. QuickExtract buffer was added to 1 / 5 of the original cell suspension volume. The cells were incubated at 65 °C for 15 minutes, 68 °C for 15 minutes, and 98 °C for 10 minutes.

[0280] Samples for next-generation sequencing were prepared by two rounds of PCR. The first round (PCR1) was used to amplify a specific genomic region according to the target. The PCR1 products were purified by column purification. The second PCR round (PCR2) was performed to add Illumina adapters and indexes. Next, the reactions were pooled and purified by column purification. The sequencing run was performed with a 150-cycle NextSeq v2.5 mid-output or high-output kit.

[0281] Figures 11A, 11B, 11C, and 11D each show the percent indel of the AAVS1, VEGFA, and EMX1 target loci in HEK293T cells after transfection with the effector of SEQ ID NO: 4 or SEQ ID NO: 10. Bars reflect the average percent indel measured in two biological replicates. For the effectors of SEQ ID NO: 4 and SEQ ID NO: 10, the percent indel was higher than the percent indel of the negative control at each of the targets.

[0282] As shown in FIG. 11A, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 205 was active at the AAVS1 target of SEQ ID NO: 206, and the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 207 was active at the VEGFA target of SEQ ID NO: 208. As shown in FIG. 11B, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 252 was active at the AAVS1 target of SEQ ID NO: 253, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 254 was active at the AAVS1 target of SEQ ID NO: 255, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 256 was active at the AAVS1 target of SEQ ID NO: 257, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 258 was active at the AAVS1 target of SEQ ID NO: 259, and the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 274 was active at the AAVS1 target of SEQ ID NO: 275. As shown in FIG. 11B, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 260 was active at the EMX1 target of SEQ ID NO: 261.Similarly, as shown in FIG. 11B, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 262 is active against the VEGFA1 target of SEQ ID NO: 263, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 264 is active against the VEGFA1 target of SEQ ID NO: 265, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 266 is active against the VEGFA1 target of SEQ ID NO: 267, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 268 is active against the VEGFA1 target of SEQ ID NO: 269, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 270 is active against the VEGFA1 target of SEQ ID NO: 271, the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 272 is active against the VEGFA1 target of SEQ ID NO: 273, and the complex formed by the effector of SEQ ID NO: 4 and the crRNA of SEQ ID NO: 274 was active against the VEGFA1 target of SEQ ID NO: 275. The effector of SEQ ID NO: 4 utilized the 5'-TTTG-3' PAM for each of the targets in FIGS. 11A and 11B.

[0283] As shown in FIG. 11C, the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 209 is active at the AAVS1 target of SEQ ID NO: 210, the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 211 is active at the AAVS1 target of SEQ ID NO: 212, and the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 214 is active at the VEGFA target of SEQ ID NO: 215. As shown in FIG. 11D, the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 278 is active at the AAVS1 target of SEQ ID NO: 279, the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 280 is active at the AAVS1 target of SEQ ID NO: 281, the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 284 is active at the AAVS1 target of SEQ ID NO: 285, and the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 286 is active at the AAVS1 target of SEQ ID NO: 287. Similarly, as shown in FIG. 11D, the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 288 is active at the EMX1 target of SEQ ID NO: 289, and the complex formed by the effector of SEQ ID NO: 10 and the crRNA of SEQ ID NO: 282 is active at the VEGFA target of SEQ ID NO: 283. The effector of SEQ ID NO: 10 utilized 5'-ATTG-3' PAM and 5'-GTTA-3' PAM for the targets in FIGS. 11C and 11D.

[0284] This example suggests that nucleases of the CLUST.091979 family have activity in mammalian cells.

[0285] Other embodiments Although the present invention has been described with its detailed description, it should be understood that the foregoing description is illustrative and not intended to limit the scope of the present invention defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims. In certain embodiments, for example, the following items are provided. (Item 1) An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-Cas system of CLUST.091979, (a) a CRISPR-associated protein or a nucleic acid encoding said CRISPR-associated protein, wherein the CRISPR-associated protein comprises the amino acid sequence of SEQ ID NO: 241; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid comprising a CRISPR-Cas system, wherein the CRISPR-associated protein is capable of binding to the RNA guide and modifying the target nucleic acid sequence complementary to the spacer sequence. (Item 2) The system according to Item 1, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4, SEQ ID NO: 10, SEQ ID NO: 12, or SEQ ID NO: 14. (Item 3) An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-Cas system of CLUST.091979, (a) a CRISPR-associated protein or a nucleic acid encoding said CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1-56; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid comprising A CRISPR-Cas system in which the CRISPR-related protein can bind to the RNA guide and modify the target nucleic acid sequence complementary to the spacer sequence. (Item 4) The system according to item 3, wherein the CRISPR-related protein comprises at least one RuvC domain or at least one split RuvC domain. (Item 5) The CRISPR-related protein is one or more of the following sequences: (a) PX 1 X 2 X 3 X 4 F (SEQ ID NO: 216) (where X 1 is L or M or I or C or F, X 2 is Y or W or F, X 3 is K or T or C or R or W or Y or H or V, X 4 is I or L or M); (b) RX 1 X 2 X 3 L (SEQ ID NO: 217) (where X 1 is I or L or M or Y or T or F, X 2 is R or Q or K or E or S or T, X 3 is L or I or T or C or M or K); (c) NX 1 YX 2 (SEQ ID NO: 218) (where X 1 is I or L or F, X 2 is K or R or V or E); (d) KX 1 X 2 X 3 FAX 4 X 5 KD (SEQ ID NO: 219) (where X 1 is T or I or N or A or S or F or V, X 2 is I or V or L or S, X 3 is H or S or G or R, X 4 is D or S or E, X 5 is I or V or M or T or N); (e) LX 1 NX2( SEQ ID NO: 220) (where X 1 is G or S or C or T, X 2 is N or Y or K or S); (f) PX 1 X 2 X 3 X 4 SQX 5 DS (SEQ ID NO: 221) (where X 1 is S or P or A, X 2 is Y or S or A or P or E or Y or Q or N, X 3 is F or Y or H, X 4 is T or S, X 5 is M or T or I); (g) KX 1 X 2 VRX 3 X 4 QEX 5 H (SEQ ID NO: 222) (where X 1 is N or K or W or R or E or T or Y, X 2 is M or R or L or S or K or V or E or T or I or D, X 3 is L or R or H or P or T or K or P's Q or S or A, X 4 is G or Q or N or R or K or E or I or T or S or C, X 5 is R or W or Y or K or T or F or S or Q); and (h) X 1 NGX 2 X 3 X 4 DX 5 NX 6 X 7 X 8 N (SEQ ID NO: 223) (where X 1 is I or K or V or L, X 2 is L or M, X 3 is N or H or P, X 4 is A or S or C, X 5 is V or Y or I or F or T or N, X 6 is A or S, X 7 is S or A or P, X 8 is M or C or L or R or N or S or K or L) The system according to item 3 or 4, comprising (Item 6) The system according to any one of items 3 to 5, wherein the directory repeat array comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. (Item 7) The system according to item 6, wherein the directory repeat array comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. (Item 8) The directory repeat array comprises one or more of the following sequences: (a) X 1 X 2 TX 3 X 4 X 5 X 6 X 7 X 8 (SEQ ID NO: 224) (where X 1 is A or C or G, X 2 is T or C or A, X 3 is T or G or A, X 4is T or G, X 5 is T or G or A, X 6 is G or T or A, X 7 is T or G or A, X 8 is A or G or T); (b) X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 (SEQ ID NO: 226) (where X 1 is T or C or A, X 2 is T or A or G, X 3 is T or C or A, X 4 is T or A, X 5 is T or A or G, X 6 is T or A, X 7 is A or T, X 8 is A or G or C or T, X 9 is G or A or C); and (c) X 1 X 2 X 3 AC (SEQ ID NO: 228) (where X 1 is A or C or G, X 2 is C or A, X 3 is A or C) The system according to any one of items 3 to 7, comprising (Item 9) The system according to any one of items 3 to 8, wherein the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. (Item 10) The system according to item 9, wherein the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. (Item 11) The system according to any one of items 3 to 8, wherein the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, the CRISPR-related protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G. (Item 12) The CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G, the system according to item 11. (Item 13) The CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence contains a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60, the system according to any one of items 3 to 8. (Item 14) The CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence contains a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60, the system according to item 13. (Item 15) The CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence includes a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G, the system according to any one of items 3 to 8. (Item 16) The CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein can recognize a PAM sequence, and the PAM sequence includes a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G, the system according to item 15. (Item 17) The CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence includes a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213, the system according to any one of items 3 to 8. (Item 18) The CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence contains a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213, the system according to item 17. (Item 19) The CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G, the system according to any one of items 3 to 8. (Item 20) The CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G, the system according to item 19. (Item 21) The spacer sequence of the RNA guide contains from about 15 nucleotides to about 55 nucleotides, the system according to any one of items 1 to 20. (Item 22) The spacer sequence of the RNA guide contains 20 to 45 nucleotides, the system according to item 21. (Item 23) The CRISPR-related protein contains a catalytic residue (e.g., aspartic acid or glutamic acid), the system according to any one of items 1 to 22. (Item 24) The system according to any one of items 1 to 23, wherein the CRISPR-related protein cleaves the target nucleic acid. (Item 25) The system according to any one of items 1 to 24, wherein the CRISPR-related protein further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcription modification factor, a light gate control factor, a chemical inducibility factor, or a chromatin visualization factor. (Item 26) The system according to any one of items 1 to 25, wherein the nucleic acid encoding the CRISPR-related protein is codon-optimized for expression in cells. (Item 27) The system according to any one of items 1 to 26, wherein the nucleic acid encoding the CRISPR-related protein is operably linked to a promoter. (Item 28) The system according to any one of items 1 to 27, wherein the nucleic acid encoding the CRISPR-related protein is within a vector. (Item 29) The system according to item 28, wherein the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector. (Item 30) The system according to any one of items 1 to 29, wherein the target nucleic acid is a DNA molecule. (Item 31) The system according to any one of items 1 to 30, wherein the CRISPR-related protein comprises non-specific nuclease activity. (Item 32) The system according to any one of items 1 to 31, wherein the modification of the target nucleic acid occurs by recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide. (Item 33) The system according to item 32, wherein the modification of the target nucleic acid is a double-strand break event. (Item 34) The system according to item 32, wherein the modification of the target nucleic acid is a single-strand break event. (Item 35) The system according to item 32, wherein an insertion event occurs by the modification of the target nucleic acid. (Item 36) The system according to item 32, wherein a deletion event occurs by the modification of the target nucleic acid. (Item 37) The system according to any one of items 32 to 36, wherein cytotoxicity or cell death occurs by the modification of the target nucleic acid. (Item 38) The system according to any one of items 1 to 30, further comprising a donor template nucleic acid. (Item 39) The system according to item 38, wherein the donor template nucleic acid is a DNA molecule. (Item 40) The system according to item 38, wherein the donor template nucleic acid is an RNA molecule. (Item 41) The system according to any one of items 1 to 40, wherein the RNA guide optionally includes a tracrRNA. (Item 42) The system according to any one of items 1 to 40, wherein the system does not include a tracrRNA. (Item 43) The system according to any one of items 1 to 42, wherein the CRISPR-related protein is self-processing. (Item 44) The system according to any one of items 1 to 43, wherein the system is present in a delivery composition comprising nanoparticles, liposomes, exosomes, microvesicles, or a gene gun. (Item 45) The system according to any one of items 1 to 43, which is intracellular. (Item 46) The system according to item 45, wherein the cell is a eukaryotic cell. (Item 47) The system according to item 45, wherein the cell is a prokaryotic cell. (Item 48) (a) A CRISPR-related protein or a nucleic acid encoding the CRISPR-related protein, wherein the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 56; and (b) An RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid A cell comprising the above. (Item 49) The CRISPR-related protein is one or more of the following sequences: (a) PX 1 X 2 X 3 X 4 F (SEQ ID NO: 216) (where X 1 is L or M or I or C or F, X 2 is Y or W or F, X 3 is K or T or C or R or W or Y or H or V, X 4 is I or L or M); (b) RX 1 X 2 X 3 L (SEQ ID NO: 217) (where X 1 is I or L or M or Y or T or F, X 2 is R or Q or K or E or S or T, X 3 is L or I or T or C or M or K); (c) NX 1 YX 2 (SEQ ID NO: 218) (where X 1 is I or L or F, X 2 is K or R or V or E); (d) KX 1 X 2 X 3 FAX 4 X 5 KD (SEQ ID NO: 219) (where X 1 is T or I or N or A or S or F or V, X 2 is I or V or L or S, X 3 is H or S or G or R, and X 4 is D or S or E, and X 5 is I or V or M or T or N); (e) LX 1 NX 2 (SEQ ID NO: 220) (where X 1 is G or S or C or T, and X 2 is N or Y or K or S); (f) PX 1 X 2 X 3 X 4 SQX 5 DS (SEQ ID NO: 221) (where X 1 is S or P or A, and X 2 is Y or S or A or P or E or Y or Q or N, and X 3 is F or Y or H, and X 4 is T or S, and X 5 is M or T or I); (g) KX 1 X 2 VRX 3 X 4 QEX 5 H (SEQ ID NO: 222) (where X 1 is N or K or W or R or E or T or Y, and X 2 is M or R or L or S or K or V or E or T or I or D, and X 3 is L or R or H or P or T or K or P's Q or S or A, and X 4 is G or Q or N or R or K or E or I or T or S or C, and X 5 is R or W or Y or K or T or F or S or Q); and (h) X 1 NGX 2 X 3 X 4 DX 5 NX 6 X 7 X 8 N (SEQ ID NO: 223) (where X 1 is I or K or V or L, and X 2 is L or M, and X 3 is N or H or P, and X 4 is A or S or C, and X 5 is V or Y or I or F or T or N, and X 6 is A or S, and X 7 is S or A or P, and X 8 is M or C or L or R or N or S or K or L) The cell according to item 48, comprising (Item 50) The cell according to item 48 or 49, wherein the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. (Item 51) The cell according to item 50, wherein the direct repeat sequence comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. (Item 52) The direct repeat sequence is one or more of the following sequences: (a) X 1 X 2 TX 3 X 4 X 5 X 6 X 7 X 8 (SEQ ID NO: 224) (where X 1 is A or C or G, and X 2 is T or C or A, and X 3 is T or G or A, and X 4 is T or G, and X 5 is T or G or A, X 6 is G or T or A, X 7 is T or G or A, X 8 is A or G or T); (b) X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 (SEQ ID NO: 226) (where X 1 is T or C or A, X 2 is T or A or G, X 3 is T or C or A, X 4 is T or A, X 5 is T or A or G, X 6 is T or A, X 7 is A or T, X 8 is A or G or C or T, X 9 is G or A or C); and (c) X 1 X 2 X 3 AC (SEQ ID NO: 228) (where X 1 is A or C or G, X 2 is C or A, X 3 is A or C) The cell according to any one of items 48 to 51, comprising (Item 53) The cell according to any one of items 48 to 52, wherein the CRISPR-related protein has at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. (Item 54) The cell according to item 53, wherein the CRISPR-related protein has at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. (Item 55) The cell according to any one of items 48 to 52, wherein the CRISPR-related protein is a protein having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, the CRISPR-related protein has the ability to recognize a PAM sequence, the PAM sequence contains a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', "N" is any nucleotide, and "R" is A or G. (Item 56) The cell according to item 55, wherein the CRISPR-related protein is a protein having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, the CRISPR-related protein has the ability to recognize a PAM sequence, the PAM sequence contains a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', "N" is any nucleotide, and "R" is A or G. (Item 57) The cell according to any one of items 48 to 52, wherein the CRISPR-related protein is a protein having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence contains a nucleotide sequence having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. (Item 58) The cell according to item 57, wherein the CRISPR-related protein is a protein having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence contains a nucleotide sequence having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. (Item 59) The CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein has the ability to recognize a PAM sequence, the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G, the cell according to any one of items 48 to 52. (Item 60) The CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein has the ability to recognize a PAM sequence, the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G, the cell according to item 59. (Item 61) The CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the direct repeat sequence contains a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213, the cell according to any one of items 48 to 52. (Item 62) The cell according to item 61, wherein the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 63) The cell according to any one of items 48 to 52, wherein the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G. (Item 64) The cell according to item 63, wherein the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G. (Item 65) The cell according to any one of items 48 to 64, wherein the spacer sequence comprises from about 15 nucleotides to about 55 nucleotides. (Item 66) The cell according to item 65, wherein the spacer sequence comprises 20 to 45 nucleotides. (Item 67) The cell according to any one of items 48 to 66, wherein the cell further comprises tracrRNA. (Item 68) The cell according to any one of items 48 to 66, wherein the system does not contain tracrRNA. (Item 69) The cell according to any one of items 48 to 68, wherein the cell is a eukaryotic cell, such as a mammalian cell, such as a human cell. (Item 70) The cell according to any one of items 48 to 69, wherein the cell is a prokaryotic cell. (Item 71) A method of binding a system according to any one of items 1 to 47 to a target nucleic acid in a cell, comprising: (a) providing the system; and (b) delivering the system to the cell wherein the cell contains the target nucleic acid, the CRISPR-related protein binds to the RNA guide, and the spacer sequence binds to the target nucleic acid. (Item 72) The method according to item 71, wherein the cell is a eukaryotic cell, such as a mammalian cell, such as a human cell. (Item 73) A method of modifying a target nucleic acid, comprising: (a) a CRISPR-related protein having an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 56, or a nucleic acid encoding the CRISPR-related protein; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to the target nucleic acid; comprising: the CRISPR-related protein having the ability to bind to the RNA guide; modification of the target nucleic acid occurs by recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide; a method comprising delivering an engineered, non-naturally occurring CRISPR-Cas system to a target nucleic acid. (Item 74) The CRISPR-related protein is one or more of the following sequences: (a) PX 1 X 2 X 3 X 4 F (SEQ ID NO: 216) (where X 1 is L or M or I or C or F, X 2 is Y or W or F, X 3 is K or T or C or R or W or Y or H or V, X 4 is I or L or M); (b) RX 1 X 2 X 3 L (SEQ ID NO: 217) (where X 1 is I or L or M or Y or T or F, X 2 is R or Q or K or E or S or T, X 3 is L or I or T or C or M or K); (c) NX 1 YX 2 (SEQ ID NO: 218) (where X 1 is I or L or F, X 2 is K or R or V or E); (d) KX 1 X 2 X 3 FAX 4 X 5 KD (SEQ ID NO: 219) (where X 1 is T or I or N or A or S or F or V, and X 2 is I or V or L or S, and X 3 is H or S or G or R, and X 4 is D or S or E, and X 5 is I or V or M or T or N); (e) LX 1 NX 2 (SEQ ID NO: 220) (where X 1 is G or S or C or T, and X 2 is N or Y or K or S); (f) PX 1 X 2 X 3 X 4 SQX 5 DS (SEQ ID NO: 221) (where X 1 is S or P or A, and X 2 is Y or S or A or P or E or Y or Q or N, and X 3 is F or Y or H, and X 4 is T or S, and X 5 is M or T or I); (g) KX 1 X 2 VRX 3 X 4 QEX 5 H (SEQ ID NO: 222) (where X 1 is N or K or W or R or E or T or Y, and X 2 is M or R or L or S or K or V or E or T or I or D, and X 3 is L or R or H or P or T or K or P's Q or S or A, and X 4 is G or Q or N or R or K or E or I or T or S or C, and X 5 is R or W or Y or K or T or F or S or Q); and (h) X 1 NGX 2X 3 X 4 DX 5 NX 6 X 7 X 8 N (SEQ ID NO: 223) (where X 1 is I or K or V or L, and X 2 is L or M, and X 3 is N or H or P, and X 4 is A or S or C, and X 5 is V or Y or I or F or T or N, and X 6 is A or S, and X 7 is S or A or P, and X 8 is M or C or L or R or N or S or K or L) The method according to item 73, comprising (Item 75) The method according to item 73 or 74, wherein the directory repeat array comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. (Item 76) The method according to item 75, wherein the directory repeat array comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213. (Item 77) The directory repeat array is one or more of the following sequences: (a) X 1 X 2 TX 3 X 4 X 5 X 6 X 7 X 8 (SEQ ID NO: 224) (where X 1 is A or C or G, and X 2 is T or C or A, and X 3 is T or G or A, and X 4 is T or G, and X 5 is T or G or A, and X 6 is G or T or A, and X 7 is T or G or A, and X 8 is A or G or T); (b) X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 (SEQ ID NO: 226) (where X 1 is T or C or A, and X 2 is T or A or G, and X 3 is T or C or A, and X 4 is T or A, and X 5 is T or A or G, and X 6 is T or A, and X 7 is A or T, and X 8 is A or G or C or T, and X 9 is G or A or C); and (c) X 1 X 2 X 3 AC (SEQ ID NO: 228) (where X 1 is A or C or G, and X 2is C or A, and X 3 is A or C) The method according to any one of items 73 to 76, comprising (Item 78) The method according to any one of items 73 to 77, wherein the CRISPR-related protein is a protein having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. (Item 79) The method according to item 78, wherein the CRISPR-related protein is a protein having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 57. (Item 80) The method according to any one of items 73 to 77, wherein the CRISPR-related protein is a protein having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence includes a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G. (Item 81) The method according to item 80, wherein the CRISPR-related protein is a protein having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 1, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence includes a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G. (Item 82) The method according to any one of items 73 to 77, wherein the CRISPR-related protein is a protein having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence includes a nucleotide sequence having at least 80% (for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. (Item 83) The method according to item 82, wherein the CRISPR-related protein is a protein having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence includes a nucleotide sequence having at least 95% (for example, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 60. (Item 84) The method according to any one of items 73 to 77, wherein the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G. (Item 85) The method according to item 84, wherein the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G. (Item 86) The method according to any one of items 73 to 77, wherein the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence contains a nucleotide sequence having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 87) The method according to item 86, wherein the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence contains a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 88) The method according to any one of items 73 to 77, wherein the CRISPR-related protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G. (Item 89) The method according to item 88, wherein the CRISPR-related protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity with the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G. (Item 90) The method according to any one of items 73 to 89, wherein the spacer sequence contains from about 15 nucleotides to about 55 nucleotides. (Item 91) The method according to item 90, wherein the spacer sequence contains 20 to 45 nucleotides. (Item 92) The method according to any one of items 73 to 91, wherein the system further comprises a tracrRNA. (Item 93) The method according to any one of items 73 to 91, wherein the system does not contain a tracrRNA. (Item 94) The method according to any one of items 73 to 93, wherein the target nucleic acid is a DNA molecule. (Item 95) The method according to any one of items 73 to 94, wherein the CRISPR-related protein comprises non-specific nuclease activity. (Item 96) The method according to any one of items 73 to 95, wherein the modification of the target nucleic acid is a double-strand break event. (Item 97) The method according to any one of items 73 to 96, wherein the modification of the target nucleic acid is a single-strand break event. (Item 98) The method according to any one of items 73 to 97, wherein an insertion event occurs due to the modification of the target nucleic acid. (Item 99) The method according to any one of items 73 to 98, wherein a deletion event occurs due to the modification of the target nucleic acid. (Item 100) The method according to any one of items 73 to 99, wherein cytotoxicity or cell death occurs due to the modification of the target nucleic acid. (Item 101) A method for editing a target nucleic acid, comprising contacting the target nucleic acid with the system according to any one of items 1 to 47. (Item 102) A method for modifying the expression of a target nucleic acid, comprising contacting the target nucleic acid with the system according to any one of items 1 to 47. (Item 103) A method for targeting the insertion of a payload nucleic acid at a site of a target nucleic acid, comprising contacting the target nucleic acid with the system according to any one of items 1 to 47. (Item 104) A method for targeting the excision of a payload nucleic acid from a site of a target nucleic acid, comprising contacting the target nucleic acid with the system according to any one of items 1 to 47. (Item 105) A method for non-specifically degrading single-stranded DNA upon recognition of a DNA target nucleic acid, comprising contacting the target nucleic acid with the system according to any one of items 1 to 47. (Item 106) A method for detecting a target nucleic acid in a sample, (a) contacting the sample with the system according to any one of items 1 to 47 and a labeled reporter nucleic acid, wherein when the spacer sequence hybridizes to the target nucleic acid, cleavage of the labeled reporter nucleic acid occurs; and (b) measuring a detectable signal generated by the cleavage of the labeled reporter nucleic acid, thereby detecting the presence of the target nucleic acid in the sample. (Item 107) (a) A method for targeting and editing a target nucleic acid; (b) A method for non-specifically degrading single-stranded nucleic acids in response to the recognition of nucleic acids; (c) Method for targeting and nicking the non-spacer complementary strand of a double-stranded target according to the recognition of the spacer complementary strand of the double-stranded target; (d) Method for targeting and cleaving a double-stranded target nucleic acid; (e) Method for detecting a target nucleic acid in a sample; (f) Method for specifically editing a double-stranded nucleic acid; (g) Method for base editing of a double-stranded nucleic acid; (h) Method for inducing genotype-specific or transcription state-specific cell death or dormancy in cells; (i) Method for creating indels in a double-stranded nucleic acid target; (j) Method for inserting a sequence into a double-stranded nucleic acid target; or (k) Method for deleting or inverting a sequence in a double-stranded nucleic acid target Use of the system according to any one of items 1 to 47 in an in vitro or ex vivo method. (Item 108) A method for introducing an insertion or deletion into a target nucleic acid in mammalian cells, comprising: (a) a nucleic acid sequence encoding a CRISPR-related protein, wherein the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 56; and (b) an RNA guide (or a nucleic acid encoding the RNA guide) comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to the target nucleic acid comprising transfection, wherein the CRISPR-related protein has the ability to bind to the RNA guide; a method in which the modification of the target nucleic acid occurs by the recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide. (Item 109) The method according to item 108, wherein the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4. (Item 110) The method according to item 109, wherein the CRISPR-related protein comprises an amino acid sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4. (Item 111) The method according to item 108, wherein the directory repeat comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 112) The method according to item 111, wherein the directory repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 113) The method according to item 108, wherein the target nucleic acid is adjacent to a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G. (Item 114) The method according to item 108, wherein the CRISPR-related protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. (Item 115) The method according to item 114, wherein the CRISPR-related protein comprises an amino acid sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. (Item 116) The method according to item 108, wherein the directory repeat comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 117) The method according to item 116, wherein the directory repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 118) The method according to item 108, wherein the target nucleic acid is adjacent to a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5’-NTTN-3’ or 5’-RTTR-3’ (for example, 5’-ATTG-3’ or 5’-GTTA-3’), where “N” is any nucleotide and “R” is A or G. (Item 119) The method according to any one of items 108 to 118, wherein the transfection is transient transfection. (Item 120) The method according to any one of items 108 to 119, wherein the cell is a human cell. (Item 121) (a) A CRISPR-related protein or a nucleic acid encoding the CRISPR-related protein, and (b) An RNA guide comprising a direct repeat sequence and a spacer sequence A composition comprising: The CRISPR-related protein is one or more of the following amino acid sequences: (i) PX 1 X 2 X 3 X 4 F (SEQ ID NO: 216) (where X 1 is L or M or I or C or F, X 2 is Y or W or F, X 3is K or T or C or R or W or Y or H or V, X 4 is I or L or M); (ii) RX 1 X 2 X 3 L (SEQ ID NO: 217) (where X 1 is I or L or M or Y or T or F, X 2 is R or Q or K or E or S or T, X 3 is L or I or T or C or M or K); (iii) NX 1 YX 2 (SEQ ID NO: 218) (where X 1 is I or L or F, X 2 is K or R or V or E); (iv) KX 1 X 2 X 3 FAX 4 X 5 KD (SEQ ID NO: 219) (where X 1 is T or I or N or A or S or F or V, X 2 is I or V or L or S, X 3 is H or S or G or R, X 4 is D or S or E, X 5 is I or V or M or T or N); (v) LX 1 NX 2 (SEQ ID NO: 220) (where X 1 is G or S or C or T, X 2 is N or Y or K or S); (vi) PX 1 X 2 X 3 X 4 SQX 5 DS (SEQ ID NO: 221) (where X 1 is S or P or A, X 2 is Y or S or A or P or E or Y or Q or N, X 3 is F or Y or H, X 4 is T or S, X 5 is M or T or I); (vii) KX 1 X 2 VRX 3 X 4 QEX 5 H (SEQ ID NO: 222) (where X 1 is N or K or W or R or E or T or Y, X 2 is M or R or L or S or K or V or E or T or I or D, X 3 is L or R or H or P or T or K or P of Q or S or A, X 4 is G or Q or N or R or K or E or I or T or S or C, and X 5 is R or W or Y or K or T or F or S or Q); and (viii) X 1 NGX 2 X 3 X 4 DX 5 NX 6 X 7 X 8 N (SEQ ID NO: 223) (where X 1 is I or K or V or L, X 2 is L or M, X 3 is N or H or P, X 4 is A or S or C, X 5is V or Y or I or F or T or N, X 6 is A or S, X 7 is S or A or P, X 8 is M or C or L or R or N or S or K or L) comprising a composition, wherein the CRISPR-related protein binds to the RNA guide and the spacer binds to a target nucleic acid.

Claims

**Claim 1**: An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-Cas system, comprising: (a) a CRISPR-associated protein or a nucleic acid encoding said CRISPR-associated protein, wherein the CRISPR-associated protein comprises the amino acid sequence of SEQ ID NO: 241; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid ; wherein said CRISPR-associated protein is capable of binding to said RNA guide and modifying said target nucleic acid sequence complementary to said spacer sequence; a CRISPR-Cas system, wherein said CRISPR-associated protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence set forth in SEQ ID NO: 10, SEQ ID NO: 4, SEQ ID NO: 12, or SEQ ID NO:

14. **Claim 2** The system according to claim 1, wherein said CRISPR-associated protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence set forth in SEQ ID NO:

10. **Claim 3**: The system according to claim 1, wherein said CRISPR-associated protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence set forth in SEQ ID NO:

4. **Claim 4** The system according to any one of claims 1 to 3, wherein said CRISPR-associated protein comprises at least one RuvC domain or at least one split RuvC domain. **Claim 5** The CRISPR-associated protein (a) contains catalytic residues; (b) cleaves said target nucleic acid; (c) further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcriptional modification factor, a light gate control factor, a chemically inducible factor, or a chromatin visualization factor; or (d) is self-processing. The system according to any one of claims 1 to 4. **Claim 6** The nucleic acid encoding said CRISPR-associated protein (a) is codon-optimized for expression in cells; (b) is operably linked to a promoter; or (c) is within a vector. The system according to any one of claims 1 to 5. **Claim 7** The system according to any one of claims 1 to 6, wherein said target nucleic acid is a DNA molecule. **Claim 8** The system according to any one of claims 1 to 7, wherein the target nucleic acid is modified by recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide.

9. The system according to any one of claims 1 to 8, further comprising a donor template nucleic acid.

10. The system according to claim 9, wherein the donor template nucleic acid is a DNA molecule.

11. The system is a) present in a delivery composition comprising a nanoparticle, liposome, exosome, microvesicle, or gene gun; or b) intracellular, The system according to any one of claims 1 to 10.

12. a) a CRISPR-related protein or a nucleic acid encoding the CRISPR-related protein, wherein the CRISPR-related protein comprises the amino acid sequence of SEQ ID NO: 241, and the CRISPR-related protein comprises at least 95% identical to the amino acid sequence set forth in SEQ ID NO: 10, SEQ ID NO: 4, SEQ ID NO: 12, or SEQ ID NO: 14 A CRISPR-related protein or a nucleic acid encoding the CRISPR-related protein; and b) an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid A cell comprising.

13. A method of binding the system according to any one of claims 1 to 11 to a target nucleic acid in a cell ex vivo, a) providing the system; and b) delivering the system to the cell Including, wherein the cell contains the target nucleic acid, the CRISPR-related protein binds to the RNA guide, and the spacer sequence binds to the target nucleic acid.

14. The cell according to claim 12, wherein the cell is a prokaryotic cell, eukaryotic cell, mammalian cell, or human cell.

15. A method of modifying a target nucleic acid, a) a CRISPR-related protein or a nucleic acid encoding the CRISPR-related protein, wherein the CRISPR-related protein comprises the amino acid sequence of SEQ ID NO: 241, and the CRISPR-related protein comprises at least 95% identical to the amino acid sequence set forth in SEQ ID NO: 10, SEQ ID NO: 4, SEQ ID NO: 12, or SEQ ID NO: 14 A CRISPR-related protein or a nucleic acid encoding the CRISPR-related protein; and An RNA guide comprising a (b) a directory repeat sequence and a spacer sequence having the ability to hybridize to the target nucleic acid; comprising, wherein the CRISPR-related protein has the ability to bind to the RNA guide, recognition of the target nucleic acid by the CRISPR-related protein and the RNA guide results in modification of the target nucleic acid, A method comprising delivering an engineered naturally-occurring CRISPR-Cas system to a target nucleic acid.

16. The CRISPR-related protein comprises one or more of the following sequences: (a) PX 1 X 2 X 3 X 4 F (SEQ ID NO: 216) (where X 1 is L or M or I or C or F, X 2 is Y or W or F, X 3 is K or T or C or R or W or Y or H or V, X 4 is I or L or M); (b)RX 1 X 2 X 3 L (SEQ ID NO: 217) (where X 1 is I or L or M or Y or T or F, and X 2 is R or Q or K or E or S or T, and X 3 is L or I or T or C or M or K); (c) NX 1 YX 2 (SEQ ID NO: 218) (where X 1 is I or L or F, and X 2 is K or R or V or E); (d) KX 1 X 2 X 3 FAX 4 X 5 KD (SEQ ID NO: 219) (where X 1 is T or I or N or A or S or F or V, and X 2 is I or V or L or S, and X 3 is H or S or G or R, and X 4 is D or S or E, and X 5 is I or V or M or T or N); (e) LX 1 NX 2 (SEQ ID NO: 220) (where X 1 is G or S or C or T, and X 2 is N or Y or K or S); (f) PX 1 X 2 X 3 X 4 SQX 5 DS (SEQ ID NO: 221) (where X 1 is S or P or A, X 2 is Y or S or A or P or E or Y or Q or N, X 3 is F or Y or H, X 4 is T or S, X 5 is M or T or I); (g) KX 1 X 2 VRX 3 X 4 QEX 5 H (SEQ ID NO: 222) (where X 1 is N or K or W or R or E or T or Y, and X 2 is M or R or L or S or K or V or E or T or I or D, and X 3 is L or R or H or P or T or K or P or Q or S or A, and X 4 is G or Q or N or R or K or E or I or T or S or C, and X 5 is R or W or Y or K or T or F or S or Q); and (h)X 1 NGX 2 X 3 X 4 DX 5 NX 6 X 7 X 8 N (sequence number 223) (where X 1 is I or K or V or L, X 2 is L or M, X 3 is N or H or P, X 4 is A or S or C, X 5 is V or Y or I or F or T or N, X 6 is A or S, X 7 is S or A or P, X 8 is M or C or L or R or N or S or K or L) The system according to any one of claims 1 to 11, comprising.

17. The directory repeat sequence is i) comprises a nucleotide sequence that is at least 80% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, SEQ ID NOs: 118-151, or SEQ ID NO: 213; or ii) one or more of the following sequences: (a) X 1 X 2 TX 3 X 4 X 5 X 6 X 7 X 8 (Array No. 224) (where X 1 is A or C or G, X 2 is T or C or A, X 3 is T or G or A, X 4 is T or G, X 5 is T or G or A, X 6 is G or T or A, X 7 is T or G or A, X 8 is A or G or T); (b) X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 (SEQ ID NO: 226) (wherein X 1 is T or C or A, X 2 is T or A or G, X 3 is T or C or A, X 4 is T or A, X 5 is T or A or G, X 6 is T or A, X 7 is A or T, X 8 is A or G or C or T, X 9 is G or A or C); and (c) X 1 X 2 X 3 AC (SEQ ID NO: 228) (where X 1 is A or C or G, X 2 is C or A, X 3 is A or C) comprising, The system according to any one of claims 1 to 11 or 16.

18. The CRISPR-related protein is a) a protein having at least 95% identity to the amino acid sequence set forth in SEQ ID NO: 4, and the directory repeat sequence comprises a nucleotide sequence that is at least 80% identical to the nucleotide sequence set forth in SEQ ID NO: 60; b) a protein having at least 95% identity to the amino acid sequence set forth in SEQ ID NO: 4, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3', or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G; c) a protein having at least 95% identity to the amino acid sequence set forth in SEQ ID NO: 10, and the directory repeat sequence comprises a nucleotide sequence that is at least 95% identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213; or d) a protein having at least 95% identity to the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-related protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3', where "N" is any nucleotide and "R" is A or G, The system according to any one of claims 1 to 11, 16 or 17.

19. The system according to any one of claims 1 to 11 or 16 to 18, wherein the directory repeat array contains uracil at one or more positions shown as thymine.

20. The system according to any one of claims 1 to 11 or 16 to 19, wherein the spacer array contains 15 to 55 nucleotides.

21. The system according to any one of claims 1 to 11 or 16 to 20, wherein the system does not contain tracrRNA.

22. The system according to any one of claims 1 to 11 or 16 to 21, wherein the target nucleic acid is a DNA molecule.

23. (a) The modification of the target nucleic acid is a double-strand cleavage event; (b) The modification of the target nucleic acid is a single-strand cleavage event; (c) An insertion event occurs due to the modification of the target nucleic acid; or (d) A deletion event occurs due to the modification of the target nucleic acid, The system according to any one of claims 1 to 11 or 16 to 22.

24. (a) A method for targeting and editing a target nucleic acid; (b) A method for non-specific degradation of single-stranded nucleic acids according to the recognition of nucleic acids; (c) A method for targeting and nicking the non-spacer complementary strand of a double-stranded target according to the recognition of the spacer complementary strand of the double-stranded target; (d) A method for targeting and cleaving a double-stranded target nucleic acid; (e) A method for detecting a target nucleic acid in a sample; (f) A method for specifically editing a double-stranded nucleic acid; (g) A method for base editing of a double-stranded nucleic acid; (h) A method for inducing genotype-specific or transcription state-specific cell death or dormancy in cells; (i) A method for creating indels in a double-stranded nucleic acid target; (j) A method for inserting a sequence into a double-stranded nucleic acid target; or (k) A method for deleting or inverting a sequence in a double-stranded nucleic acid target is the use of the system according to any one of claims 1 to 11 or 16 to 23 in an in vitro or ex vivo method.

25. A composition comprising the engineered non-naturally occurring CRISPR-Cas system according to any one of claims 1 to 11 or 16 to 23 for use in a method for introducing an insertion or deletion into a target nucleic acid in mammalian cells.

26. The composition according to claim 25, wherein the CRISPR-related protein contains an amino acid sequence that is at least 95% identical to the amino acid sequence set forth in SEQ ID NO: 4 or 10.

27. The composition according to claim 25 or 26, wherein the directory repeat contains a nucleotide sequence that is at least 80% identical to the nucleotide sequence set forth in SEQ ID NO:

60. **Claim 28** The composition according to any one of claims 25 to 27, wherein the target nucleic acid is adjacent to a PAM sequence, and the PAM sequence contains a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3', or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G. **Claim 29** (i) the transfection is transient transfection; or (ii) the cell is a human cell, The composition according to any one of claims 25 to 28.