Novel crispr DNA targeting enzymes and systems
Engineered CRISPR-Cas systems with novel CRISPR-associated proteins and RNA guides address the limitations of existing systems by enabling targeted nucleic acid modification in diverse cellular environments, enhancing applicability and specificity.
Patent Information
- Application Number
- JP2025109210
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-09
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-24
AI Technical Summary
Current CRISPR-Cas systems lack programmable effectors with unique PAM sequence requirements for modifying nucleic acids, limiting their applications in non-native environments such as bacteria and eukaryotic cells.
Development of engineered, non-naturally occurring CRISPR-Cas systems, designated CLUST.091979, comprising CRISPR-associated proteins with specific amino acid sequences and RNA guides capable of hybridizing to target nucleic acids, allowing for efficient modification and recognition of protospacer adjacent motifs (PAMs) in diverse cellular environments.
Enables targeted nucleic acid modification in various cellular contexts, including eukaryotic cells, with enhanced specificity and versatility beyond traditional CRISPR-Cas systems, facilitating novel biotechnological applications.
Smart Images

Figure 2025161809000045 
Figure 2025161809000046 
Figure 2025161809000047
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 897,859, filed September 9, 2019, the entire contents of which are hereby incorporated by reference.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy, created on September 9, 2020, has the file name A2186-7028WO_SL.txt and is 475,511 bytes in size.
[0003] The present disclosure relates to systems and methods for genome editing and regulation of gene expression using novel Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and CRISPR-associated (Cas) genes. [Background technology]
[0004] Recent advances in genome sequencing technology and analysis have provided important insights into the genetic basis of biological activity in a wide variety of natural domains, ranging from prokaryotic biosynthetic pathways to human pathologies. To fully understand and evaluate the vast amount of information generated, corresponding improvements in the scale, efficacy, and ease of sequencing technologies for genome and epigenome manipulation are required. These novel technologies will accelerate the development of novel applications in numerous areas, including biotechnology, agriculture, and human therapeutics.
[0005] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes, collectively known as the CRISPR-Cas or CRISPR / Cas system, are adaptive immune systems in archaea and bacteria that defend certain species against foreign genetic elements. CRISPR-Cas systems contain a highly diverse set of protein effectors, non-coding elements, and locus architectures, several examples of which have been engineered and adapted to produce important biotechnological advances.
[0006] Components of this system involved in host defense include one or more effector proteins capable of modifying nucleic acids and an RNA guide element responsible for targeting the effector proteins to specific sequences on the phage nucleic acid. CRISPR systems are composed of a single RNA (crRNA) and may require an additional transactivating RNA (tracrRNA) to manipulate the target nucleic acid with one or more effector proteins. The crRNA consists of direct repeats responsible for protein binding to the crRNA and a spacer sequence complementary to the desired nucleic acid target sequence. The CRISPR system can be reprogrammed to target alternative DNA or RNA targets by modifying the spacer sequence of the crRNA.
[0007] CRISPR-Cas systems can be broadly divided into two classes: Class 1 systems consist of multiple effector proteins that together form a complex around the crRNA, whereas Class 2 systems consist of a single effector protein that complexes with an RNA guide to target a nucleic acid substrate. The single-subunit effector composition of Class 2 systems provides a more convenient set of components for engineering and application translation and has thus far represented an important source of programmable effectors. Nevertheless, there remains a need for additional programmable effectors and systems beyond current CRISPR-Cas systems, such as smaller effectors and / or effectors with unique PAM sequence requirements that enable novel applications through their unique properties for modifying nucleic acids and polynucleotides (i.e., DNA, RNA, or any hybrid, derivative, or modified form). Summary of the Invention [Means for solving the problem]
[0008] This disclosure provides non-naturally occurring engineered systems and compositions for novel single-effector Class 2 CRISPR-Cas systems that were initially computationally identified from genome databases, then engineered, and experimentally validated. In particular, the identification of these CRISPR-Cas system components enables their use in non-native environments, such as bacteria other than those in which the systems were originally discovered, or eukaryotic cells, such as mammalian cells. These novel effectors differ in sequence and function compared to orthologs and homologs of existing Class 2 CRISPR effectors.
[0009] In one aspect, the present disclosure provides an engineered, non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-Cas system, designated CLUST.091979, comprising: a CRISPR-associated protein comprising an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid, wherein the CRISPR-associated protein is capable of binding to the RNA guide and modifying a target nucleic acid sequence complementary to the spacer sequence. In one aspect, the disclosure provides an engineered, non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-Cas system of CLUST.091979, comprising: a CRISPR-associated protein or a nucleic acid encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to an amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide, or a nucleic acid encoding the RNA guide, comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid, wherein the CRISPR-associated protein is capable of binding to the RNA guide and modifying a target nucleic acid sequence complementary to the spacer sequence.
[0010] In some aspects, the disclosure provides an engineered, non-naturally occurring, clustered regularly interspaced short palindromic repeats (CRISPR)-Cas system of CLUST.091979, comprising a CRISPR-associated protein or a nucleic acid encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises the amino acid sequence of SEQ ID NO: 241; and an RNA guide having a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid, wherein the CRISPR-associated protein is capable of binding to the RNA guide and modifying a target nucleic acid sequence that is complementary to the spacer sequence. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO:4, SEQ ID NO:10, SEQ ID NO:12, or SEQ ID NO:14.
[0011] In some embodiments of any of the systems described herein, the CRISPR-associated protein comprises at least one (e.g., one, two, or three) RuvC domain or at least one split RuvC domain.
[0012] In some embodiments of any of the systems described herein, the CRISPR-associated protein has one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216), where XI is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M; (b) RX1X2X3L (SEQ ID NO: 217), where XI is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or is C or M or K; (c) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E and X5 is I or V or M or T or N); (e) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or is S; (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, and X3 is L or R or H or P or T or K or P, Q or S or A) wherein X4 is G or Q or N or R or K or E or I or T or S or C and X5 is R or W or Y or K or T or F or S or Q; and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P and X8 is M or C or L or R or N or S or K or L).In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 216 is the N-terminal sequence. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 219 is the C-terminal sequence. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 220 is the C-terminal sequence. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 221 is the C-terminal sequence. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 222 is the C-terminal sequence. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 223 is the C-terminal sequence.
[0013] In some embodiments of any of the systems described herein, the CRISPR-associated protein is one or more of the following sequences: (a) ECPITKDVINEYK (SEQ ID NO: 290); (b) NLTSITIG (SEQ ID NO: 231); (c) NYRTKIRTLN (SEQ ID NO: 232); (d) ISYIENVEN (SEQ ID NO: 233); (e) ELLSVEQLK (SEQ ID NO: 234); (f) HINSMTINIQDFKIE (SEQ ID NO: 235); (g) KENSLGFIL (SEQ ID NO: 236); (h) GNRQIKKG (SEQ ID NO: 237); (i) DVNFKHA (SEQ ID NO: 238); (j) GYINLYKYLLEH (SEQ ID NO: 239); (k) KEQVLSKLLY (SEQ ID NO: 240); (l) EYIYVSCVNKLRAKYVSYFILKE (u) LLSNNGKTQIALVPSE (SEQ ID NO:250); (v) HINGLNADFNAANNIKYI (SEQ ID NO:251), or a sequence having no more than one, two, or three sequence differences (e.g., substitutions) relative to any of the foregoing. In some embodiments, the CRISPR-associated protein has a sequence at least 70% identical to SEQ ID NO:4. In some embodiments, the CRISPR-associated protein has a sequence that is at least 70% identical to SEQ ID NO:10.
[0014] In some embodiments of any of the systems described herein, the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, 118-151, or 213. In some embodiments of any of the systems described herein, the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, 118-151, or 213.
[0015] In some embodiments of any of the systems described herein, the direct repeat sequence is one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (wherein Xi is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 227)); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein Xi is A or C or G, X2 is C or A, and X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any of the systems described herein, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments of any of the systems described herein, SEQ ID NO: 228 is proximal to the 3' end of the direct repeat.
[0016] In some embodiments of any of the systems described herein, the CRISPR-associated protein has the ability to recognize a protospacer adjacent motif (PAM), wherein the PAM sequence comprises a nucleic acid sequence, which is 5'-NTTN-3', 5'-NTTR-3', 5'-RTTR-3', 5'-TNNT-3', 5'-TNRT-3', 5'-TSRT-3', 5'-TGRT-3', 5'-TNRY-3', 5'-TTNR-3', 5'-TTYR-3', 5'-TTTR-3', 5'-TTCV-3', 5'-DTYR-3' , 5'-WTTR-3', 5'-NNR-3', 5'-NYR-3', 5'-YYR-3', 5'-TYR-3', 5'-TTN-3', 5'-TTR-3', 5'-CNT-3', 5'-NGG-3', 5'-BGG-3', or 5'-R-3', wherein "N" is any nucleotide, "B" is C or G or T, "D" is A or G or T, "R" is A or G, "S" is G or C, "V" is A or C or G, "W" is A or T, and "Y" is C or T.
[0017] In some embodiments of any of the systems described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the systems described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the systems described herein, the CRISPR-associated protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence set forth as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G.
[0018] In some embodiments of any of the systems described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:60. In some embodiments of any of the systems described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the systems described herein, the CRISPR-associated protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G.
[0019] In some embodiments of any of the systems described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the systems described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the systems described herein, the CRISPR-associated protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G.
[0020] In some embodiments of any of the systems described herein, the spacer sequence of the RNA guide comprises between about 15 nucleotides and about 55 nucleotides. In some embodiments of any of the systems described herein, the spacer sequence of the RNA guide comprises between 20 and 45 nucleotides.
[0021] In some embodiments of any of the systems described herein, the CRISPR-associated protein comprises a catalytic residue (e.g., aspartic acid or glutamic acid). In some embodiments of any of the systems described herein, the CRISPR-associated protein cleaves the target nucleic acid. In some embodiments of any of the systems described herein, the CRISPR-associated protein further comprises a peptide tag, a fluorescent protein, a base-editing domain, a DNA methylation domain, a histone residue-modifying domain, a localization factor, a transcription modifier, a photogating factor, a chemical-inducible factor, or a chromatin visualization factor.
[0022] In some embodiments of any of the systems described herein, the nucleic acid encoding the CRISPR-associated protein is codon-optimized for expression in a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell). In some embodiments of any of the systems described herein, the nucleic acid encoding the CRISPR-associated protein is operably linked to a promoter. In some embodiments of any of the systems described herein, the nucleic acid encoding the CRISPR-associated protein is in a vector. In some embodiments, the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.
[0023] In some embodiments of any of the systems described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the systems described herein, the target nucleic acid comprises a PAM sequence.
[0024] In some embodiments of any of the systems described herein, the CRISPR-associated protein has non-specific nuclease activity.
[0025] In some embodiments of any of the systems described herein, recognition of the target nucleic acid by the CRISPR-associated protein and the RNA guide results in modification of the target nucleic acid. In some embodiments of any of the systems described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the systems described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the systems described herein, the modification of the target nucleic acid results in an insertion event. In some embodiments of any of the systems described herein, the modification of the target nucleic acid results in a deletion event. In some embodiments of any of the systems described herein, the modification of the target nucleic acid results in cytotoxicity or cell death.
[0026] In some embodiments of any of the systems described herein, the system further comprises a donor template nucleic acid. In some embodiments of any of the systems described herein, the donor template nucleic acid is a DNA molecule. In some embodiments of any of the systems described herein, the donor template nucleic acid is an RNA molecule.
[0027] In some embodiments of any of the systems described herein, the RNA guide optionally comprises a tracrRNA and / or a modulator RNA. In some embodiments of any of the systems described herein, the system further comprises a tracrRNA. In some embodiments of any of the systems described herein, the system does not comprise a tracrRNA. In some embodiments of any of the systems described herein, the CRISPR-associated protein is self-processing. In some embodiments of any of the systems described herein, the system further comprises a modulator RNA.
[0028] In some embodiments of any of the systems described herein, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the amino acid sequence of SEQ ID NO:1, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the nucleotide sequence of SEQ ID NO:152, SEQ ID NO:153, or SEQ ID NO:154.
[0029] In some embodiments of any of the systems described herein, the system is present in a delivery composition comprising a nanoparticle, a liposome, an exosome, a microvesicle, or a gene gun.
[0030] In some embodiments of any of the systems described herein, the system is within a cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a prokaryotic cell.
[0031] In another aspect, the present disclosure provides a cell, wherein the cell comprises a CRISPR-associated protein comprising an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid. In another aspect, the disclosure provides a cell, wherein the cell comprises a CRISPR-associated protein or a nucleic acid encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide, or a nucleic acid encoding the RNA guide, comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid.
[0032] In some embodiments of any of the cells described herein, the CRISPR-associated protein comprises at least one (e.g., one, two, or three) RuvC domain or at least one split RuvC domain.
[0033] In some embodiments of any of the cells described herein, the CRISPR-associated protein has one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216), where XI is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M; (b) RX1X2X3L (SEQ ID NO: 217), where XI is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K; (c) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (e) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, and X3 is L or R or H or P or T or K or P, Q or S or A) wherein X4 is G or Q or N or R or K or E or I or T or S or C and X5 is R or W or Y or K or T or F or S or Q; and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P and X8 is M or C or L or R or N or S or K or L).In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 216 is the N-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 219 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 220 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 221 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 222 is the C-terminal sequence. In some embodiments of any of the cells described herein, the sequence of SEQ ID NO: 223 is the C-terminal sequence.
[0034] In some embodiments of any of the cells described herein, the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, 118-151, or 213. In some embodiments of any of the cells described herein, the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, 118-151, or 213.
[0035] In some embodiments of any of the cells described herein, the direct repeat sequence is one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 ( X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 226)); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any of the cells described herein, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments of any of the cells described herein, SEQ ID NO: 228 is proximal to the 3' end of the direct repeat.
[0036] In some embodiments of any of the cells described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the cells described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the cells described herein, the CRISPR-associated protein is capable of recognizing a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G.
[0037] In some embodiments of any of the cells described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:60. In some embodiments of any of the cells described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the cells described herein, the CRISPR-associated protein is capable of recognizing a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G.
[0038] In some embodiments of any of the cells described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the cells described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the cells described herein, the CRISPR-associated protein is capable of recognizing a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G.
[0039] In some embodiments of any of the cells described herein, the spacer sequence comprises from about 15 nucleotides to about 55 nucleotides. In some embodiments of any of the cells described herein, the spacer sequence comprises from 20 to 45 nucleotides.
[0040] In some embodiments of any of the cells described herein, the CRISPR-associated protein comprises a catalytic residue (e.g., aspartic acid or glutamic acid). In some embodiments of any of the cells described herein, the CRISPR-associated protein cleaves a target nucleic acid. In some embodiments of any of the cells described herein, the CRISPR-associated protein further comprises a peptide tag, a fluorescent protein, a base-editing domain, a DNA methylation domain, a histone residue-modifying domain, a localization factor, a transcription modifier, a light-gating factor, a chemical-inducible factor, or a chromatin visualization factor.
[0041] In some embodiments of any of the cells described herein, the nucleic acid encoding the CRISPR-associated protein is codon-optimized for expression in a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell). In some embodiments of any of the cells described herein, the nucleic acid encoding the CRISPR-associated protein is operably linked to a promoter. In some embodiments of any of the cells described herein, the nucleic acid encoding the CRISPR-associated protein is in a vector. In some embodiments, the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.
[0042] In some embodiments of any of the cells described herein, the RNA guide optionally comprises tracrRNA and / or modulator RNA. In some embodiments of any of the cells described herein, the cell further comprises tracrRNA. In some embodiments of any of the cells described herein, the cell does not comprise tracrRNA. In some embodiments of any of the cells described herein, the CRISPR-associated protein is self-processing. In some embodiments of any of the cells described herein, the cell further comprises modulator RNA.
[0043] In some embodiments of any of the cells described herein, the cell is a eukaryotic cell. In some embodiments of any of the cells described herein, the cell is a mammalian cell. In some embodiments of any of the cells described herein, the cell is a human cell. In some embodiments of any of the cells described herein, the cell is a prokaryotic cell.
[0044] In some embodiments of any of the cells described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the cells described herein, the target nucleic acid comprises a PAM sequence.
[0045] In some embodiments of any of the cells described herein, the CRISPR-associated protein has non-specific nuclease activity.
[0046] In some embodiments of any of the cells described herein, recognition of the target nucleic acid by the CRISPR-associated protein and the RNA guide results in modification of the target nucleic acid. In some embodiments of any of the cells described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the cells described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the cells described herein, the modification of the target nucleic acid results in an insertion event. In some embodiments of any of the cells described herein, the modification of the target nucleic acid results in a deletion event. In some embodiments of any of the cells described herein, the modification of the target nucleic acid results in cytotoxicity or cell death.
[0047] In another aspect, the present disclosure provides a method of binding a system described herein to a target nucleic acid in a cell, comprising: (a) providing the system; and (b) delivering the system to a cell, wherein the cell contains the target nucleic acid, the CRISPR-associated protein binds to the RNA guide, and the spacer sequence binds to the target nucleic acid. In some embodiments, the cell is a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell.
[0048] In another aspect, the present disclosure provides a method of modifying a target nucleic acid, the method comprising delivering to the target nucleic acid an engineered, non-naturally occurring CRISPR-Cas system comprising: a CRISPR-associated protein comprising an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid, wherein the CRISPR-associated protein is capable of binding to the RNA guide, and wherein recognition of the target nucleic acid by the CRISPR-associated protein and the RNA guide results in modification of the target nucleic acid. In another aspect, the disclosure provides a method of modifying a target nucleic acid, the method comprising delivering to a target nucleic acid an engineered non-naturally occurring CRISPR-Cas system comprising: a CRISPR-associated protein or a nucleic acid encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid, wherein the CRISPR-associated protein is capable of binding to the RNA guide, and wherein recognition of the target nucleic acid by the CRISPR-associated protein and the RNA guide results in modification of the target nucleic acid.
[0049] In some embodiments of any of the methods described herein, the CRISPR-associated protein has one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216), where XI is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M; (b) RX1X2X3L (SEQ ID NO: 217), where XI is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K; (c) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (e) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, and X3 is L or R or H or P or T or K or P, Q or S or A) wherein X4 is G or Q or N or R or K or E or I or T or S or C and X5 is R or W or Y or K or T or F or S or Q; and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P and X8 is M or C or L or R or N or S or K or L).In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 216 is the N-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 219 is the C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 220 is the C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 221 is the C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 222 is the C-terminal sequence. In some embodiments of any of the methods described herein, the sequence of SEQ ID NO: 223 is the C-terminal sequence.
[0050] In some embodiments of any of the methods described herein, the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, 118-151, or 213. In some embodiments of any of the methods described herein, the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57-90, 118-151, or 213.
[0051] In some embodiments of any of the methods described herein, the direct repeat sequence is one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 ( X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 226)); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any of the methods described herein, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments of any of the methods described herein, SEQ ID NO: 228 is proximal to the 3' end of the direct repeat.
[0052] In some embodiments of any of the methods described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the methods described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. In some embodiments of any of the methods described herein, the CRISPR-associated protein is capable of recognizing a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-TNNT-3' or 5'-TNRT-3', where "N" is any nucleotide and "R" is A or G.
[0053] In some embodiments of any of the methods described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:60. In some embodiments of any of the methods described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods described herein, the CRISPR-associated protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, where the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G.
[0054] In some embodiments of any of the methods described herein, the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods described herein, the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods described herein, the CRISPR-associated protein is capable of recognizing a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G.
[0055] In some embodiments of any of the methods described herein, the spacer sequence comprises about 15 nucleotides to about 55 nucleotides. In some embodiments of any of the methods described herein, the spacer sequence comprises 20 to 45 nucleotides.
[0056] In some embodiments of any of the methods described herein, the RNA guide optionally comprises a tracrRNA and / or a modulator RNA. In some embodiments of any of the methods described herein, the system further comprises a tracrRNA. In some embodiments of any of the methods described herein, the system does not comprise a tracrRNA. In some embodiments of any of the methods described herein, the CRISPR-associated protein is self-processing. In some embodiments of any of the methods described herein, the system further comprises a modulator RNA.
[0057] In some embodiments of any of the methods described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the methods described herein, the target nucleic acid comprises a PAM sequence.
[0058] In some embodiments of any of the methods described herein, the CRISPR-associated protein has non-specific nuclease activity.
[0059] In some embodiments of any of the methods described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid results in an insertion event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid results in a deletion event. In some embodiments of any of the methods described herein, the modification of the target nucleic acid results in cytotoxicity or cell death.
[0060] In another aspect, the present disclosure provides a method for editing a target nucleic acid, the method comprising contacting a system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for modifying expression of a target nucleic acid, the method comprising contacting a system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for targeting insertion of a payload nucleic acid at a site in a target nucleic acid, the method comprising contacting a system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for targeting excision of a payload nucleic acid from a site in a target nucleic acid, the method comprising contacting a system described herein with the target nucleic acid. In another aspect, the present disclosure provides a method for nonspecifically degrading single-stranded DNA upon recognition of a DNA target nucleic acid, the method comprising contacting a system described herein with the target nucleic acid.
[0061] In some embodiments of any of the systems or methods provided herein, contacting includes direct contacting or indirect contacting. In some embodiments of any of the systems or methods provided herein, indirect contacting includes administering one or more nucleic acids encoding an RNA guide or CRISPR-associated protein described herein under conditions that allow production of the RNA guide and / or CRISPR-associated protein. In some embodiments of any of the systems or methods provided herein, contacting includes in vivo contacting or in vitro contacting. In some embodiments of any of the systems or methods provided herein, contacting a target nucleic acid with the system includes contacting a cell containing the nucleic acid with the system under conditions that allow the CRISPR-associated protein and guide RNA to reach the target nucleic acid. In some embodiments of any of the systems or methods provided herein, contacting a cell with the system in vivo includes administering the system to a subject containing the cell under conditions that allow the CRISPR-associated protein and guide RNA to reach or be produced in the cell.
[0062] In another aspect, the disclosure provides a system provided herein for use in an in vitro or ex vivo method that is: (a) a method for targeting and editing a target nucleic acid; (b) a method for non-specific degradation of a single-stranded nucleic acid in response to recognition of a nucleic acid; (c) a method for targeting and nicking a non-spacer complement of a double-stranded target in response to recognition of a spacer complement of a double-stranded target; (d) a method for targeting and cleaving a double-stranded target nucleic acid; (e) a method for detecting a target nucleic acid in a sample; (f) a method for specifically editing a double-stranded nucleic acid; (g) a method for base editing a double-stranded nucleic acid; (h) a method for inducing genotype-specific or transcriptional state-specific cell death or dormancy in a cell; (i) a method for creating indels in a double-stranded nucleic acid target; (j) a method for inserting a sequence into a double-stranded nucleic acid target; or (k) a method for deleting or forming an inversion in a double-stranded nucleic acid target.
[0063] In another aspect, the disclosure provides a method of introducing an insertion or deletion into a target nucleic acid in a mammalian cell, comprising transfection of (a) a nucleic acid sequence encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and (b) an RNA guide (or a nucleic acid encoding the RNA guide) comprising a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid, wherein the CRISPR-associated protein is capable of binding to the RNA guide, and wherein recognition of the target nucleic acid by the CRISPR-associated protein and the RNA guide results in modification of the target nucleic acid.
[0064] In some embodiments of any of the methods provided herein, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments of any of the methods provided herein, the CRISPR-associated protein comprises an amino acid sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO:4. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. In some embodiments of any of the methods provided herein, the target nucleic acid is flanked by a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G.
[0065] In some embodiments of any of the methods provided herein, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. In some embodiments of any of the methods provided herein, the CRISPR-associated protein comprises an amino acid sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods provided herein, the direct repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments of any of the methods provided herein, the target nucleic acid is flanked by a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G.
[0066] In some embodiments of any of the methods provided herein, the transfection is transient transfection. In some embodiments of any of the methods provided herein, the cell is a human cell.
[0067] In another aspect, the disclosure provides a composition comprising (a) a CRISPR-associated protein or a nucleic acid encoding a CRISPR-associated protein and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence; wherein the CRISPR-associated protein has one or more of the following amino acid sequences: (i) PX1X2X3X4F (SEQ ID NO: 216), where XI is L or M or I or C or F, X2 is Y or W or F, and X3 is K or T or C or R or W or Y or H or V; and X4 is I or L or M; (ii) RX1X2X3L (SEQ ID NO: 217) (wherein XI is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (iii) NX1YX2 (SEQ ID NO: 218) (wherein XI is I or L or F and X2 is K or R or V or E); (iv) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein XI is T or I or N or A or S or F or is V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N; (v) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or S); (vi) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, and X3 is F or Y or H wherein X4 is T or S and X5 is M, T or I; (vii) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein X1 is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P, Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q);and (viii) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223), wherein X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L, and wherein the CRISPR-associated protein can bind to the RNA guide and modify a target nucleic acid sequence complementary to the spacer sequence;
[0068] In some embodiments of any of the compositions described herein, the direct repeat sequence is one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T) (e.g., ATTGTTGDA (SEQ ID NO: 225)); (b) X1X2X3X4X5X6X7X8X9 ( X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C) (e.g., TTTTWTARG (SEQ ID NO: 226)); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) (e.g., ACAAC (SEQ ID NO: 229)). In some embodiments of any of the compositions described herein, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments of any of the compositions described herein, SEQ ID NO: 228 is proximal to the 3' end of the direct repeat.
[0069] In some embodiments of any of the compositions described herein, the CRISPR-associated protein comprises at least one (e.g., one, two, or three) RuvC domain or at least one split RuvC domain.
[0070] In some embodiments of any of the compositions described herein, the spacer sequence of the RNA guide comprises from about 15 nucleotides to about 55 nucleotides. In some embodiments of any of the compositions described herein, the spacer sequence of the RNA guide comprises from 20 to 45 nucleotides.
[0071] In some embodiments of any of the compositions described herein, the CRISPR-associated protein comprises a catalytic residue (e.g., aspartic acid or glutamic acid). In some embodiments of any of the compositions described herein, the CRISPR-associated protein cleaves a target nucleic acid. In some embodiments of any of the compositions described herein, the CRISPR-associated protein further comprises a peptide tag, a fluorescent protein, a base-editing domain, a DNA methylation domain, a histone residue-modifying domain, a localization factor, a transcription modifier, a photogating factor, a chemical-inducible factor, or a chromatin visualization factor.
[0072] In some embodiments of any of the compositions described herein, the nucleic acid encoding the CRISPR-associated protein is codon-optimized for expression in a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell). In some embodiments of any of the compositions described herein, the nucleic acid encoding the CRISPR-associated protein is operably linked to a promoter. In some embodiments of any of the compositions described herein, the nucleic acid encoding the CRISPR-associated protein is in a vector. In some embodiments, the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.
[0073] In some embodiments of any of the compositions described herein, the target nucleic acid is a DNA molecule. In some embodiments of any of the compositions described herein, the target nucleic acid comprises a PAM sequence.
[0074] In some embodiments of any of the compositions described herein, the CRISPR-associated protein has non-specific nuclease activity.
[0075] In some embodiments of any of the compositions described herein, recognition of the target nucleic acid by the CRISPR-associated protein and the RNA guide results in modification of the target nucleic acid. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid is a double-strand break event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid is a single-strand break event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid results in an insertion event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid results in a deletion event. In some embodiments of any of the compositions described herein, the modification of the target nucleic acid results in cytotoxicity or cell death.
[0076] In some embodiments of any of the compositions described herein, the system further comprises a donor template nucleic acid. In some embodiments of any of the compositions described herein, the donor template nucleic acid is a DNA molecule. In some embodiments of any of the compositions described herein, the donor template nucleic acid is an RNA molecule.
[0077] In some embodiments of any of the compositions described herein, the RNA guide optionally comprises a tracrRNA. In some embodiments of any of the compositions described herein, the system further comprises a tracrRNA. In some embodiments of any of the compositions described herein, the system does not comprise a tracrRNA. In some embodiments of any of the compositions described herein, the CRISPR-associated protein is self-processing.
[0078] In some embodiments of any of the compositions described herein, the system is present in a delivery composition comprising a nanoparticle, a liposome, an exosome, a microvesicle, or a gene gun.
[0079] In some embodiments of any of the compositions described herein, the composition is intracellular. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a prokaryotic cell.
[0080] The effectors described herein offer additional features, including, but not limited to, 1) novel nucleic acid editing properties and control mechanisms, 2) smaller size relative to greater versatility in delivery strategies, 3) genotype-driven cellular processes such as cell death, 4) programmable RNA-guided DNA insertion, excision, and mobilization, and 5) differentiated profiles of pre-existing immunity via non-human commensal sources. See, e.g., Examples 1, 4, and 5 and Figures 1-3 and 5-11D. The novel DNA targeting systems described herein add to the toolbox of genome and epigenome engineering techniques, enabling broad applications for specific, programmed perturbations.
[0081] Other features and advantages of the invention will be apparent from the following detailed description, and from the claims.
[0082] The figures are a series of schematic representations showing the results of an analysis of a protein cluster called CLUST.091979. [Brief explanation of the drawings]
[0083] [Figure 1A] Figures 1A, 1B, 1C, 1D, 1E, 1F, 1G, 1H, 1I, 1J, 1K, and 1L show the alignment of effectors of SEQ ID NOs: 1 to 4, 14, 15, 17 to 19, 21 to 25, 27 to 33, 35 to 49, and 51 to 56. [Figure 1B] Same as above. [Figure 1C] Same as above. [Figure 1D] Same as above. [Figure 1E] Same as above. [Figure 1F] Same as above. [Figure 1G] Same as above. [Figure 1H] Same as above. [Figure 1I] Same as above. [Figure 1J] Same as above. [Figure 1K] Same as above. [Figure 1L] Same as above. [Figure 2] Schematic diagram showing the RuvC domain of the CLUST.091979 effector, which is based on the consensus sequence of the sequences shown in Table 6. [Figure 3] Shown is an alignment of the direct repeat sequences of SEQ ID NOs: 57, 58, 60, 62, 63, 70, 72-74, 76, 77, 80, 83, 84, 86-88, 90, 128, 130, 139, and 213. The consensus sequence (SEQ ID NO: 230) is shown above the alignment. [Figure 4A] 1 is a schematic representation of the components of the in vivo negative selection screening assay described in Example 4. A CRISPR array library was designed containing non-representative spacers flanked by two DRs and uniformly sampled from both strands of pACYC184 or E. coli essential genes expressed by J23119. [Figure 4B] Schematic representation of the in vivo negative selection screening workflow described in Example 4. The CRISPR array library was cloned into an effector plasmid. The effector and non-coding plasmids were transformed into E. coli and subsequently grown for negative selection of CRISPR arrays that confer interference with transcripts from essential E. coli genes or pACYC184. Targeted sequencing of the effector plasmid was used to identify depleted CRISPR arrays. Small RNA sequencing was further performed to identify the requirements for mature crRNAs and potential tracrRNAs. [Figure 5] 1 is a graph of CLUST.091979 AUXO013988882 (effector set forth in SEQ ID NO: 1) showing the degree of depletion activity of engineered compositions for pACYC184 with non-coding sequences and spacers targeting the direct repeat transcriptional orientation. The degree of depletion of the direct repeat in the "forward" orientation (5'-ACTA...AACT-[spacer]-3') and the "reverse" orientation (5'-AGTT...TAGT-[spacer]-3') is shown. [Figure 6A] Figure 6A is a graphical representation showing the density of depleted and non-depleted targets of CLUST.091979 AUXO013988882, which have non-coding sequences, by location on the pACYC184 plasmid. Figure 6B is a graphical representation showing the density of depleted and non-depleted targets of CLUST.091979 AUXO013988882, which have non-coding sequences, by location on the Escherichia coli (E. coli) strain, E. Cloni. Targets on the top and bottom strands are shown separately with respect to the annotated gene orientation. Band size indicates the degree of depletion, with bright bands closer to the hit threshold of 3. The gradient is an RNA-sequencing heat map showing relative transcript abundance. [Figure 6B] Same as above. [Figure 7]WebLogo of sequences flanking the depleted target in E. Cloni as predicted PAM sequences for CLUST.091979 AUXO013988882 (including non-coding sequences). [Figure 8] 1 is a graph of CLUST.091979 SRR3181151 (effector set forth in SEQ ID NO: 4) showing the degree of depletion activity of engineered compositions for pACYC184 with non-coding sequences and spacers targeting the direct repeat transcriptional orientation. The degree of depletion of the "forward" orientation direct repeat (5'-GTTG...CAGG-[spacer]-3') and the "reverse" orientation direct repeat (5'-CCTG...CAAC-[spacer]-3') is shown as solid and dashed lines, respectively. [Figure 9A] Figure 9A is a graphical representation showing the density of depleted and non-depleted targets of CLUST.091979 SRR3181151, which have non-coding sequences, by location on the pACYC184 plasmid. Figure 9B is a graphical representation showing the density of depleted and non-depleted targets of CLUST.091979 SRR3181151, which have non-coding sequences, by location on the Escherichia coli (E. coli) strain, E. Cloni. Targets on the top and bottom strands are shown separately with respect to the annotated gene orientation. Band size indicates the degree of depletion, with bright bands closer to the hit threshold of 3. The gradient is an RNA-sequencing heat map showing relative transcript abundance. [Figure 9B] Same as above. [Figure 10] WebLogo of sequences flanking the depleted target in E. cloni as predicted PAM sequences for CLUST.091979 SRR3181151 (with non-coding sequences). [Figure 11A]Figure 11A shows indels induced by the effector of SEQ ID NO: 4 at the AAVS1 target locus of SEQ ID NO: 206 and the VEGFA target locus of SEQ ID NO: 208 in HEK293 cells. Figure 11B shows indels induced by the effector of SEQ ID NO: 4 at the AAVS1 target loci of SEQ ID NOs: 253, 255, 257, 259, and 275, the VEGFA target loci of SEQ ID NOs: 263, 265, 267, 269, 271, 273, and 277, and the EMX1 target locus of SEQ ID NO: 261 in HEK293 cells. Figure 11C shows indels induced by the effector of SEQ ID NO: 10 at the AAVS1 target locus of SEQ ID NO: 210, the AAVS1 target locus of SEQ ID NO: 212, and the VEGFA target locus of SEQ ID NO: 215 in HEK293 cells. FIG. 11D shows indels induced by the effector of SEQ ID NO: 10 at the AAVS1 target loci of SEQ ID NO: 279, 281, 285, and 287, the VEGFA target locus of SEQ ID NO: 283, and the EMX1 target locus of SEQ ID NO: 289 in HEK293 cells. [Figure 11B] Same as above. [Figure 11C] Same as above. [Figure 11D] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0084] CRISPR-Cas systems are naturally diverse and contain a variety of active mechanisms and functional elements that can be exploited in programmable biotechnology. In nature, these systems provide efficient defense against foreign DNA and viruses while providing self-nonself discrimination to avoid self-targeting. In engineered settings, these systems offer a diverse toolbox of molecular techniques and define the boundaries of targeting space. Using the methods described herein, additional mechanisms and parameters have been discovered within single-subunit class 2 effector systems that expand the capabilities of RNA-programmable nucleic acid manipulation.
[0085] Unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. Applicant reserves the right to alternatively claim any disclosed inventions using the transitional phrases "comprising," "consisting essentially of," or "consisting of," in accordance with standard practice in patent law.
[0086] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, reference to a "nucleic acid" means one or more nucleic acids.
[0087] It should be noted that terms such as "preferably," "suitably," "generally," and "typically" are not used herein to limit the scope of the claimed invention or to suggest that a particular feature is critical, essential, or even essential to the structure or function of the claimed invention. Rather, these terms are merely intended to highlight alternative or additional features that may or may not be utilized in a particular embodiment of the invention.
[0088] It should be noted that for purposes of describing and defining the present invention, the term "substantially" is used herein to represent the inherent degree of uncertainty that can be attributed to any quantitative comparison, value, measurement, or other representation. The term "substantially" is also used herein to represent the extent to which a quantitative representation can vary from the stated reference without resulting in a change in the basic functionality of the subject matter at issue.
[0089] The term "CRISPR-Cas system," as used herein, refers to nucleic acids and / or proteins involved in the expression of or directing the activity of CRISPR effectors, including sequences encoding CRISPR effectors, RNA guides, and other sequences and transcripts from the CRISPR locus.
[0090] The terms "CRISPR-associated protein," "CRISPR-Cas effector," "CRISPR effector," "effector," "effector protein," "CRISPR enzyme," and the like, when used interchangeably herein, refer to a protein that carries out an enzymatic activity or binds to a target site on a nucleic acid specified by an RNA guide. In some embodiments, a CRISPR effector has endonuclease activity, nickase activity, and / or exonuclease activity.
[0091] The terms "RNA guide," "guide RNA," "gRNA," and "guide sequence," as used herein, refer to any RNA molecule that facilitates targeting of an effector described herein to a target nucleic acid, such as DNA and / or RNA. Exemplary "RNA guides" include, but are not limited to, crRNA and crRNA hybridized or fused with either tracrRNA and / or modulator RNA. In some embodiments, an RNA guide comprises both crRNA and tracrRNA, either fused into a single RNA molecule or as separate RNA molecules. In some embodiments, an RNA guide comprises crRNA and modulator RNA, either fused into a single RNA molecule or as separate RNA molecules. In some embodiments, an RNA guide comprises crRNA, tracrRNA, and modulator RNA, either fused into a single RNA molecule or as separate RNA molecules.
[0092] The terms "CRISPR effector complex," "effector complex," or "surveillance complex," as used herein, refer to a complex comprising a CRISPR effector and an RNA guide. A CRISPR effector complex may further comprise one or more accessory proteins. The one or more accessory proteins may be non-catalytic and / or non-target binding. The crRNA may comprise a sequence that hybridizes to the tracrRNA. The crRNA:tracrRNA duplex may then bind to the CRISPR effector. As used herein, the term "pre-crRNA" refers to an unprocessed RNA molecule comprising a DR-spacer-DR sequence. As used herein, the term "mature crRNA" refers to the processed form of the pre-crRNA. The mature crRNA may comprise a DR-spacer sequence, where the DR is a truncated version of the DR of the pre-crRNA and / or the spacer is a truncated version of the spacer of the pre-crRNA.
[0093] The terms "CRISPR RNA" and "crRNA" as used herein refer to an RNA molecule containing a guide sequence that is used by a CRISPR effector to specifically recognize a nucleic acid sequence. The crRNA "spacer" sequence is complementary to the nucleic acid target sequence and has the ability to bind partially or completely to the nucleic acid target sequence.
[0094] The term "trans-activating crRNA" or "tracrRNA," as used herein, refers to an RNA molecule that contains sequences that form the structural and / or sequence motifs required for a CRISPR effector to bind to a specific target nucleic acid.
[0095] As used herein, the term "CRISPR array" refers to a nucleic acid (e.g., DNA) segment that includes CRISPR repeats and spacers, starting from the first nucleotide of the first CRISPR repeat and ending with the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeats. As used herein, the terms "CRISPR repeat", "CRISPR direct repeat" and "direct repeat" refer to multiple short, directional repeat sequences, which show little or no sequence variation within a CRISPR array.
[0096] The term "modulator RNA," as used herein, refers to any RNA molecule that modulates (e.g., increases or decreases) the activity of a CRISPR effector or a nucleoprotein complex that includes a CRISPR effector. In some embodiments, the modulator RNA modulates the nuclease activity of a CRISPR effector or a nucleoprotein complex that includes a CRISPR effector.
[0097] As used herein, the term "target nucleic acid" refers to a nucleic acid comprising a nucleotide sequence complementary to all or part of a spacer in an RNA guide. In some embodiments, the target nucleic acid comprises a gene. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter). In some embodiments, the target nucleic acid is single-stranded. In some embodiments, the target nucleic acid is double-stranded. "Transcriptionally active site," as used herein, refers to a site in a nucleic acid sequence that is actively being transcribed.
[0098] As used herein, the term "protospacer adjacent motif" or "PAM" refers to a DNA sequence adjacent to a target sequence to which a complex comprising an effector and an RNA guide binds. In some embodiments, a PAM is required for enzymatic activity. As used herein, the term "adjacent" includes cases where the RNA guide of the complex specifically binds, interacts with, or associates with a target sequence directly adjacent to the PAM. In such cases, there are no nucleotides between the target sequence and the PAM. The term "adjacent" also includes cases where there are a small number of nucleotides (e.g., one, two, three, four, or five) between the target sequence to which the targeting moiety binds and the PAM. As used herein, the term "recognizing a PAM sequence" refers to binding of a complex comprising a CRISPR-associated protein and crRNA to a target nucleic acid, where the target nucleic acid is adjacent to the PAM sequence.
[0099] The terms "activated CRISPR effector complex," "activated CRISPR complex," and "activated complex," as used herein, refer to a CRISPR effector complex that can modify a target nucleic acid. In some embodiments, an activated CRISPR complex can modify a target nucleic acid after the activated CRISPR complex binds to the target nucleic acid. In some embodiments, binding of the activated CRISPR complex to the target nucleic acid results in an additional cleavage event, such as collateral cleavage.
[0100] The term "cleavage event" as used herein refers to the cleavage of nucleic acid such as DNA and / or RNA.In some embodiments, the cleavage event as used herein refers to the cleavage in the target nucleic acid created by the nuclease of the CRISPR system described herein.In some embodiments, the cleavage event is a double-stranded DNA cleavage.In some embodiments, the cleavage event is a single-stranded DNA cleavage.In some embodiments, the cleavage event refers to the cleavage of collateral nucleic acid.
[0101] The term " collateral nucleic acid " as used herein refers to the nucleic acid substrate that is non-specifically cut by activated CRISPR complex.The term " collateral DNase activity " as used herein in reference to CRISPR effector refers to the non-specific DNase activity of activated CRISPR complex.The term " collateral RNase activity " as used herein in reference to CRISPR effector refers to the non-specific RNase activity of activated CRISPR complex.
[0102] The term "donor template nucleic acid," as used herein, refers to a nucleic acid molecule that can be used to make templated changes to a target sequence or a target-proximal sequence after a CRISPR effector described herein modifies the target nucleic acid. In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome).
[0103] As used herein, the terms "polynucleotide," "nucleotide," "oligonucleotide," and "nucleic acid" may be used interchangeably to refer to nucleic acids, including DNA, RNA, derivatives thereof, or combinations thereof. Methods well known to those skilled in the art can be used to construct gene expression constructs and recombinant cells according to the present invention. These methods include in vitro recombinant DNA techniques, synthetic techniques, in vivo recombinant techniques, and polymerase chain reaction (PCR) techniques. See, e.g., Maniatis et al., 1989, MOLECULAR CLONING: A LABORATORY MANUAL, Cold Spring Harbor Laboratory, New York; Ausubel et al., 1989, CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, Greene Publishing. Associates and Wiley Interscience, New York, and the techniques described in PCR Protocols: A Guide to Methods and Applications (Innis et al., 1990, Academic Press, San Diego, Calif.).
[0104] The terms "genetic modification" or "genetic engineering" broadly refer to the manipulation of a cell's genome or nucleic acid. Similarly, the terms "genetically engineered" and "engineered" refer to a cell containing an engineered genome or nucleic acid. Methods of genetic modification include, for example, heterologous gene expression, gene or promoter insertion or deletion, nucleic acid mutation, altered gene expression or inactivation, enzyme engineering, directed evolution, knowledge-based design, random mutagenesis, gene shuffling, and codon optimization.
[0105] The term "recombinant" indicates that a nucleic acid, protein, or cell is the product of genetic modification, engineering, or recombination. Generally, the term "recombinant" refers to a nucleic acid, protein, or cell that contains or is encoded by genetic material derived from multiple sources. As used herein, the term "recombinant" can also be used to describe a cell that contains a mutant nucleic acid or protein, including a mutant version of an endogenous nucleic acid or protein. The terms "recombinant cell" and "recombinant host" can be used interchangeably. In some embodiments, the recombinant cell contains a CRISPR effector disclosed herein. The CRISPR effector can be codon-optimized for expression in the recombinant cell. In some embodiments, the recombinant cell disclosed herein further contains an RNA guide. In some embodiments, the RNA guide of the recombinant cell disclosed herein comprises a tracrRNA. In some embodiments, the recombinant cell disclosed herein contains a modulator RNA. In some embodiments, the recombinant cell is a prokaryotic cell, such as an E. coli cell. In some embodiments, the recombinant cell is a eukaryotic cell, such as a mammalian cell, including a human cell.
[0106] Identification of CLUST.091979 This application relates to the identification, engineering, and use of a novel protein family referred to herein as "CLUST.091979." As shown in Figure 2, CLUST.091979 proteins contain RuvC domains (designated RuvC I, RuvC II, and RuvC III). As shown in Table 5, the size of CLUST.091979 effectors ranges from about 700 amino acids to about 800 amino acids. Thus, as shown below, CLUST.091979 effectors are smaller than effectors known in the art. See, e.g., Table 1.
[0107] [Table 1]
[0108] The effectors of CLUST.091979 were identified using computational methods and algorithms to search for and identify proteins that exhibit strong co-occurrence patterns with other specific functions. In certain embodiments, these computational methods were directed to identifying proteins that co-occur in close proximity to CRISPR arrays. The methods disclosed herein are also useful for identifying proteins that naturally occur in close proximity to other features, both non-coding and protein-coding (e.g., fragments of phage sequences in non-coding regions of bacterial loci; or CRISPR Cas1 proteins). It is understood that the methods and calculations described herein may be performed on one or more computing devices.
[0109] A set of genome sequences was obtained from a genome or metagenomic database. The database included short reads, contig-level data, assembled scaffolds, or complete genome sequences of organisms. Similarly, the database may include genome sequence data from prokaryotes, eukaryotes, or data from metagenomic environmental samples. Examples of database repositories include the National Center for Biotechnology Information (NCBI) RefSeq, NCBI's GenBank, NCBI's Whole Genome Shotgun (WGS), and the Joint Genome Institute's (JGI) Integrated Microbial Genomes (IMG).
[0110] In some embodiments, the selection of genome sequence data of a specified minimum length imposes a minimum size requirement, in certain exemplary embodiments, the minimum contig length may be 100 nucleotides, 500 nt, 1 kb, 1.5 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 40 kb, or 50 kb.
[0111] In some embodiments, known or predicted proteins are extracted from a complete or selected set of genome sequence data. In some embodiments, known or predicted proteins are taken from extracting coding sequence (CDS) annotations provided by a source database. In some embodiments, predicted proteins are determined by applying computational methods to identify proteins from nucleotide sequences. In some embodiments, proteins are predicted from genome sequences using the GeneMark suite. In some embodiments, proteins are predicted from genome sequences using Prodigal. In some embodiments, multiple protein prediction algorithms may be used on the same set of sequence data, and redundancies may be removed from the resulting set of proteins.
[0112] In some embodiments, CRISPR arrays are identified from genome sequence data. In some embodiments, CRISPR arrays are identified using PILER-CR. In some embodiments, CRISPR arrays are identified using CRISPR Recognition Tool (CRT). In some embodiments, CRISPR arrays are identified by a heuristic method that identifies a nucleotide motif that repeats a minimum number of times (for example, 2, 3, or 4 times), and the interval between consecutive occurrences of the repeated motif does not exceed a specified length (for example, 50, 100, or 150 nucleotides). In some embodiments, multiple CRISPR array identification tools can be used on the same set of sequence data, and duplicates can be removed from the resulting set of CRISPR arrays.
[0113] In some embodiments, proteins that are in close proximity to a CRISPR array (referred to herein as "CRISPR-proximal protein clusters") are identified. In some embodiments, proximity is defined as a nucleotide distance, and may be within 20 kb, 15 kb, or 5 kb. In some embodiments, proximity is defined as the number of open reading frames (ORFs) between the protein and the CRISPR array, and certain exemplary distances may be 10, 5, 4, 3, 2, 1, or 0 ORFs. Proteins identified as being within close proximity to the CRISPR array are then grouped into homologous protein clusters. In some embodiments, CRISPR-proximal protein clusters are formed using blastclust. In certain other embodiments, CRISPR-proximal protein clusters are formed using mmseqs2.
[0114] To establish strong co-occurrence patterns between members of a CRISPR-proximal protein cluster, a BLAST search of each member of the protein cluster against a complete, pre-organized set of known and predicted proteins may be performed. In some embodiments, UBLAST or mmseqs2 may be used to search for similar proteins. In some embodiments, a search may be performed on only a representative subset of proteins within the family.
[0115] In some embodiments, CRISPR-proximal protein clusters are ranked or filtered by a metric to determine co-occurrence.One exemplary metric is the ratio of the number of elements in a protein cluster to the number of BLAST matches up to a certain E-value threshold.In some embodiments, a fixed E-value threshold may be used.In other embodiments, the E-value threshold may be determined by the most distant member of the protein cluster.In some embodiments, a global set of proteins is clustered, and the co-occurrence metric is the ratio of the number of elements of CRISPR-proximal proteins to the number of elements of one or more global clusters they contain.
[0116] In some embodiments, a manual review process is used to evaluate the potential functionality and minimal set of components of an engineered system based on the naturally occurring locus structure of proteins in the cluster. In some embodiments, the manual review can be aided by a graphical representation of the protein cluster, which can include information such as pairwise sequence similarity, phylogenetic tree, source organism / environment, predicted functional domains, and a graphical representation of the locus structure. In some embodiments, the graphical representation of the locus structure can filter for highly representative neighboring protein families. In some embodiments, the representativeness can be calculated by the ratio of the number of related neighboring proteins to one or more sizes of the containing global cluster(s). In certain exemplary embodiments, the graphical representation of the protein cluster can include a representation of the CRISPR array structure of the naturally occurring locus. In some embodiments, the graphical representation of the protein cluster can include a representation of the number of conserved direct repeats relative to the length of the estimated CRISPR array, or the number of unique spacer sequences relative to the length of the estimated CRISPR array. In some embodiments, the graphical representation of the protein clusters may include depictions of various co-occurrence metrics of putative effectors with CRISPR arrays to predict novel CRISPR-Cas systems and identify their components.
[0117] Pooled screening of CLUST.091979 To efficiently validate the activity, mechanism, and functional parameters of the engineered CLUST.091979 CRISPR-Cas system identified herein, we used a pooled screening approach in Escherichia coli (E. coli), as described in Example 4. First, from computational identification of conserved proteins and non-coding elements of the CLUST.091979 CRISPR-Cas system, we used DNA synthesis and molecular cloning to assemble the individual components into a single artificial expression vector (based in one embodiment on a pET-28a+ backbone). In a second embodiment, the effectors and non-coding elements are transcribed into mRNA transcripts, and individual effectors are translated using distinct ribosome binding sites.
[0118] Second, the natural crRNA and targeting spacer are replaced with a library of unprocessed crRNAs containing a non-natural spacer that targets a second plasmid, pACYC184. This crRNA library is cloned into a vector backbone (e.g., pET-28a+) containing effector and non-coding elements, and then this library is subsequently transformed into E. coli together with the pACYC184 plasmid target. As a result, each resulting E. coli cell contains a single targeting array. In an alternative embodiment, a library of unprocessed crRNAs containing non-natural spacers further targets essential E. coli genes cited from sources such as those described in Baba et al. (2006) Mol. Syst. Biol. 2:2006.0008; and Gerdes et al. (2003) J. Bacteriol. 185(19):5673-84, the entire contents of each of which are incorporated herein by reference. In this embodiment, positive targeted activation of the novel CRISPR-Cas system to disrupt essential gene function results in cell death or growth arrest. In some embodiments, essential gene targeting spacers can be combined with pACYC184 targeting.
[0119] Third, E. coli is grown under antibiotic selection. In one embodiment, triple antibiotic selection is used: kanamycin to confirm successful transformation of the pET-28a+ vector containing the engineered CRISPR effector system, and chloramphenicol and tetracycline to confirm successful co-transformation of the pACYC184 targeting vector. Because pACYC184 normally confers resistance to chloramphenicol and tetracycline, under antibiotic selection, positive activity of the novel CRISPR-Cas system targeting this plasmid will eliminate cells actively expressing the effector, non-coding elements, and specific active elements of the crRNA library. Typically, the population of surviving cells is analyzed 12-14 hours after transformation. In some embodiments, analysis of surviving cells is performed 6-8 hours after transformation, 8-12 hours after transformation, up to 24 hours after transformation, or more than 24 hours after transformation. Examination of the surviving cell population at later time points compared to earlier time points results in a depletion of signal compared to inactive crRNA.
[0120] In some embodiments, dual antibiotic selection is used. Removing the selective pressure by withdrawing either chloramphenicol or tetracycline can provide new information about targeting substrates, sequence specificity, and efficacy. For example, dsDNA cleavage in selected or non-selected genes can result in negative selection in E. coli, where depletion of both selected and non-selected genes is observed. If the CRISPR-Cas system interferes with transcription or translation (e.g., by binding or cleaving transcripts), selection is observed only against targets of the selected resistance gene, rather than non-selected resistance genes.
[0121] In some embodiments, successful transformation of the pET-28a+ vector containing the engineered CRISPR-Cas system is confirmed using only kanamycin. This embodiment is suitable for libraries containing spacers targeting essential E. coli genes, as no further selection beyond kanamycin is required to observe growth changes. In this embodiment, chloramphenicol and tetracycline dependence is eliminated, and those targets (if present) in the library provide an additional source of negative or positive information regarding targeting substrate, sequence specificity, and potency.
[0122] Because the pACYC184 plasmid contains a diverse set of features and sequences that can affect the activity of CRISPR-Cas systems, mapping active crRNAs from pooled screens to pACYC184 provides activity patterns that may suggest different mechanisms of activity and functional parameters. In this way, the features required for reconstitution of novel CRISPR-Cas systems in heterologous prokaryotic species can be more comprehensively tested and studied.
[0123] Key advantages of the in vivo pooled screens described herein include: (1) Versatility - the plasmid design allows for the expression of multiple effectors and / or non-coding elements; the library cloning strategy allows for the expression of both computationally predicted crRNA transcriptional directions; (2) Comprehensive testing of activity mechanisms and functional parameters allows for evaluation of diverse interference mechanisms, including nucleic acid cleavage; co-occurrence of features such as transcription and plasmid DNA replication; and examination of flanking sequences for crRNA libraries to reliably determine 4N complexity-equivalent PAMs; (3) Sensitivity—pACYC184 is a low-copy plasmid, allowing high sensitivity for CRISPR-Cas activity because even small interference rates can eliminate the antibiotic resistance encoded by the plasmid; and (4) Efficiency - Optimized molecular biology steps that enable higher speed and throughput for RNA sequencing allow protein expression samples to be taken directly from surviving cells in the screen.
[0124] This in vivo pooled screen was used to evaluate the novel CRIST.091979 CRISPR-Cas family described herein by assessing its operable elements, mechanisms and parameters, and its ability to be active and reprogrammed in an engineered system outside of its endogenous cellular environment.
[0125] CRISPR effector activity and modification In some embodiments, the CRISPR effector of CRIST.091979 and the RNA guide form a binary complex, which may include other components. The binary complex is activated upon binding to a nucleic acid substrate (i.e., a sequence-specific substrate or a target nucleic acid) that is complementary to the spacer sequence in the RNA guide. In some embodiments, the sequence-specific substrate is double-stranded DNA. In some embodiments, the sequence-specific substrate is single-stranded DNA. In some embodiments, the sequence-specific substrate is single-stranded RNA. In some embodiments, the sequence-specific substrate is double-stranded RNA. In some embodiments, sequence specificity requires a perfect match between the spacer sequence in the RNA guide (e.g., crRNA) and the target substrate. In other embodiments, sequence specificity requires a partial (contiguous or non-contiguous) match between the spacer sequence in the RNA guide (e.g., crRNA) and the target substrate.
[0126] In some embodiments, the CRISPR effectors of the present invention have enzymatic activity, e.g., nuclease activity, over a wide range of pH conditions. In some embodiments, the nuclease has enzymatic activity, e.g., nuclease activity, at a pH of about 3.0 to about 12.0. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 4.0 to about 10.5. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 5.5 to about 8.5. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 6.0 to about 8.0. In some embodiments, the CRISPR effector has enzymatic activity at a pH of about 7.0.
[0127] In some embodiments, the CRISPR effectors of the invention have enzymatic activity, e.g., nuclease activity, in a temperature range of about 10° C. to about 100° C. In some embodiments, the CRISPR effectors of the invention have enzymatic activity in a temperature range of about 20° C. to about 90° C. In some embodiments, the CRISPR effectors of the invention have enzymatic activity at a temperature of about 20° C. to about 25° C. or at a temperature of about 37° C.
[0128] In some embodiments, the binary complex becomes activated upon binding to the target substrate. In some embodiments, the activated complex exhibits "multiple turnover" activity, such that upon acting on the target substrate (e.g., cleaving it), the activated complex remains in an activated state. In some embodiments, the activated binary complex exhibits "single turnover" activity, such that upon acting on the target substrate, the binary complex returns to an inactive state. In some embodiments, the activated binary complex exhibits non-specific (i.e., "collateral") cleavage activity, such that the complex cleaves a non-target nucleic acid. In some embodiments, the non-target nucleic acid is a DNA molecule (e.g., single- or double-stranded DNA). In some embodiments, the non-target nucleic acid is an RNA molecule (e.g., single- or double-stranded RNA).
[0129] In some embodiments in which the CRISPR effector of the present invention induces double-strand or single-strand breaks in a target nucleic acid (e.g., genomic DNA), the double-strand breaks can stimulate cell-intrinsic DNA repair pathways, including homologous recombination (HDR), non-homologous end joining (NHEJ), or alternative non-homologous end joining (A-NHEJ). NHEJ can repair a cleaved target nucleic acid without the need for a homologous template. This can result in the deletion or insertion of one or more nucleotides at the target locus. HDR can occur using a homologous template, such as donor DNA. The homologous template can include a sequence homologous to a sequence flanking the target nucleic acid cleavage site. In some instances, HDR can insert an exogenous polynucleotide sequence into the cleaved target locus. Modifications of the target DNA resulting from NHEJ and / or HDR can result in, for example, mutations, deletions, alterations, integrations, gene corrections, gene replacements, gene tagging, transgene knock-ins, gene disruption, and / or gene knockouts.
[0130] In some embodiments, the CRISPR effectors described herein can be fused to one or more peptide tags, including His tags, GST tags, FLAG tags, or myc tags. In some embodiments, the CRISPR effectors described herein can be fused to a detectable moiety, such as a fluorescent protein (e.g., green fluorescent protein or yellow fluorescent protein). In some embodiments, the CRISPR effectors and / or accessory proteins of the present disclosure are fused to a peptide or non-peptide moiety that directs or localizes the protein to a tissue, cell, or region of a cell. For example, the CRISPR effectors of the present disclosure can include a nuclear localization sequence (NLS), such as the SV40 (Simian Virus 40) NLS, c-Myc NLS, or other suitable monopartite NLS. The NLS can be fused to the N-terminus and / or C-terminus of the CRISPR effector and can be fused singly (i.e., a single NLS) or concatenated (e.g., a chain of two, three, four, etc. NLSs).
[0131] In some embodiments, at least one nuclear export signal (NES) is attached to the nucleic acid sequence encoding the CRISPR effector. In some embodiments, a C-terminal and / or N-terminal NLS or NES is attached for optimal expression and nuclear targeting in eukaryotic cells, such as human cells.
[0132] In embodiments in which a tag is fused to a CRISPR effector, such a tag can facilitate affinity-based or charge-based purification of the CRISPR effector, for example, by liquid chromatography or bead separation utilizing immobilized affinity or ion exchange reagents. As a non-limiting example, a recombinant CRISPR effector of the present disclosure may contain a polyhistidine (His) tag and, for purification, be loaded onto a chromatography column containing an immobilized metal ion (e.g., Zn chelated by a chelating ligand immobilized on the resin). 2+ , Ni 2+ , Cu 2+ The resin may be an individually prepared resin or a commercially available resin or pre-made column, such as the HisTrap FF column commercialized by GE Healthcare Life Sciences, Marlborough, Massachusetts. After the loading step, the column is optionally rinsed, for example, using one or more suitable buffer solutions, and the His-tagged protein is then eluted using a suitable elution buffer. Alternatively or additionally, if a recombinant CRISPR effector of the present disclosure utilizes a FLAG tag, such protein may be purified using immunoprecipitation methods known in the art. Other suitable purification methods for tagged CRISPR effectors or accessory proteins of the present disclosure will be apparent to those of skill in the art.
[0133] The proteins (e.g., CRISPR effector or accessory proteins) described herein can be delivered or used as either nucleic acid molecules or polypeptides. When nucleic acid molecules are used, the nucleic acid molecule encoding the CRISPR effector can be codon-optimized. Nucleic acids can be codon-optimized for use in any organism of interest, particularly human cells or bacteria. For example, nucleic acids can be codon-optimized for any non-human eukaryote, including mouse, rat, rabbit, dog, livestock, or non-human primates. Codon usage tables are readily available, for example, in the "Codon Usage Database," available at www.kazusa.orjp / codon / , and these tables can be adapted in a number of ways. See Nakamura et al. Nucl. Acids Res. 28:292 (2000) (incorporated herein by reference in its entirety). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available in Genet. Forge (Aptagen; Jacobus, PA) and other similar products are available.
[0134] In some examples, nucleic acids of the present disclosure encoding CRISPR effectors for expression in eukaryotic (e.g., human or other mammalian) cells include one or more introns, i.e., one or more non-coding sequences including a splice donor sequence at a first end (e.g., the 5' end) and a splice acceptor sequence at a second end (e.g., the 3' end). In various embodiments of the present disclosure, any suitable splice donor / splice acceptor can be used, including, without limitation, a Simian Virus 40 (SV40) intron, a β-globin intron, and a synthetic intron. Alternatively or additionally, nucleic acids of the present disclosure encoding CRISPR effectors or accessory proteins can include a transcription termination signal, such as a polyadenylation (polyA) signal, at the 3' end of the DNA coding sequence. In some examples, the polyA signal is located in close proximity to or adjacent to an intron, such as an SV40 intron.
[0135] Deactivated / inactivated CRISPR effectors The CRISPR enzymes described herein can be modified to have reduced nuclease activity, for example, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation compared to wild-type CRISPR effectors. Nuclease activity can be reduced by several methods well known in the art, for example, by introducing mutations into the nuclease domain of the protein. In some embodiments, catalytic residues of nuclease activity can be identified, and nuclease activity can be reduced by substituting these amino acid residues with different amino acid residues (e.g., glycine or alanine).
[0136] The inactivated CRISPR effector may comprise or be associated with one or more functional domains (e.g., via a fusion protein, a linker peptide, a "GS" linker, etc.). Such functional domains may have various activities, such as methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and switch activity (e.g., light-inducible). In some embodiments, the functional domains are Krüppel-associated box (KRAB), VP64, VP16, Fok1, P65, HSF1, MyoD1, and biotin-APEX.
[0137] Positioning one or more functional domains on an inactivated CRISPR effector allows the functional domain to be in the correct spatial orientation to affect the target with the functional effect it ascribes. For example, if the functional domain is a transcriptional activator (e.g., VP16, VP64, or p65), the transcriptional activator is positioned in a spatial orientation that allows it to affect the transcription of the target. Similarly, a transcriptional repressor is positioned to affect the transcription of the target, and a nuclease (e.g., Fok1) is positioned to cleave or partially cleave the target. In some embodiments, the functional domain is located at the N-terminus of the CRISPR effector. In some embodiments, the functional domain is located at the C-terminus of the CRISPR effector. In some embodiments, the inactivated CRISPR effector is modified to include a first functional domain at the N-terminus and a second functional domain at the C-terminus.
[0138] Split Enzyme The present disclosure also provides the split version of CRISPR effector described herein.Split version CRISPR effector can be advantageous for delivery.In some embodiments, CRISPR effector is split into two parts of enzyme, and they together comprise substantially functional CRISPR effector.
[0139] The division can be done in such a way that one or more catalytic domains are unaffected. The CRISPR effector may function as a nuclease or may be an inactivated enzyme that is essentially an RNA-binding protein with little or no catalytic activity (e.g., due to one or more mutations in its catalytic domain).
[0140] In some embodiments, the nuclease lobe and the α-helical lobe are expressed as separate polypeptides. These lobes do not interact with each other, but the RNA guide recruits them into a complex that mimics the activity of a full-length CRISPR effector and catalyzes site-specific DNA cleavage. The use of modified RNA guides abolishes split enzyme activity by preventing dimerization, allowing the development of an inducible dimerization system. Split enzymes are described, for example, in Wright et al., "Rational design of a split-Cas9 enzyme complex," Proc. Nat'l. Acad. Sci., 112.10 (2015): 2984-2989 (incorporated herein by reference in its entirety).
[0141] In some embodiments, the split enzyme can be fused to a dimerization partner, for example, by utilizing a rapamycin-sensitive dimerization domain. This allows the creation of a chemically inducible CRISPR effector for temporally controlling the activity of the CRISPR effector. By being split into two fragments in this way, the CRISPR effector can be chemically inducible, and the rapamycin-sensitive dimerization domain can be used for the controlled reassembly of the CRISPR effector.
[0142] The split point is typically designed in silico and cloned into the construct. During this process, mutations may be introduced into the split enzyme, and non-functional domains may be removed. In some embodiments, the two portions or fragments (i.e., N-terminal and C-terminal fragments) of the split CRISPR effector can form a complete CRISPR effector that contains, for example, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% of the sequence of the wild-type CRISPR effector.
[0143] Self-activating or inactivating enzymes The CRISPR effectors described herein may be designed to be self-activating or self-inactivating. In some embodiments, the CRISPR effector is self-inactivating. For example, a target sequence can be introduced into a construct encoding the CRISPR effector. The CRISPR effector can then cleave the target sequence, thereby causing the construct encoding the enzyme to self-inactivate its expression. Methods for constructing self-inactivating CRISPR systems are described, for example, in Epstein et al., "Engineering a Self-Inactivating CRISPR System for AAV Vectors," Mol. Ther., 24 (2016): S50 (incorporated herein by reference in its entirety).
[0144] In some other embodiments, an additional RNA guide expressed under the control of a weak promoter (e.g., a 7SK promoter) can target the nucleic acid sequence encoding the CRISPR effector and interfere with and / or prevent its expression (e.g., by preventing transcription and / or translation of the nucleic acid). Transfecting a cell with a vector expressing a CRISPR effector, an RNA guide, and an RNA guide targeting the nucleic acid encoding the CRISPR effector can lead to efficient destruction of the nucleic acid encoding the CRISPR effector, reducing CRISPR effector levels and thus limiting genome editing activity.
[0145] In some embodiments, the genome editing activity of CRISPR effectors can be regulated through endogenous RNA signatures (e.g., miRNAs) in mammalian cells. A CRISPR effector switch can be created by using an miRNA-complementary sequence in the 5'-UTR of an mRNA encoding a CRISPR effector. This switch selectively and efficiently responds to miRNAs in target cells. Thus, this switch can differentially control genome editing by sensing endogenous miRNA activity within a heterogeneous cell population. Therefore, this switch system may provide a framework for cell-type-selective genome editing and cell engineering based on intracellular miRNA information (Hirosawa et al., "Cell-type-specific genome editing with a microRNA-responsive CRISPR-Cas9 switch," Nucl. Acids Res., 2017 Jul. 27;45(13):e118).
[0146] Inducible CRISPR effectors CRISPR effectors can be inducible, e.g., light- or chemically inducible. This mechanism allows for activation of functional domains in the CRISPR enzyme. Light-inducibility can be achieved by various methods known in the art, for example, by designing a fusion complex in which the CRY2PHR / CIBN pair is used in a split CRISPR effector (see, e.g., Konermann et al., "Optical control of mammalian endogenous transcription and epigenetic states," Nature, 500.7463 (2013):472). Chemical inducibility can be achieved, for example, by designing a fusion complex in which the FKBP / FRB (FK506-binding protein / FKBP rapamycin-binding domain) pair is used in a split CRISPR effector. Rapamycin is required for the formation of the fusion complex, thus activating the CRISPR effector (see, e.g., Zetsche et al., "A split-Cas9 architecture for inducible genome editing and transcription modulation," Nature, 500.7463 (2013):472). Biotech.,33.2(2015):139-142).
[0147] Furthermore, expression of CRISPR effectors can be regulated by inducible promoters, such as tetracycline- or doxycycline-controlled transcriptional activation (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. When delivered as RNA, expression of RNA-targeting effector proteins may be regulated by riboswitches capable of sensing small molecules like tetracycline (see, e.g., Goldfless et al., "Direct and specific chemical control of eukaryotic translation with a synthetic RNA-protein interaction," Nucl. Acids Res., 40.9 (2012):e64-e64).
[0148] Various embodiments of inducible CRISPR effectors and inducible CRISPR systems are described, for example, in U.S. Pat. No. 8,871,445, U.S. Patent Application Publication No. 20160208243, and WO 2016205764, each of which is herein incorporated by reference in its entirety.
[0149] Functional mutations Various mutations or modifications can be introduced into the CRISPR effectors described herein to improve specificity and / or robustness. In some embodiments, amino acid residues that recognize a protospacer adjacent motif (PAM) are identified. The CRISPR effectors described herein can be further modified to recognize different PAMs, for example, by substituting the amino acid residues that recognize the PAM with other amino acid residues. In some embodiments, the CRISPR effectors can be further modified to recognize different PAMs, for example, 5'-NTTN-3', 5'-NTTR-3', 5'-RTTR-3', 5'-TNNT-3', 5'-TNRT-3', 5'-TSRT-3', 5'-TGRT-3', 5'-TNRY-3', 5'-TTNR-3', 5'-TTYR-3', 5'-TTTR-3', 5'-TTCV-3', 5'-DTYR-3', 5'-WTTR-3', 5'-NNR-3', 5'-NYR- 3', 5'-YYR-3', 5'-TYR-3', 5'-TTN-3', 5'-TTR-3', 5'-CNT-3', 5'-NGG-3', 5'-BGG-3', or 5'-R-3', where "N" is any nucleotide, "B" is C or G or T, "D" is A or G or T, "R" is A or G, "S" is G or C, "V" is A or C or G, "W" is A or T, and "Y" is C or T.
[0150] In some embodiments, the CRISPR effectors described herein may have one or more functional activities altered by mutating one or more amino acid residues. For example, in some embodiments, the helicase activity of a CRISPR effector is altered by mutating one or more amino acid residues. In some embodiments, the nuclease activity (e.g., endonuclease activity or exonuclease activity) of a CRISPR effector is altered by mutating one or more amino acid residues. In some embodiments, the ability of a CRISPR effector to functionally associate with an RNA guide is altered by mutating one or more amino acid residues. In some embodiments, the ability of a CRISPR effector to functionally associate with a target nucleic acid is altered by mutating one or more amino acid residues.
[0151] In some embodiments, the CRISPR effectors described herein have the ability to cleave a target nucleic acid molecule. In some embodiments, the CRISPR effector cleaves both strands of a target nucleic acid molecule. However, in some embodiments, the cleavage activity of a CRISPR effector is modified by mutating one or more amino acid residues. For example, in some embodiments, a CRISPR effector may contain one or more mutations that increase the ability of the CRISPR effector to cleave a target nucleic acid. In another example, in some embodiments, a CRISPR effector may contain one or more mutations that result in the enzyme being incapable of cleaving a target nucleic acid. In other embodiments, a CRISPR effector may contain one or more mutations that result in the enzyme having the ability to cleave a strand of a target nucleic acid (i.e., nickase activity). In some embodiments, a CRISPR effector has the ability to cleave a strand of a target nucleic acid that is complementary to the strand to which an RNA guide is hybridized. In some embodiments, a CRISPR effector has the ability to cleave a strand of a target nucleic acid to which an RNA guide is hybridized.
[0152] In some embodiments, one or more residues of the CRISPR effector disclosed herein are mutated to arginine moiety.In some embodiments, one or more residues of the CRISPR effector disclosed herein are mutated to glycine moiety.In some embodiments, one or more residues of the CRISPR effector disclosed herein are mutated based on the consensus residue of the systematic alignment of the CRISPR effector disclosed herein.
[0153] In some embodiments, the CRISPR effectors described herein may be engineered to contain deletions of one or more amino acid residues to reduce the size of the enzyme while retaining one or more desired functional activities (e.g., nuclease activity and the ability to functionally interact with an RNA guide). This truncated CRISPR effector may be advantageously used in combination with a payload-limited delivery system.
[0154] In one aspect, the present disclosure provides nucleic acid sequences that are at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the nucleic sequences described herein while maintaining the domain organization shown in FIG. 2. In another aspect, the present disclosure also provides amino acid sequences that are at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to an amino acid sequence described herein while maintaining the domain organization shown in FIG. 2.
[0155] In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., contiguous or non-contiguous nucleotides) that is identical to a sequence described herein. In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., contiguous or non-contiguous nucleotides) that differs from a sequence described herein.
[0156] In some embodiments, the amino acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., contiguous or non-contiguous amino acid residues) that is identical to a sequence described herein. In some embodiments, the amino acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., contiguous or non-contiguous amino acid residues) that differs from a sequence described herein.
[0157] To determine the percent identity of two amino acid sequences or two nucleic acid sequences, the sequences are aligned for optimal comparison (e.g., gaps may be introduced into one or both of the first and second amino acid or nucleic acid sequences to ensure optimal alignment, and non-homologous sequences may be ignored for comparison purposes). Generally, the length of the reference sequence aligned for comparison purposes should be at least 80% of the length of the reference sequence, and in some embodiments, at least 90%, 95%, or 100% of the length of the reference sequence. The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between two sequences is a function of the number of gaps that need to be introduced for optimal alignment of the two sequences and the number of identical positions shared by the sequences, taking into account the length of each gap. For purposes of this disclosure, comparison of sequences and determination of percent identity between two sequences may be accomplished using a Blossum 62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5.
[0158] In some embodiments, the nuclease comprises a sequence set forth as PX1X2X3X4F (SEQ ID NO:216), where X1 is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M. In some embodiments, the sequence set forth in SEQ ID NO:216 is an N-terminal sequence. In some embodiments, the nuclease comprises a sequence set forth as RX1X2X3L (SEQ ID NO:217), where X1 is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K. In some embodiments, the nuclease comprises a sequence set forth as NX1YX2 (SEQ ID NO:218), where X1 is I or L or F and X2 is K or R or V or E. In some embodiments, the nuclease comprises a sequence set forth as KX1X2X3FAX4X5KD (SEQ ID NO: 219), where X1 is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 219 is the C-terminal sequence. In some embodiments, the nuclease comprises a sequence set forth as LX1NX2 (SEQ ID NO: 220), where X1 is G or S or C or T and X2 is N or Y or K or S. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 220 is the C-terminal sequence. In some embodiments, the nuclease comprises the sequence set forth as PX1X2X3X4SQX5DS (SEQ ID NO: 221), where X1 is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M or T or I. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 221 is the C-terminal sequence.In some embodiments, the nuclease comprises the sequence set forth as KX1X2VRX3X4QEX5H (SEQ ID NO: 222), where Xi is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 222 is the C-terminal sequence. In some embodiments, the nuclease comprises the sequence set forth as X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223), where X1 is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L. In some embodiments of any of the systems described herein, the sequence of SEQ ID NO: 223 is the C-terminal sequence.
[0159] RNA-guided and RNA-guided modifications In some embodiments, the RNA guides described herein comprise uracil (U). In some embodiments, the RNA guides described herein comprise thymine (T). In some embodiments, the direct repeat sequences of the RNA guides described herein comprise uracil (U). In some embodiments, the direct repeat sequences of the RNA guides described herein comprise thymine (T). In some embodiments, the direct repeat sequences according to Table 2 or Table 8 comprise a sequence comprising uracil at one or more positions shown as thymine in the corresponding sequence in Table 2 or Table 8.
[0160] In some embodiments, direct repeat comprises only one copy of the sequence that repeats in endogenous CRISPR array.In some embodiments, direct repeat is the full-length sequence that is adjacent (for example, flanking) to one or more spacer sequences found in endogenous CRISPR array.In some embodiments, direct repeat is the part (for example, processed part) of the full-length sequence that is adjacent (for example, flanking) to one or more spacer sequences found in endogenous CRISPR array.
[0161] Spacers and direct repeats The RNA guide spacer length may range from about 15 to 55 nucleotides. The RNA guide spacer length may range from about 20 to 45 nucleotides. In some embodiments, the RNA guide spacer length is at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, or at least 22 nucleotides. In some embodiments, the spacer length is 15-17 nucleotides, 15-23 nucleotides, 16-22 nucleotides, 17-20 nucleotides, 20-24 nucleotides (e.g., 20, 21, 22, 23, or 24 nucleotides), 23-25 nucleotides (e.g., 23, 24, or 25 nucleotides), 24-27 nucleotides, 27-30 nucleotides, 30-45 nucleotides (e.g., 30, 31, 32, 33, 34, 35, 40, or 45 nucleotides), 30 or 35-40 nucleotides, 41-45 nucleotides, 45-50 nucleotides, or more.
[0162] In some embodiments, the direct repeat length of the RNA guide is at least 16 nucleotides, or is 16 to 20 nucleotides (e.g., 16, 17, 18, 19, or 20 nucleotides). In some embodiments, the direct repeat length of the RNA guide is about 19 to about 40 nucleotides.
[0163] Exemplary direct repeat sequences (e.g., direct repeat sequences of pre-crRNA (e.g., unprocessed crRNA) or mature crRNA (e.g., processed crRNA)) are shown in Table 2. See also Table 8.
[0164] [Table 2]
[0165] In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:1, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:57. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:2, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:58. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:3, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:59.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:60. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 62 or SEQ ID NO: 213. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 14, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 128.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 15, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 63. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 17, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 130. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 18, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:70.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:21, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:72. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:22, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:73. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:23, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:74.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:24, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:63. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:27, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:76. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:28, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:77.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:29, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:139. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 31, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 58. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 32, and the direct repeat sequence is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 80. 7%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical nucleotide sequence. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 35, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 77. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 36, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 139. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 38, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 80.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 39, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 58. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:41, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:83. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:42, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:84.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:44, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:86. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:45, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:130. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:46, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:84.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:47, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:87. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:48, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:88. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:51, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:84.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:53, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:84. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:55, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:88. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 56, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 90.
[0166] In some embodiments, the RNA guide comprises a direct repeat sequence as set forth in Figure 3. For example, in some embodiments, the RNA guide comprises a direct repeat of the consensus sequence shown in Figure 3 or a portion of the consensus sequence shown in Figure 3. In some embodiments, the RNA guide comprises a direct repeat having a sequence set forth as X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224), where X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T. For example, in some embodiments, the RNA guide comprises a direct repeat having a sequence set forth as ATTGTTGDA (SEQ ID NO: 225). In some embodiments, SEQ ID NO: 224 is proximal to the 5' end of the direct repeat. In some embodiments, SEQ ID NO: 225 is proximal to the 5' end of the direct repeat. In some embodiments, the RNA guide comprises a direct repeat having a sequence set forth as X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226), where X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C. For example, in some embodiments, the RNA guide comprises a direct repeat having a sequence set forth as TTTTWTARG (SEQ ID NO: 227). In some embodiments, the RNA guide comprises a direct repeat having a sequence set forth as X1X2X3AC (SEQ ID NO: 228), where X1 is A or C or G, X2 is C or A, and X3 is A or C. For example, in some embodiments, the RNA guide comprises a direct repeat having the sequence set forth as ACAAC (SEQ ID NO: 229). In some embodiments, SEQ ID NO: 228 is adjacent to the 3' end of the direct repeat. In some embodiments, SEQ ID NO: 229 is adjacent to the 3' end of the direct repeat.
[0167] In some embodiments, the spacer of the RNA guide binds to a target nucleic acid adjacent to a PAM sequence in Table 3. For example, in some embodiments, a complex of an effector and an RNA guide disclosed herein binds to a target nucleic acid adjacent to a PAM sequence shown in Table 3.
[0168] [Table 3]
[0169] [Table 4]
[0170] In some embodiments, the RNA guide further comprises a tracrRNA. In some embodiments, the tracrRNA is not required (e.g., the tracrRNA is optional). In some embodiments, the tracrRNA is part of a non-coding sequence set forth in Table 9. For example, in some embodiments, the tracrRNA is a sequence in Table 4.
[0171] [Table 5]
[0172] [Table 6]
[0173] [Table 7]
[0174] [Table 8]
[0175] [Table 9]
[0176]
Table 10
[0177] In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:1, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:152, SEQ ID NO:153, or SEQ ID NO:154. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:2, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:155, SEQ ID NO:156, SEQ ID NO:157, or SEQ ID NO:158. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:3, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:159, SEQ ID NO:160, or SEQ ID NO:161.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 14, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 162. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 17, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, or SEQ ID NO: 166. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 18, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 167 or SEQ ID NO: 168.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:21, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:169, SEQ ID NO:170, or SEQ ID NO:171. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:22, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, or SEQ ID NO:175. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:23, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:176, SEQ ID NO:177, SEQ ID NO:178, or SEQ ID NO:179.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:27, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:180 or SEQ ID NO:181. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:29, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:182, SEQ ID NO:183, or SEQ ID NO:184. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:31, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:185, SEQ ID NO:186, SEQ ID NO:187, or SEQ ID NO:188.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 32, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 189 or SEQ ID NO: 190. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:36, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:182, SEQ ID NO:183, or SEQ ID NO:184. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO: 38, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO: 189 or SEQ ID NO: 190.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:39, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:185, SEQ ID NO:186, SEQ ID NO:187, or SEQ ID NO:188. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:41, and the tracrRNA sequence is at least 80% (e.g., 81%, 82%, 83%, 84%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:191, SEQ ID NO:192, SEQ ID NO:193, or SEQ ID NO:194. 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical nucleotide sequence. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:43, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:197, SEQ ID NO:198, or SEQ ID NO:199. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:44, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:195 or SEQ ID NO:196. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:45, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:163, SEQ ID NO:164, SEQ ID NO:165, or SEQ ID NO:166.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:48, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:200, SEQ ID NO:201, or SEQ ID NO:202. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:52, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:197, SEQ ID NO:198, or SEQ ID NO:199. In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:55, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:200, SEQ ID NO:201, or SEQ ID NO:202.In some embodiments, the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of SEQ ID NO:56, and the tracrRNA sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence of SEQ ID NO:203 or SEQ ID NO:204.
[0178] RNA guide sequences may be modified in a way that allows for successful CRISPR complex formation and target binding, but simultaneously disallows successful nuclease activity (i.e., no nuclease activity / no indels). Such modified guide sequences are referred to as "dead guides" or "dead guide sequences." Such dead guides or dead guide sequences may be catalytically inactive or conformationally inactive with respect to nuclease activity. Dead guide sequences are typically shorter than the respective guide sequences that result in active RNA cleavage. In some embodiments, dead guides are 5%, 10%, 20%, 30%, 40%, or 50% shorter than the respective guide RNAs that have nuclease activity. Dead guide sequences for RNA guides may be 13-15 nucleotides in length (e.g., 13, 14, or 15 nucleotides in length), 15-19 nucleotides in length, or 17-18 nucleotides in length (e.g., 17 nucleotides in length).
[0179] Thus, in one aspect, the present disclosure provides a non-naturally occurring or engineered CRISPR system comprising a functional CLUST.091979 CRISPR effector as described herein and an RNA guide, wherein the RNA guide comprises a dead guide sequence, such that the RNA guide is capable of hybridizing to a target sequence such that the CRISPR system is directed to a genomic locus of interest in a cell without detectable cleavage activity. A detailed description of dead guides is provided, for example, in International Publication No. WO2016094872, which is incorporated herein by reference in its entirety.
[0180] Inducible RNA guides The RNA guide can be made as a component of an inducible system. The inducible nature of this system allows for spatiotemporal control of gene editing or gene expression. In some embodiments, the stimulus for the inducible system can include, for example, electromagnetic radiation, acoustic energy, chemical energy, and / or thermal energy.
[0181] In some embodiments, transcription of the RNA guide can be regulated by an inducible promoter, such as tetracycline- or doxycycline-controlled transcriptional activation (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. Other examples of inducible systems include, for example, small molecule two-hybrid transcription activation systems (FKBP, ABA, etc.), light-inducible systems (phytochrome, LOV domain, or cryptochrome), or light-inducible transcription effectors (LITEs). These inducible systems are described, for example, in WO2016205764 and U.S. Pat. No. 8,795,965, each of which is incorporated herein by reference in its entirety.
[0182] chemical modification Chemical modifications can be applied to the phosphate backbone, sugar, and / or base of the guide RNA. Backbone modifications such as phosphorothioates modify the charge on the phosphate backbone and aid in oligonucleotide delivery and nuclease resistance (see, e.g., Eckstein, "Phosphorothioates, essential components of therapeutic oligonucleotides," Nucl. Acid Ther., 24 (2014), pp. 374-387); sugar modifications such as 2'-O-methyl (2'-OMe), 2'-F, and locked nucleic acid (LNA) enhance both base pairing and nuclease resistance (see, e.g., Allerson et al., "Fully 2'-modified oligonucleotide duplexes with improved in vitro potency and stability compared to unmodified small interfering RNA," J. Med. Chem., 48.4 (2005): 901-904). Chemically modified bases, such as 2-thiouridine or N6-methyladenosine, among others, can allow for either stronger or weaker base pairing (see, e.g., Bramsen et al., "Development of therapeutic-grade small interfering RNAs by chemical engineering," Front. Genet., 2012 Aug. 20;3:154). In addition, RNA is suitable for conjugation at both the 5' and 3' ends with various functional moieties, including fluorescent dyes, polyethylene glycol, or proteins.
[0183] A wide variety of modifications can be applied to chemically synthesized RNA guide molecules. For example, oligonucleotides can be modified with 2'-OMe to improve nuclease resistance and alter the binding energy of Watson-Crick base pairing. Furthermore, 2'-OMe modifications can affect how the oligonucleotide interacts with transfection reagents, proteins, or any other molecules in the cell. The effects of these modifications can be determined through empirical testing.
[0184] In some embodiments, the RNA guide comprises one or more phosphorothioate modifications, hi some embodiments, the RNA guide comprises one or more locked nucleic acids for the purposes of enhancing base pairing and / or increasing nuclease resistance.
[0185] For an overview of these chemical modifications, see, for example, Kelley et al., "Versatility of chemically synthesized guides." See, for example, “RNAs for CRISPR-Cas9 genome editing,” J. Biotechnol. 2016 Sep 10;233:74-83; International Publication No. WO 2016205764; and U.S. Patent No. 8,795,965 (each of which is incorporated by reference in its entirety).
[0186] Array Modification The sequences and lengths of the RNA guides, tracrRNA, and crRNA described herein can be optimized. In some embodiments, the optimized length of the RNA guide may be determined by identifying processed forms of the tracrRNA and / or crRNA or by empirical length studies of the crRNA RNA guide.
[0187] The RNA guide can also include one or more aptamer sequences. Aptamers are oligonucleotide or peptide molecules capable of binding to specific target molecules. Aptamers may be specific for gene effectors, gene activators, or gene repressors. In some embodiments, the aptamer may be specific for a protein that, in turn, is specific for and recruits / binds to a specific gene effector, gene activator, or gene repressor. The effector, activator, or repressor can exist in the form of a fusion protein. In some embodiments, the RNA guide has two or more aptamer sequences specific for the same adaptor protein. In some embodiments, the two or more aptamer sequences are specific for different adaptor proteins. Examples of adaptor proteins include MS2, PP7, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. Thus, in some embodiments, the aptamer is selected from binding proteins that specifically bind to any one of the adaptor proteins described herein. In some embodiments, the aptamer sequence is the MS2 loop. For a detailed description of aptamers, see, for example, Nowak et al., "Guide RNA engineering for versatile Cas9 functionality," Nucl. Acid. Res., 2016 Nov 16;44(20):9555-9564; and WO 2016205764 (each of which is incorporated herein by reference in its entirety).
[0188] Guide: Target sequence match requirements In a CRISPR system, the degree of complementarity between a guide sequence and its corresponding target sequence can be about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100%. To reduce off-target interactions, e.g., to reduce guide interactions with target sequences with low complementarity, mutations can be introduced into the CRISPR system to enable the CRISPR system to distinguish between target sequences and off-target sequences with greater than 80%, 85%, 90%, or 95% complementarity. In some embodiments, the degree of complementarity is 80%-95%, e.g., about 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, or 95% (e.g., to distinguish between a target having 18 nucleotides and an 18-nucleotide off-target with 1, 2, or 3 mismatches). Thus, in some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is greater than 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 99.9%, hi some embodiments, the degree of complementarity is 100%.
[0189] It is known in the art that perfect complementarity is not a requirement, as long as there is sufficient complementarity to be functional. The introduction of mismatches, for example, one or more mismatches, for example, one or two mismatches, between the spacer sequence and the target sequence, including the location of the mismatch along the spacer / target, can be used to adjust the cleavage efficiency. The more centrally located the mismatch, for example, a double mismatch (i.e., not at the 3' or 5' end), the greater the impact on cleavage efficiency. Thus, the selection of the mismatch location along the spacer sequence can adjust the cleavage efficiency. For example, if less than 100% cleavage of the target is desired (e.g., in a cell population), one or two mismatches between the spacer and the target sequence may be introduced into the spacer sequence.
[0190] How to use the CRISPR system The CRISPR system described herein has a wide variety of utilities, including modification (e.g., deletion, insertion, rearrangement, inactivation, or activation) of target polynucleotides in numerous cell types. The CRISPR system has a wide range of applications, for example, in DNA / RNA detection (e.g., specific high sensitivity enzymatic reporter unlocking (SHERLOCK)), nucleic acid tracking and labeling, enrichment assays (extraction of desired sequences from background), detection of circulating tumor DNA, next-generation library preparation, drug screening, disease diagnosis and prognosis, and treatment of various genetic disorders.
[0191] DNA / RNA detection In one embodiment, the CRISPR system described herein can be used in DNA / RNA detection.By reprogramming single-effector RNA-guided DNase with CRISPR RNA (crRNA), it can provide a platform for specific single-stranded DNA (ssDNA) sensing.When recognizing its DNA target, the activated V-type single-effector DNA-guided DNase is involved in the "collateral" cleavage of neighboring non-target ssDNA.This collateral cleavage activity programmed by crRNA allows the CRISPR system to detect the presence of specific DNA through the non-specific degradation of labeled ssDNA.
[0192] In DNA detection applications, collateral ssDNA activity can be combined with a reporter, such as in a method called DNA Endonuclease-Targeted CRISPR trans reporter (DETECTR), which achieves attomolar DNA detection sensitivity (see, e.g., Chen et al., Science, 360(6387):436-439, 2018, which is incorporated herein by reference in its entirety). One application using the enzymes described herein is the degradation of nonspecific ssDNA in an in vitro environment. A "reporter" ssDNA molecule linked to a fluorophore and a quencher can also be added to this in vitro system along with an unknown DNA sample (either single-stranded or double-stranded). Upon recognition of the target sequence in the unknown piece of DNA, the effector complex cleaves the reporter ssDNA, producing a fluorescent readout.
[0193] In another embodiment, the SHERLOCK method (Specific High Sensitivity Enzymatic Reporter Unlocking) also provides an in vitro nucleic acid detection platform with attomolar (or single molecule) sensitivity based on nucleic acid amplification and collateral cleavage of a reporter ssDNA, enabling real-time detection of targets. The use of CRISPR in SHERLOCK is described in detail, for example, in Gootenberg, et al. "Nucleic acid detection with CRISPR-Cas13a / C2c2," Science, 356(6336):438-442(2017), which is incorporated herein by reference in its entirety.
[0194] In some embodiments, the CRISPR systems described herein can be used in multiplexed error-robust fluorescence in situ hybridization (MERFISH). Such methods are described, for example, in Chen et al., "Spatially resolved, highly multiplexed RNA profiling in single cells,” Science, 2015 Apr 24;348(6233):aaa6090, which is incorporated herein by reference in its entirety.
[0195] Nucleic Acid Tracking and Labeling Cellular processes depend on a network of molecular interactions between proteins, RNA, and DNA. Accurate detection of protein-DNA and protein-RNA interactions is key to understanding such processes. In vitro proximity labeling techniques use affinity tags combined with reporter groups, such as photoactivatable groups, to label polypeptides and RNAs in vitro near a protein or RNA of interest. After ultraviolet irradiation, the photoactivatable groups react with proteins and other molecules in close proximity to the tagged molecule, thereby labeling them. The labeled interacting molecules can then be recovered and identified. This RNA targeting effector protein can be used, for example, to target probes to selected RNA sequences. Such applications can also be applied to in vivo imaging of disease or difficult-to-culture cell types in animal models. Methods for tracking and labeling nucleic acids are described, for example, in U.S. Pat. No. 8,795,965; WO 2016205764; and WO 2017070605 (each of which is incorporated by reference in its entirety).
[0196] High-throughput screening The CRISPR system described herein can be used to prepare next-generation sequencing (NGS) libraries. For example, to create cost-effective NGS libraries, the CRISPR system can be used to disrupt the coding sequence of target genes, and simultaneously screen the clones transfected with CRISPR effectors by next-generation sequencing (for example, by Ion Torrent PGM system). For detailed descriptions of NGS library preparation methods, see, for example, Bell et al., "A high-throughput screening strategy for detecting CRISPR-Cas9 induced mutations using next-generation sequencing," BMC Genomics, 15.1 (2014): 1002 (which is incorporated herein by reference in its entirety).
[0197] Engineered Cells Microorganisms (e.g., E. coli, yeast, and microalgae) are widely used in synthetic biology. Developments in synthetic biology have broad utility, including various clinical applications. For example, programmable CRISPR systems can be used to split proteins of toxic domains for targeted cell death using, for example, cancer-associated RNAs as target transcripts. Furthermore, fusion complexes with appropriate effectors, such as kinases or enzymes, can affect pathways involving protein-protein interactions in synthetic biology systems.
[0198] In some embodiments, an RNA guide sequence that targets a phage sequence can be introduced into a microorganism. Thus, the present disclosure also provides methods for "vaccinating" a microorganism (e.g., a production strain) against phage infection.
[0199] In some embodiments, the CRISPR systems provided herein can be used to engineer microorganisms to, for example, improve yields or improve fermentation efficiency. For example, the CRISPR systems described herein can be used to engineer microorganisms such as yeast to produce biofuels or biopolymers from fermentable sugars, or to degrade plant-derived lignocellulose from agricultural waste as a source of fermentable sugars. More specifically, the methods described herein can be used to modify the expression of endogenous genes required for biofuel production and / or to modify endogenous genes that may interfere with biofuel synthesis. These microbial engineering methods are described, for example, in Verwaal et al., "CRISPR / Cpf1 enables fast and simple genome editing of Saccharomyces cerevisiae,”Yeast,2017 Sep 8.doi:10.1002 / yea.3278; and Hlavova et al., “Improving microalgae for biotechnology-from Genetics to Synthetic Biology,” Biotechnol. Adv., 2015 Nov 1;33:1194-203 (each of which is incorporated herein by reference in its entirety).
[0200] In some embodiments, the CRISPR system provided herein can be used to engineer eukaryotic cells or eukaryotic organisms.For example, the CRISPR system described herein can be used to engineer eukaryotic cells, including but not limited to plant cells, fungal cells, mammalian cells, reptile cells, insect cells, bird cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, or human cells.In some embodiments, the eukaryotic cells are in vitro culture.In some embodiments, the eukaryotic cells are in vivo.In some embodiments, the eukaryotic cells are ex vivo.
[0201] In some embodiments, the cells are derived from a cell line. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, 293T, MF7, K562, HeLa, and transgenic variants thereof. Cell lines are available from a variety of sources known to those skilled in the art (e.g., American (See American Type Culture Collection (ATCC) (Manassas, Va.)). In some embodiments, cells transfected with one or more nucleic acids (e.g., a nuclease polypeptide-encoding vector and an RNA guide) are used to establish new cell lines that contain one or more vector-derived sequences to establish new cell lines that contain modifications of the target nucleic acid or target locus. In some embodiments, the cells are immortal or immortalized cells.
[0202] In some embodiments, the cell is a primary cell. In some embodiments, the cell is a stem cell, such as a totipotent stem cell (e.g., multipotent), pluripotent stem cell, multipotent stem cell, oligopotent stem cell, or unipotent stem cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC) or is derived from an iPSC. In some embodiments, the cell is a differentiated cell. For example, in some embodiments, the differentiated cell is a muscle cell (e.g., myocyte), adipocyte (e.g., adipocyte), bone cell (e.g., osteoblast, osteocyte, osteoclast), blood cell (e.g., monocyte, lymphocyte, neutrophil, eosinophil, basophil, macrophage, erythrocyte, or platelet), nerve cell (e.g., neuron), epithelial cell, immune cell (e.g., lymphocyte, neutrophil, monocyte, or macrophage), hepatocyte (e.g., hepatocyte), fibroblast, or germ cell. In some embodiments, the cell is a terminally differentiated cell. For example, in some embodiments, the terminally differentiated cell is a neuron, adipocyte, cardiomyocyte, skeletal muscle cell, epidermal cell, or intestinal cell. In some embodiments, the cell is a mammalian cell, e.g., a human cell or a murine cell. In some embodiments, the murine cell is derived from a wild-type mouse, an immunosuppressed mouse, or a disease-specific mouse model.
[0203] gene drive Gene drives are a phenomenon that favors the inheritance of a particular gene or group of genes. Gene drives can be constructed using the CRISPR systems described herein. For example, a CRISPR system can be designed to target and destroy a specific allele of a gene, causing cells to copy a second allele and fix the sequence. This copying converts the first allele to the second allele, increasing the likelihood that the second allele will be inherited by offspring. For detailed methods of how to construct a gene drive using the CRISPR systems described herein, see, for example, Hammond et al., "A CRISPR-Cas9 gene drive system targeting female reproduction." in the malaria mosquito vector Anopheles gambiae,” Nat. Biotechnol., 2016 Jan;34(1):78-83, which is incorporated herein by reference in its entirety.
[0204] Pooled Screening As described herein, pooled CRISPR screening is a powerful tool for identifying genes involved in biological mechanisms such as cell proliferation, drug resistance, and viral infection. Cells are bulk-transduced with a library of RNA guide-encoding vectors described herein, and the distribution of gRNAs is measured before and after a selective challenge. Pooled CRISPR screens work well for mechanisms affecting cell survival and proliferation and can be extended to measuring the activity of individual genes (e.g., by using engineered reporter cell lines). Arrayed CRISPR screens, in which only one gene is targeted at a time, allow for the use of RNA-seq as a readout. In some embodiments, the CRISPR system described herein can be used for single-cell CRISPR screens. For a detailed description of pooled CRISPR screening, see, for example, Datlinger et al., "Pooled CRISPR screening with single-cell transcriptome readout," Nat. Methods., 2017 Mar;14(3):297-301, which is incorporated herein by reference in its entirety.
[0205] Saturation mutagenesis ("bashing") The CRISPR system described herein can be used for in situ saturation mutagenesis. In some embodiments, a pooled RNA-guided library can be used to perform in situ saturation mutagenesis of specific genes or regulatory elements. Such methods can reveal the critical minimal features and individual vulnerabilities of those genes or regulatory elements (e.g., enhancers). These methods are described, for example, in Canver et al., "BCL11A enhancer dissection by Cas9-mediated in situ saturating mutagenesis," Nature, 2015 Nov 12;527(7577):192-7 (which is incorporated herein by reference in its entirety).
[0206] Therapeutic Applications In some embodiments, the CRISPR systems described herein can be used to edit a target nucleic acid to modify the target nucleic acid (e.g., by inserting, deleting, or mutating one or more amino acid residues). For example, in some embodiments, the CRISPR systems described herein include an exogenous donor template nucleic acid (e.g., a DNA molecule or an RNA molecule) comprising a desired nucleic acid sequence. Upon resolution of a cleavage event induced by a CRISPR system described herein, the cell's molecular machinery can utilize the exogenous donor template nucleic acid in repairing and / or degrading the cleavage event. Alternatively, the cell's molecular machinery can utilize an endogenous template in repairing and / or degrading the cleavage event. In some embodiments, the CRISPR systems described herein can be used to modify a target nucleic acid to produce insertions, deletions, and / or point mutations. In some embodiments, the insertion is an intact insertion (i.e., insertion of the intended nucleic acid sequence into the target nucleic acid does not result in additional unintended nucleic acid sequences upon resolution of the cleavage event). The donor template nucleic acid can be a double-stranded or single-stranded nucleic acid molecule (e.g., DNA or RNA). Methods for designing exogenous donor template nucleic acids are described, for example, in International Publication No. WO2016094874, the entire contents of which are expressly incorporated herein by reference.
[0207] In another aspect, the present disclosure provides use of the system described herein in a method selected from the group consisting of RNA sequence-specific interference; RNA sequence-specific gene regulation; screening of RNA, RNA products, lncRNA, non-coding RNA, nuclear RNA, or mRNA; mutagenesis; inhibition of RNA splicing; fluorescence in situ hybridization; breeding; inducing cell dormancy; inducing cell cycle arrest; reducing cell growth and / or cell proliferation; inducing cell anergy; inducing cell apoptosis; inducing cell necrosis; inducing cell death; or inducing programmed cell death.
[0208] The CRISPR system described herein can have various therapeutic applications. In some embodiments, the novel CRISPR system can be used to treat various diseases and disorders, for example, genetic disorders (e.g., single gene disorders) or diseases that can be treated by nuclease activity (e.g., Pcsk9 targeting or BCL11a targeting). In some embodiments, the methods described herein are used to treat subjects, for example, mammals, such as human patients. Mammalian subjects can also be domesticated mammals, such as dogs, cats, horses, monkeys, rabbits, rats, mice, cows, goats, or sheep.
[0209] The method may include wherein the condition or disease is infectious, wherein the infectious agent is selected from the group consisting of human immunodeficiency virus (HIV), herpes simplex virus type 1 (HSV1), and herpes simplex virus type 2 (HSV2).
[0210] In one aspect, the CRISPR system described herein can be used to treat diseases caused by overexpression of RNA, toxic RNA, and / or mutant RNA (e.g., splicing defects or truncations). For example, expression of toxic RNA can be associated with the formation of nuclear inclusions and delayed degenerative changes in the brain, heart, or skeletal muscle. In some embodiments, the disorder is myotonic dystrophy. In myotonic dystrophy, the primary pathogenic effect of toxic RNA is to sequester binding proteins and impair the regulation of alternative splicing (see, e.g., Osborne et al., "RNA-dominant diseases," Hum. Mol. Genet., 2009 Apr 15;18(8):1471-81). Myotonic dystrophy (dystrophia myotonica (DM)) is of particular interest to geneticists because it produces a wide range of clinical features. The classical form of DM, now called DM type 1 (DM1), is caused by an expansion of a CTG repeat in the 3' untranslated region (UTR) of DMPK, a gene encoding a cytosolic protein kinase. The CRISPR system described herein can target either overexpressed or toxic RNA, for example, the DMPK gene, or misregulated alternative splicing in DM1 skeletal muscle, heart, or brain.
[0211] The CRISPR system described herein can also target trans-acting mutations affecting RNA-dependent functions that cause various diseases, such as Prader-Willi syndrome, spinal muscular atrophy (SMA), dyskeratosis congenita, etc. Lists of diseases that can be treated using the CRISPR system described herein are summarized in Cooper et al., "RNA and disease," Cell, 136.4 (2009):777-793 and WO 2016205764 (each of which is incorporated herein by reference in its entirety).
[0212] The CRISPR system described herein can also be used to treat various tauopathies, including primary and secondary tauopathies such as primary age-related tauopathy (PART) / neurofibrillary tangle (NFT)-dominant senile dementia (with NFTs similar to those seen in Alzheimer's disease (AD), but without plaques), dementia boxing (chronic traumatic encephalopathy), and progressive supranuclear palsy. A useful list of tauopathies and methods for treating these diseases is described, for example, in International Publication No. WO2016205764 (incorporated herein by reference in its entirety).
[0213] The CRISPR system described herein can also be used to target mutations that disrupt cis-acting splicing codes that can cause splicing defects and diseases, such as motor neuron degenerative diseases caused by deletions of the SMN1 gene (e.g., spinal muscular atrophy), Duchenne muscular dystrophy (DMD), frontotemporal dementia and parkinsonism linked to chromosome 17 (FTDP-17), and cystic fibrosis.
[0214] The CRISPR system described herein can further be used for antiviral activity, particularly against RNA viruses. The effector protein can target the viral RNA using an appropriate RNA guide selected to target the viral RNA sequence.
[0215] Furthermore, in vitro RNA sensing assays can be used to detect specific RNA substrates. RNA-targeting effector proteins can be used for RNA-based sensing in living cells. An example application is diagnosis by sensing disease-specific RNA.
[0216] Detailed descriptions of therapeutic applications of the CRISPR systems described herein can be found, for example, in U.S. Pat. No. 8,795,965, EP 3,009,511, WO 2016205764, and WO 2017070605, each of which is herein incorporated by reference in its entirety.
[0217] Applications in plants The CRISPR system described herein has a wide variety of uses in plants.In some embodiments, CRISPR system can be used to engineer the genome of plants (for example, to improve production, to make products with desired post-translational modifications, or to introduce genes for producing industrial products).In some embodiments, CRISPR system can be used to introduce desired traits into plants (for example, with or without heritable modifications to genome), or to regulate the expression of endogenous genes in plant cells or whole plants.
[0218] In some embodiments, the CRISPR system can be used to identify, edit, and / or silence genes encoding specific proteins, such as allergen proteins (e.g., allergen proteins in peanuts, soybeans, lentils, peas, green beans, and mung beans). Detailed descriptions of methods for identifying, editing, and / or silencing protein-encoding genes are described, for example, in Nicolaou et al., "Molecular diagnosis of peanut and legume allergy," Curr. Opin. Allergy Clin. Immunol., 11(3):222-8 (2011), and WO 2016205764 (each of which is incorporated herein by reference in its entirety).
[0219] Delivery of the CRISPR system Through this disclosure and knowledge in the art, the CRISPR systems described herein, or their components, nucleic acid molecules, or nucleic acid molecules encoding or providing the components, can be delivered by various delivery systems, such as vectors, for example, plasmids, or viral delivery vectors. The CRISPR effectors and / or any RNAs (e.g., RNA guides) disclosed herein can be delivered using an appropriate vector, for example, a plasmid, or a viral vector, such as an adeno-associated virus (AAV), lentivirus, adenovirus, and other viral vectors, or a combination thereof. The effectors and one or more RNA guides can be packaged in one or more vectors, for example, a plasmid or viral vector.
[0220] In some embodiments, vectors, such as plasmids or viral vectors, are delivered to the target tissue by, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such delivery may be by single dose or multiple doses. Those skilled in the art will understand that the actual dosage delivered herein may vary widely depending on a variety of factors, including, but not limited to, the choice of vector, the target cell, organism, tissue, the general condition of the subject being treated, the degree of transformation / modification desired, the route and mode of administration, and the type of transformation / modification desired.
[0221] In certain embodiments, delivery is by adenovirus, which delivers at least 1 x 10 5 The dose may be in the form of a single dose containing particles (also referred to as particle units, pu) of adenovirus. In some embodiments, the dose is preferably at least about 1 x 10 6 particles, at least about 1 x 10 7 particles, at least about 1 x 10 8 particles, and at least about 1×10 9The delivery method and dosage are described, for example, in International Publication No. 2016205764 and U.S. Patent No. 8,454,972 (each of which is incorporated herein by reference in its entirety).
[0222] In some embodiments, delivery is via a plasmid. The dosage can be a sufficient number of plasmids to elicit a response. In some cases, a suitable amount of plasmid DNA in a plasmid composition can be about 0.1 to about 2 mg. The plasmid will generally include: (i) a promoter; (ii) a sequence encoding a nucleic acid-targeting CRISPR effector operably linked to the promoter; (iii) a selectable marker; (iv) an origin of replication; and (v) a transcription terminator downstream of and operably linked to (ii). The plasmid can also encode the RNA components of the CRISPR complex, although alternatively, one or more of these may be encoded on a different vector. The frequency of administration is within the purview of a medical or veterinary practitioner (e.g., physician, veterinarian) or skilled in the art.
[0223] In another embodiment, delivery is via liposome or lipofectin formulations, etc., which can be prepared by methods known to those skilled in the art, such as those described in WO2016205764 and U.S. Patent Nos. 5,593,972; 5,589,466; and 5,580,859, each of which is incorporated by reference herein in its entirety.
[0224] In some embodiments, delivery is via nanoparticles or exosomes. For example, exosomes have been shown to be particularly useful for RNA delivery.
[0225] A further means of introducing one or more components of the CRISPR systems described herein into cells is through the use of cell-penetrating peptides (CPPs). In some embodiments, the cell-penetrating peptide is linked to a CRISPR effector. In some embodiments, the CRISPR effector and / or RNA guide are coupled to one or more CPPs and transported into the cell interior (e.g., plant protoplasts). In some embodiments, the CRISPR effector and / or one or more RNA guides are encoded by one or more circular or non-circular DNA molecules that are coupled to one or more CPPs for cellular delivery.
[0226] CPPs are short peptides of less than 35 amino acids derived from either proteins or chimeric sequences that have the ability to transport biomolecules across cell membranes in a receptor-independent manner. CPPs can be cationic peptides, peptides with hydrophobic sequences, amphipathic peptides, peptides with proline-rich antimicrobial sequences, and chimeric or bipartite peptides. Examples of CPPs include Tat (a nuclear transcription activator protein required for HIV type 1 viral replication), penetratin, Kaposi's fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, and sweet arrow peptide. CPPs and methods for their use are described, for example, in Haellbrink et al., "Prediction of cell-penetrating peptides," Methods Mol. Biol., 2015;1324:39-58; Ramakrishna et al., "Gene disruption by cell-penetrating peptide-mediated delivery of Cas9 protein and guide RNA," Genome Res., 2014 Jun;24(6):1020-7; and WO 2016205764 (each of which is incorporated by reference herein in its entirety).
[0227] Various delivery methods for the CRISPR systems described herein are also described in, for example, U.S. Pat. No. 8,795,965, EP 3,009,511, WO 2016205764, and WO 2017070605, each of which is herein incorporated by reference in its entirety. [Example]
[0228] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.
[0229] Example 1 - Identification of Components of the CLUST.091979 CRISPR-Cas System This protein family was identified using the computational methods described above. The CLUST.091979 system contains single effectors associated with CRISPR systems found in uncultured metagenomic sequences collected from environments including, but not limited to, the intestine, bovine intestine, human intestine, sheep intestine, terrestrial, fecal, and mammalian digestive environments (Table 5). Exemplary CLUST.091979 effectors include those shown in Tables 5 and 6 below. As shown in Figures 1A-1L, the effector sequences set forth in SEQ ID NOS: 1-4, 14, 15, 17-19, 21-25, 27-33, 35-49, and 51-56 were aligned to identify regions of sequence similarity. Bar graphs indicate sequence similarity, with the highest bar indicating the residue with the greatest sequence similarity. Non-limiting regions of sequence similarity are shown in Table 7. The regions of sequence similarity indicate that the effectors disclosed herein are a family with a conserved C-terminal RuvC domain representative of nucleases.
[0230] [Table 11]
[0231] [Table 12]
[0232]
Table 13
[0233]
Table 14
[0234]
Table 15
[0235] Table 16
[0236] Table 17
[0237] Table 18
[0238] Table 19
[0239] Table 20
[0240] Table 21
[0241] Table 22
[0242] [Table 23]
[0243] [Table 24]
[0244] [Table 25]
[0245] [Table 26]
[0246] [Table 27]
[0247] [Table 28]
[0248] [Table 29]
[0249] Examples of direct repeat sequences and spacer lengths for these systems are shown in Table 8.
[0250] [Table 30]
[0251] [Table 31]
[0252] [Table 32]
[0253] [Table 33]
[0254] [Table 34]
[0255] [Table 35]
[0256] Example 2 - Identification of trans-activating RNA elements In addition to the effector protein and crRNA, some CRISPR systems described herein may also contain an additional small RNA called a trans-activating RNA (tracrRNA) that activates a robust enzymatic activity. Such tracrRNAs typically contain a complementary region that hybridizes to the crRNA. The crRNA-tracrRNA hybrid forms a complex with the effector, resulting in the activation of a programmable enzymatic activity. The tracrRNA sequence can be identified by searching the genomic sequence flanking the CRISPR array for short sequence motifs homologous to the direct repeat portion of the crRNA. Search methods include perfect or degenerate sequence matching for the complete direct repeat (DR) or DR subsequence. For example, a DR of length n nucleotides can be broken down into a set of overlapping 6-10 nt kmers. These kmers can be aligned with sequences flanking the CRISPR locus, and regions homologous to one or more kmer alignments can be identified as DR homology regions for experimental validation as tracrRNA. Alternatively, the RNA simultaneous folding free energy can be calculated for the complete DR or DR subsequence and short kmer sequences from the genomic sequence flanking the elements of the CRISPR system. Flanking sequence elements with low minimum free energy structures can be identified as DR homology regions for experimental validation as tracrRNA. tracrRNA elements frequently occur in close proximity to CRISPR-associated genes or CRISPR arrays. As an alternative to searching for DR homology regions to identify tracrRNA elements, non-coding sequences flanking CRISPR effectors or CRISPR arrays can be isolated by cloning or gene synthesis for direct experimental validation of tracrRNA. Experimental validation of tracrRNA elements can be performed using small RNA sequencing of the host organism for heterologously expressed CRISPR systems or synthetic sequences in non-native species. Alignment of small RNA sequences from the derived genomic locus can be used to identify expressed RNA products containing DR homology regions and anchoring processing typical of the complete tracrRNA element. Complete tracrRNA candidates identified by RNA sequencing can be validated in vitro or in vivo by expressing the crRNA and effector in combination with or without the tracrRNA candidate and monitoring activation of the effector enzyme activity. In engineered constructs, expression of tracrRNA can be driven by promoters including, but not limited to, the U6, U1, and H1 promoters for expression in mammalian cells or the J23119 promoter for expression in bacteria. In some instances, tracrRNA can be fused to crRNA and expressed as a single RNA guide. The system may include a tracrRNA contained within a non-coding sequence listed in Table 9. For example, in some embodiments, the system includes a tracrRNA set forth in any one of SEQ ID NOs: 152-204.
[0257] [Table 36]
[0258] [Table 37]
[0259] [Table 38]
[0260] [Table 39]
[0261] [Table 40]
[0262] [Table 41]
[0263] [Table 42]
[0264] Example 3 - Identification of novel RNA modulators of enzyme activity In addition to the effector protein and crRNA, some CRISPR systems described herein may also include additional small RNAs, referred to herein as RNA modulators, to activate or regulate effector activity. RNA modulators are predicted to occur in close proximity to CRISPR-associated genes or CRISPR arrays. To identify and validate RNA modulators, non-coding sequences flanking the CRISPR effector or CRISPR array can be isolated by cloning or gene synthesis for direct experimental validation. Experimental validation of RNA modulators can be performed using small RNA sequencing of the host organism for heterologously expressed CRISPR systems or synthetic sequences in non-native species. Alignment of small RNA sequences to the derived genomic locus can be used to identify expressed RNA products containing DR homology regions and their associated processing. Candidate RNA modulators identified by RNA sequencing can be validated in vitro or in vivo by expressing the crRNA and effector in combination with or without the candidate RNA modulator and monitoring changes in effector enzyme activity. In engineered constructs, the RNA modulators can be driven by promoters including, but not limited to, U6, U1, and H1 promoters for expression in mammalian cells, or the J23119 promoter for expression in bacteria. In some instances, the RNA modulator can be artificially fused to either the crRNA, the tracrRNA, or both and expressed as a single RNA element.
[0265] Example 4 - Functional validation of the engineered CLUST.091979 CRISPR-Cas system After identifying the components of the CLUST.091979 CRISPR-Cas system, loci from a metagenomic source designated AUXO013988882 (SEQ ID NO: 1) and a metagenomic source designated SRR3181151 (SEQ ID NO: 4) were selected for functional validation.
[0266] DNA synthesis and effector library cloning To test the activity of an exemplary CLUST.091979 CRISPR-Cas system, a system was designed and synthesized using the pET28a(+) vector. Briefly, the E. coli codon-optimized nucleic acid sequence encoding the CLUST.091979 AUXO013988882 effector (SEQ ID NO: 1 shown in Table 6) and the E. coli codon-optimized nucleic acid sequence encoding the CLUST.091979 SRR3181151 effector (SEQ ID NO: 4 shown in Table 6) were synthesized (Genscript) and individually cloned into a custom expression system derived from pET-28a(+) (EMD-Millipore). The vector contained nucleic acids encoding the CLUST.091979 effectors under the control of a lac promoter and an E. coli ribosome binding sequence. The vector also contained an acceptor site for a CRISPR array library driven by the J23119 promoter following the open reading frame of the CLUST.091979 effector. As shown in Table 9, the non-coding sequence used for the CLUST.091979 AUXO013988882 effector (SEQ ID NO: 1) is set forth in SEQ ID NO: 98, and the non-coding sequence used for the CLUST.091979 SRR3181151 effector (SEQ ID NO: 4) is set forth in SEQ ID NO: 99. An additional condition was tested in which the CLUST.091979 effector was cloned individually into pET28a(+) without the non-coding sequence. See Figure 4A.
[0267] Oligonucleotide library synthesis (OLS) pools containing "repeat-spacer-repeat" sequences were computationally designed, where the "repeat" corresponds to the consensus direct repeat sequence found in the CRISPR array associated with the effector, and the "spacer" corresponds to the sequence tiling the pACYC184 plasmid or the essential E. coli gene. Specifically, as shown in Table 8, the repeat sequence used for the CLUST.091979 AUXO013988882 effector (SEQ ID NO: 1) is set forth in SEQ ID NO: 57, and the repeat sequence used for the 091979 SRR3181151 effector (SEQ ID NO: 4) is set forth in SEQ ID NO: 60. The spacer length was determined by the most common spacer length found in the endogenous CRISPR array. The repeat-spacer-repeat sequences were appended with restriction sites that allow for bidirectional cloning of fragments into the CRISPR array library acceptor site described above and a unique PCR priming site that allows for specific amplification of a specific repeat-spacer-repeat library from a larger pool.
[0268] The repeat-spacer-repeat library was then cloned into a plasmid using the Golden Gate assembly method. Briefly, we first amplified each repeat-spacer-repeat from the OLS pool (Agilent Genomics) using unique PCR primers, and pre-linearized the plasmid backbone with BsaI to reduce potential background. Both DNA fragments were purified with Ampure XP (Beckman Coulter) before being added to the Golden Gate Assembly Master Mix (New England Biolabs) and incubated according to the manufacturer's instructions. The Golden Gate reaction was further purified and concentrated to achieve maximum transformation efficiency in the subsequent step of bacterial screening.
[0269] Plasmid libraries containing different repeat-spacer-repeat elements and CRISPR effectors were electroporated into E. cloni electrocompetent E. coli (Lucigen) using a Gene Pulser Xcell® (Bio-Rad) according to the protocol recommended by Lucigen. Libraries were either cotransformed with purified pACYC184 plasmid or directly transformed into E. cloni electrocompetent E. coli (Lucigen) containing pACYC184. The resulting plates were plated onto agar plates containing chloramphenicol (Fisher), tetracycline (Alfa Aesar), and kanamycin (Alfa Aesar) in BioAssay® dishes (Thermo Fisher) and incubated at 37°C for 10–12 hours. After estimating approximate colony counts to ensure sufficient library representation on bacterial plates, bacteria were harvested and plasmid DNA was extracted using a QIAprep Spin Miniprep® kit (Qiagen) to generate an "output library." Barcoded next-generation sequencing libraries were generated from both the pre-transformation "input library" and the post-recovery "output library" by PCR using custom primers containing barcodes and sites compatible with Illumina sequencing chemistry. These libraries were pooled and loaded onto a Nextseq550 (Illumina) to evaluate effectors. To ensure consistency, at least two independent biological replicates were performed for each screen. See Figure 4B.
[0270] Bacterial Screen Sequencing Analysis Next-generation sequencing data for the screen input and output libraries were demultiplexed using Illumina bcl2fastq. The reads in the resulting fastq file for each sample contained CRISPR array elements for the screening plasmid library. The array orientation was determined using the direct repeat sequence of the CRISPR array, and the corresponding targets were determined by mapping the spacer sequence to the source (pACYC184 or E. Cloni) or negative control sequence (GFP). For each sample, the total number of reads (r) for each unique array element in a given plasmid library was calculated. a ) were counted and normalized as follows: (r a +1) / total number of reads for all library array elements. The depletion score was calculated by dividing the normalized output reads for a given array element by the normalized input reads.
[0271] To identify specific parameters driving enzymatic activity and bacterial cell death, we used next-generation sequencing (NGS) to quantify and compare the representation of individual CRISPR arrays (i.e., repeat-spacer-repeat) in the PCR products of input and output plasmid libraries. The array depletion rate was defined as the normalized number of output reads divided by the normalized number of input reads. If the depletion rate was less than 0.3 (greater than three-fold depletion), the array was considered "strongly depleted," as indicated by the dashed line in Figures 5 and 8. When calculating array depletion rates across biological replicates, we took the maximum depletion rate value for a given CRISPR array across all experiments (i.e., a strongly depleted array must be strongly depleted in all biological replicates). For each spacer target, a matrix was created containing the array depletion rate and the following features: target strand, transcript targeting, ORI targeting, target sequence motif, flanking sequence motif, and target secondary structure. The extent to which different features in this matrix explain target depletion for the CLUST.091979 system was investigated.
[0272] Figures 5 and 8 show the degree of interference activity of CLUST.091979 compositions engineered with non-coding sequences by plotting the normalized ratio of sequencing reads in the screen output versus the screen input for a given target. Results are plotted for each DR transcription direction. In functional screening of the compositions, active effectors complexed with active RNA guides interfere with the ability of pACYC184 to confer resistance to chloramphenicol and tetracycline in Escherichia coli (E. coli), resulting in cell death and depletion of spacer elements within the pool. Comparing the results of deep sequencing of the initial DNA library (screen input) and surviving transformed E. coli (screen output) suggests specific target sequences and DR transcription directions that enable an active, programmable CRISPR system. The screen also shows that the effector complex is active in only one DR orientation. Thus, the screen demonstrated that the CLUST.091979 AUXO013988882 effector was active in the “forward” orientation of the DR (5′-ACTA…AACT-[spacer]-3′) (Figure 5 ), and that the CLUST.091979 We showed that the SRR3181151 effector was active in the “reverse” orientation of the DR (5′-CCTG…CAAC-[spacer]-3′) (Fig. 8 ).
[0273] Figures 6A and 6B show the locations of strongly depleted targets for pACYC184 and the CLUST.091979 AUXO013988882 effector (+non-coding sequence), which target essential E. coli (E. cloni) genes, respectively. Similarly, Figures 9A and 9B show the locations of strongly depleted targets for pACYC184 and the CLUST.091979 SRR3181151 effector, which target essential E. coli (E. cloni) genes, respectively. The flanking sequences of the depleted targets were analyzed to determine the PAM sequences of CLUST.091979 AUXO013988882 and CLUST.091979 SRR3181151. WebLogo representations (Crooks et al., Genome Research 14:1188-90, 2004) of the PAM sequences of CLUST.091979 AUXO013988882 and CLUST.091979 SRR3181151 are shown in Figures 7 and 10, respectively, with position "20" corresponding to the nucleotide adjacent to the 5' end of the target.
[0274] Thus, multiple effectors of CLUST.091979 CRISPR-Cas are active in vivo.
[0275] Example 5 - Mammalian Gene Targeting with CLUST.091979 This example describes indel assessment against multiple targets using nucleases from CLUST.091979 introduced into mammalian cells by transient transfection.
[0276] The effectors of SEQ ID NO:4, SEQ ID NO:8, and SEQ ID NO:10 were cloned into the pcda3.1 backbone (Invitrogen). The plasmids were then maxiprepped and diluted to 1 μg / μL. For RNA guide preparation, a dsDNA fragment encoding the crRNA was introduced with Ultramer containing the target sequence scaffold and U6 promoter. Ultramer was resuspended in 10 mM Tris·HCl, pH 7.5, to a final stock concentration of 100 μM. The working stock was then diluted to 10 μM and used as a template in a PCR reaction, again using 10 mM Tris·HCl. Amplification of the crRNA was performed in a 50 μL reaction using the following components: 0.02 μL of the aforementioned template, 2.5 μL of forward primer, 2.5 μL of reverse primer, 25 μL of NEB HiFi polymerase, and 20 μL of water. Cycling conditions were 1x (98°C for 30 seconds), 30x (98°C for 10 seconds, 67°C for 15 seconds), and 1x (72°C for 2 minutes). PCR products were cleaned up with 1.8x SPRI treatment and normalized to 25 ng / µL. The prepared crRNA sequences and their corresponding target sequences are shown in Table 10. The direct repeat sequences of mature crRNAs SEQ ID NO:205, SEQ ID NO:207, SEQ ID NO:252, SEQ ID NO:254, SEQ ID NO:256, SEQ ID NO:258, SEQ ID NO:260, SEQ ID NO:262, SEQ ID NO:264, SEQ ID NO:266, SEQ ID NO:268, SEQ ID NO:270, SEQ ID NO:272, SEQ ID NO:274, and SEQ ID NO:276 are set forth in SEQ ID NO:60. The direct repeats of mature crRNAs SEQ ID NO:209 and SEQ ID NO:214 are set forth in SEQ ID NO:62. The direct repeats of mature crRNAs of SEQ ID NO:211, SEQ ID NO:278, SEQ ID NO:280, SEQ ID NO:282, SEQ ID NO:284, SEQ ID NO:286, and SEQ ID NO:288 are set forth in SEQ ID NO:213.
[0277] [Table 43]
[0278] [Table 44]
[0279] Approximately 16 hours before transfection, 100 μl of 25,000 HEK293T cells in DMEM / 10% FBS + Pen / Strep were plated into each well of a 96-well plate. On the day of transfection, cells were 70–90% confluent. For each well to be transfected, a mixture of 0.5 μl of Lipofectamine 2000 and 9.5 μl of Opti-MEM was prepared and then incubated at room temperature for 5–20 minutes (Solution 1). After incubation, the Lipofectamine:OptiMEM mixture was added to another mixture containing 182 ng of effector plasmid, 14 ng of crRNA, and up to 10 μL of water (Solution 2). For the negative control, crRNA was not included in Solution 2. The mixture of Solution 1 and Solution 2 was mixed by pipetting up and down and then incubated at room temperature for 25 minutes. After incubation, 20 μL of the mixture of Solution 1 and Solution 2 was added dropwise to each well of a 96-well plate containing cells. 72 hours after transfection, cells were trypsinized by adding 10 μL of TrypLE to the center of each well and incubating for approximately 5 minutes. Next, 100 μL of D10 medium was added to each well and mixed to resuspend the cells. The cells were then spun down at 500 g for 10 minutes, and the supernatant was discarded. QuickExtract buffer was added to 1 / 5 of the original cell suspension volume. The cells were then incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.
[0280] Samples for next-generation sequencing were prepared by two rounds of PCR. The first round (PCR1) was used to amplify specific genomic regions depending on the target. The PCR1 product was purified by column purification. PCR round 2 (PCR2) was performed to add Illumina adapters and indexes. The reactions were then pooled and purified by column purification. Sequencing runs were performed using the NextSeq v2.5 medium or high throughput kit for 150 cycles.
[0281] Figures 11A, 11B, 11C, and 11D show the percent indels at the AAVS1, VEGFA, and EMX1 target loci in HEK293T cells after transfection with the effector of SEQ ID NO: 4 or SEQ ID NO: 10, respectively. Bars reflect the average percent indels measured in two biological replicates. For the effectors of SEQ ID NO: 4 and SEQ ID NO: 10, the percent indels were higher than the percent indels of the negative control at each of the targets.
[0282] As shown in Figure 11A, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 205 was active on the AAVS1 target of SEQ ID NO: 206, and the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 207 was active on the VEGFA target of SEQ ID NO: 208. As shown in Figure 11B, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 252 was active on the AAVS1 target of SEQ ID NO: 253, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 254 was active on the AAVS1 target of SEQ ID NO: 255, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 256 was active on the AAVS1 target of SEQ ID NO: 257, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 258 was active on the AAVS1 target of SEQ ID NO: 259, and the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 274 was active on the AAVS1 target of SEQ ID NO: 275. As shown in Figure 11B, the complex formed by the effector of SEQ ID NO:4 and the crRNA of SEQ ID NO:260 was active on the EMX1 target of SEQ ID NO:261.Similarly, as shown in Figure 11B, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 262 was active on the VEGFA1 target of SEQ ID NO: 263, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 264 was active on the VEGFA1 target of SEQ ID NO: 265, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 266 was active on the VEGFA1 target of SEQ ID NO: 267, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 268 was active on the VEGFA1 target of SEQ ID NO: 269, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 270 was active on the VEGFA1 target of SEQ ID NO: 271, the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 272 was active on the VEGFA1 target of SEQ ID NO: 273, and the complex formed by the effector of SEQ ID NO: 4 and crRNA of SEQ ID NO: 274 was active on the VEGFA1 target of SEQ ID NO: 275. The effector of SEQ ID NO:4 utilized 5'-TTTG-3'PAM for each of the targets in Figures 11A and 11B.
[0283] As shown in Figure 11C, the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 209 was active on the AAVS1 target of SEQ ID NO: 210, the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 211 was active on the AAVS1 target of SEQ ID NO: 212, and the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 214 was active on the VEGFA target of SEQ ID NO: 215. As shown in Figure 11D, the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 278 was active on the AAVS1 target of SEQ ID NO: 279, the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 280 was active on the AAVS1 target of SEQ ID NO: 281, the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 284 was active on the AAVS1 target of SEQ ID NO: 285, and the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 286 was active on the AAVS1 target of SEQ ID NO: 287. Similarly, as shown in Figure 11D, the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 288 was active on the EMX1 target of SEQ ID NO: 289, and the complex formed by the effector of SEQ ID NO: 10 and crRNA of SEQ ID NO: 282 was active on the VEGFA target of SEQ ID NO: 283. The effector of SEQ ID NO: 10 utilized 5'-ATTG-3'PAM and 5'-GTTA-3'PAM for the targets in Figures 11C and 11D.
[0284] This example suggests that nucleases of the CLUST.091979 family are active in mammalian cells.
[0285] Other embodiments While the present invention has been described with a detailed description thereof, it should be understood that the foregoing description is illustrative and is not intended to limit the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims. In certain embodiments, for example, the following items are provided: (Item 1) CLUST.091979, an engineered non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-Cas system, (a) a CRISPR-associated protein or a nucleic acid encoding the CRISPR-associated protein, wherein the CRISPR-associated protein comprises the amino acid sequence of SEQ ID NO: 241; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid; Including, a CRISPR-Cas system, wherein the CRISPR-associated protein is capable of binding to the RNA guide and modifying the target nucleic acid sequence complementary to the spacer sequence. (Item 2) 2. The system of claim 1, wherein the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO:4, SEQ ID NO:10, SEQ ID NO:12, or SEQ ID NO:14. (Item 3) CLUST.091979, an engineered non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-Cas system, (a) a CRISPR-associated protein or a nucleic acid encoding the CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 56; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid; Including, a CRISPR-Cas system, wherein the CRISPR-associated protein is capable of binding to the RNA guide and modifying the target nucleic acid sequence complementary to the spacer sequence. (Item 4) 4. The system of claim 3, wherein the CRISPR-associated protein comprises at least one RuvC domain or at least one split RuvC domain. (Item 5) The CRISPR-associated protein has one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216) (wherein Xi is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M); (b) RX1X2X3L (SEQ ID NO: 217) (wherein Xi is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (c) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein Xi is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (e)LX1NX 2( SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or S); (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein Xi is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M, T, or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein Xi is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q); and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L) 5. The system according to item 3 or 4, comprising: (Item 6) 6. The system according to any one of Items 3 to 5, wherein the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, 118 to 151, or 213. (Item 7) 7. The system according to Item 6, wherein the direct repeat sequence comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to a nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, 118 to 151, or 213. (Item 8) The direct repeat sequence may be one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (wherein X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) The system according to any one of items 3 to 7, comprising: (Item 9) 9. The system of any one of Items 3 to 8, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. (Item 10) 10. The system of claim 9, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:1, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:57. (Item 11) 9. The system of any one of Items 3 to 8, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1; and the CRISPR-associated protein has the ability to recognize a protospacer adjacent motif (PAM) sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 12) 12. The system of claim 11, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:1, and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 13) 9. The system of any one of Items 3 to 8, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 14) 14. The system of claim 13, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:60. (Item 15) 9. The system of any one of Items 3 to 8, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4; and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 16) 16. The system of claim 15, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99%, or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the CRISPR-associated protein is capable of recognizing a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G. (Item 17) 9. The system of any one of Items 3 to 8, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 18) 18. The system of Item 17, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 19) 9. The system of any one of Items 3 to 8, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), wherein "N" is any nucleotide, and "R" is A or G. (Item 20) 20. The system of claim 19, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), wherein "N" is any nucleotide, and "R" is A or G. (Item 21) 21. The system of any one of items 1 to 20, wherein the spacer sequence of the RNA guide comprises from about 15 nucleotides to about 55 nucleotides. (Item 22) 22. The system of item 21, wherein the spacer sequence of the RNA guide comprises 20 to 45 nucleotides. (Item 23) 23. The system of any one of paragraphs 1 to 22, wherein the CRISPR-associated protein comprises a catalytic residue (e.g., aspartic acid or glutamic acid). (Item 24) 24. The system of any one of items 1 to 23, wherein the CRISPR-associated protein cleaves the target nucleic acid. (Item 25) 25. The system of any one of items 1 to 24, wherein the CRISPR-associated protein further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcription modifier, a photogating factor, a chemical inducible factor, or a chromatin visualization factor. (Item 26) 26. The system of any one of items 1 to 25, wherein the nucleic acid encoding the CRISPR-associated protein is codon-optimized for expression in a cell. (Item 27) 27. The system of any one of items 1 to 26, wherein the nucleic acid encoding the CRISPR-associated protein is operably linked to a promoter. (Item 28) 28. The system of any one of items 1 to 27, wherein the nucleic acid encoding the CRISPR-associated protein is in a vector. (Item 29) 29. The system of item 28, wherein the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector. (Item 30) 30. The system according to any one of items 1 to 29, wherein the target nucleic acid is a DNA molecule. (Item 31) 31. The system of any one of items 1 to 30, wherein the CRISPR-associated protein comprises non-specific nuclease activity. (Item 32) 32. The system of any one of items 1 to 31, wherein recognition of the target nucleic acid by the CRISPR-associated protein and RNA guide results in modification of the target nucleic acid. (Item 33) 34. The system of claim 32, wherein the modification of the target nucleic acid is a double-strand break event. Item 35. The system of item 32, wherein the modification of the target nucleic acid is a single-strand break event. 36. The system of claim 32, wherein the modification of the target nucleic acid results in an insertion event. 37. The system of claim 32, wherein the modification of the target nucleic acid results in a deletion event. 37. The system of any one of items 32 to 36, wherein the modification of the target nucleic acid results in cytotoxicity or cell death. (Item 38) 31. The system of any one of items 1 to 30, further comprising a donor template nucleic acid. (Item 39) 39. The system of claim 38, wherein the donor template nucleic acid is a DNA molecule. (Item 40) 39. The system of claim 38, wherein the donor template nucleic acid is an RNA molecule. (Item 41) 41. The system of any one of items 1 to 40, wherein the RNA guide optionally comprises tracrRNA. (Item 42) 41. The system of any one of items 1 to 40, wherein the system does not comprise tracrRNA. (Item 43) 43. The system of any one of items 1 to 42, wherein the CRISPR-associated protein is self-processing. (Item 44) 44. The system according to any one of items 1 to 43, wherein the system is present in a delivery composition comprising a nanoparticle, a liposome, an exosome, a microvesicle, or a gene gun. (Item 45) 44. The system according to any one of items 1 to 43, which is intracellular. (Item 46) Item 46. The system of item 45, wherein the cell is a eukaryotic cell. (Item 47) Item 46. The system of item 45, wherein the cells are prokaryotic cells. (Item 48) (a) a CRISPR-associated protein or a nucleic acid encoding the CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 56; and (b) A cell comprising an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid. (Item 49) The CRISPR-associated protein has one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216) (wherein Xi is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M); (b) RX1X2X3L (SEQ ID NO: 217) (wherein Xi is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (c) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein Xi is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (e) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or S); (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein Xi is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M, T, or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein Xi is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q); and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L) Item 49. The cell according to Item 48, comprising: (Item 50) 50. The cell according to item 48 or 49, wherein the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, 118 to 151, or 213. (Item 51) 51. The cell according to Item 50, wherein the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to a nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, 118 to 151, or 213. (Item 52) The direct repeat sequence may be one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (wherein X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) 52. The cell according to any one of Items 48 to 51, comprising: (Item 53) 53. The cell of any one of Items 48 to 52, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. (Item 54) 54. The cell of claim 53, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. (Item 55) 53. The cell of any one of Items 48 to 52, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1; and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 56) 56. The cell of claim 55, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:1, and the CRISPR-associated protein is capable of recognizing a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', wherein "N" is any nucleotide and "R" is A or G. (Item 57) 53. The cell of any one of Items 48 to 52, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 58) 58. The cell of claim 57, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:60. (Item 59) 53. The cell of any one of Items 48 to 52, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4; and wherein the CRISPR-associated protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', wherein "N" is any nucleotide and "R" is A or G. (Item 60) 60. The cell of claim 59, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the CRISPR-associated protein is capable of recognizing a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', wherein "N" is any nucleotide and "R" is A or G. (Item 61) 53. The cell of any one of Items 48 to 52, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 62) 62. The cell of claim 61, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 63) 53. The cell of any one of Items 48 to 52, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), wherein "N" is any nucleotide, and "R" is A or G. (Item 64) 64. The cell of item 63, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the CRISPR-associated protein is capable of recognizing a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), wherein "N" is any nucleotide and "R" is A or G. (Item 65) 65. The cell according to any one of items 48 to 64, wherein the spacer sequence comprises from about 15 nucleotides to about 55 nucleotides. (Item 66) 66. The cell according to item 65, wherein the spacer sequence comprises 20 to 45 nucleotides. (Item 67) 67. The cell of any one of items 48 to 66, wherein the cell further comprises tracrRNA. (Item 68) 67. The cell of any one of items 48 to 66, wherein the system does not comprise tracrRNA. (Item 69) 69. The cell of any one of items 48 to 68, wherein the cell is a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell. (Item 70) 70. The cell according to any one of items 48 to 69, wherein the cell is a prokaryotic cell. (Item 71) A method for binding the system according to any one of Items 1 to 47 to a target nucleic acid in a cell, comprising: (a) providing said system; and (b) delivering said system to said cell. wherein the cell comprises the target nucleic acid, the CRISPR-associated protein binds to the RNA guide, and the spacer sequence binds to the target nucleic acid. (Item 72) 72. The method of item 71, wherein the cell is a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell. (Item 73) 1. A method for modifying a target nucleic acid, comprising: (a) a CRISPR-associated protein or a nucleic acid encoding the CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 56; and (b) an RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid; Including, the CRISPR-associated protein is capable of binding to an RNA guide; Recognition of the target nucleic acid by the CRISPR-associated protein and RNA guide results in modification of the target nucleic acid. A method comprising delivering an engineered non-naturally occurring CRISPR-Cas system to a target nucleic acid. (Item 74) The CRISPR-associated protein has one or more of the following sequences: (a) PX1X2X3X4F (SEQ ID NO: 216) (wherein Xi is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M); (b) RX1X2X3L (SEQ ID NO: 217) (wherein Xi is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (c) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (d) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein Xi is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (e) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or S); (f) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein Xi is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M, T, or I); (g) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein Xi is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q); and (h) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L) Item 74. The method according to Item 73, comprising: (Item 75) 75. The method of claim 73 or 74, wherein the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, 118 to 151, or 213. (Item 76) 76. The method of Item 75, wherein the direct repeat sequence comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 57 to 90, 118 to 151, or 213. (Item 77) The direct repeat sequence may be one or more of the following sequences: (a) X1X2TX3X4X5X6X7X8 (SEQ ID NO: 224) (wherein X1 is A or C or G, X2 is T or C or A, X3 is T or G or A, X4 is T or G, X5 is T or G or A, X6 is G or T or A, X7 is T or G or A, and X8 is A or G or T); (b) X1X2X3X4X5X6X7X8X9 (SEQ ID NO: 226) (wherein X1 is T or C or A, X2 is T or A or G, X3 is T or C or A, X4 is T or A, X5 is T or A or G, X6 is T or A, X7 is A or T, X8 is A or G or C or T, and X9 is G or A or C); and (c) X1X2X3AC (SEQ ID NO: 228) (wherein X1 is A or C or G, X2 is C or A, and X3 is A or C) 77. The method according to any one of items 73 to 76, comprising: (Item 78) 78. The method of any one of Aspects 73 to 77, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 57. (Item 79) 79. The method of claim 78, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:1, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:57. (Item 80) 78. The method of any one of Aspects 73 to 77, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 1; and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 81) 81. The method of claim 80, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:1, and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-TNNT-3' or 5'-TNRT-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 82) 78. The method of any one of Aspects 73 to 77, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 83) 83. The method of claim 82, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO:60. (Item 84) 78. The method of any one of Aspects 73 to 77, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 4, wherein the CRISPR-associated protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', wherein "N" is any nucleotide, and "R" is A or G. (Item 85) 85. The method of claim 84, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO:4, and the CRISPR-associated protein has the ability to recognize a PAM sequence, wherein the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', wherein "N" is any nucleotide and "R" is A or G. (Item 86) 78. The method of any one of Aspects 73 to 77, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 87) 87. The method of claim 86, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, and the direct repeat sequence comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 88) 78. The method of any one of Aspects 73 to 77, wherein the CRISPR-associated protein is a protein having at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, wherein the CRISPR-associated protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), wherein "N" is any nucleotide, and "R" is A or G. (Item 89) 89. The method of claim 88, wherein the CRISPR-associated protein is a protein having at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identity to the amino acid sequence set forth in SEQ ID NO: 10, the CRISPR-associated protein has the ability to recognize a PAM sequence, and the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), wherein "N" is any nucleotide and "R" is A or G. (Item 90) 90. The method according to any one of items 73 to 89, wherein the spacer sequence comprises from about 15 nucleotides to about 55 nucleotides. (Item 91) Item 91. The method of item 90, wherein the spacer sequence comprises 20 to 45 nucleotides. (Item 92) 92. The method of any one of items 73 to 91, wherein the system further comprises tracrRNA. (Item 93) 92. The method of any one of items 73 to 91, wherein the system does not comprise tracrRNA. (Item 94) 94. The method of any one of items 73 to 93, wherein the target nucleic acid is a DNA molecule. (Item 95) 95. The method of any one of items 73 to 94, wherein the CRISPR-associated protein comprises non-specific nuclease activity. (Item 96) 96. The method of any one of items 73 to 95, wherein the modification of the target nucleic acid is a double-strand break event. (Item 97) 97. The method of any one of items 73 to 96, wherein the modification of the target nucleic acid is a single-strand break event. (Item 98) 98. The method of any one of items 73 to 97, wherein the modification of the target nucleic acid results in an insertion event. (Item 99) 99. The method of any one of items 73 to 98, wherein the modification of the target nucleic acid results in a deletion event. (Item 100) 99. The method according to any one of items 73 to 99, wherein the modification of the target nucleic acid results in cytotoxicity or cell death. (Item 101) 48. A method for editing a target nucleic acid, comprising contacting the system according to any one of items 1 to 47 with the target nucleic acid. (Item 102) 48. A method for modifying expression of a target nucleic acid, comprising contacting the system according to any one of items 1 to 47 with the target nucleic acid. (Item 103) Item 104. A method for targeting insertion of a payload nucleic acid at a site in a target nucleic acid, the method comprising contacting the target nucleic acid with the system of any one of items 1 to 47. Item 105. A method for targeting excision of a payload nucleic acid from a site in a target nucleic acid, the method comprising contacting the target nucleic acid with the system described in any one of Items 1 to 47. 48. A method for non-specifically degrading single-stranded DNA upon recognition of a DNA target nucleic acid, the method comprising contacting said target nucleic acid with the system according to any one of items 1 to 47. (Item 106) 1. A method for detecting a target nucleic acid in a sample, comprising: (a) contacting the sample with the system according to any one of items 1 to 47 and a labeled reporter nucleic acid, wherein cleavage of the labeled reporter nucleic acid occurs when a spacer sequence hybridizes to the target nucleic acid; and (b) measuring a detectable signal produced by cleavage of the labeled reporter nucleic acid, thereby detecting the presence of the target nucleic acid in the sample. (Item 107) (a) Methods for targeting and editing target nucleic acids; (b) a method for non-specific degradation of single-stranded nucleic acids in response to nucleic acid recognition; (c) a method for targeting and nicking a non-spacer complementary strand of a double-stranded target in response to recognition of a spacer complementary strand of the double-stranded target; (d) methods for targeting and cleaving double-stranded target nucleic acids; (e) a method for detecting a target nucleic acid in a sample; (f) a method for specific editing of double-stranded nucleic acids; (g) A method for base editing of double-stranded nucleic acids; (h) a method for inducing genotype-specific or transcriptional state-specific cell death or dormancy in a cell; (i) A method for creating indels in double-stranded nucleic acid targets; (j) a method for inserting a sequence into a double-stranded nucleic acid target; or (k) Method for forming a sequence deletion or inversion in a double-stranded nucleic acid target 48. Use of the system according to any one of items 1 to 47 in an in vitro or ex vivo method, (Item 108) 1. A method for introducing an insertion or deletion into a target nucleic acid in a mammalian cell, comprising: (a) a nucleic acid sequence encoding a CRISPR-associated protein, wherein the CRISPR-associated protein comprises an amino acid sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to an amino acid sequence set forth in any one of SEQ ID NOs: 1-56; and (b) an RNA guide (or a nucleic acid encoding an RNA guide) comprising a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid; comprising transfection of the CRISPR-associated protein is capable of binding to the RNA guide; wherein recognition of the target nucleic acid by the CRISPR-associated protein and RNA guide results in modification of the target nucleic acid. (Item 109) 109. The method of claim 108, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO:4. (Item 110) 110. The method of claim 109, wherein the CRISPR-associated protein comprises an amino acid sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO:4. (Item 111) 109. The method of claim 108, wherein the direct repeat comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 112) 112. The method of claim 111, wherein the direct repeat comprises a nucleotide sequence that is at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 60. (Item 113) 109. The method of claim 108, wherein the target nucleic acid is flanked by a PAM sequence, and the PAM sequence comprises a nucleic acid sequence described as 5'-NTTN-3', 5'-NTTR-3' (e.g., 5'-TTTG-3'), or 5'-NNR-3', where "N" is any nucleotide and "R" is A or G. (Item 114) 109. The method of claim 108, wherein the CRISPR-associated protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. (Item 115) 115. The method of claim 114, wherein the CRISPR-associated protein comprises an amino acid sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence set forth in SEQ ID NO: 10. (Item 116) 109. The method of claim 108, wherein the direct repeat comprises a nucleotide sequence at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 117) 117. The method of claim 116, wherein the direct repeat comprises a nucleotide sequence at least 95% (e.g., 95%, 96%, 97%, 98%, 99% or 100%) identical to the nucleotide sequence set forth in SEQ ID NO: 62 or SEQ ID NO: 213. (Item 118) 109. The method of claim 108, wherein the target nucleic acid is flanked by PAM sequences, and the PAM sequence comprises a nucleic acid sequence set forth as 5'-NTTN-3' or 5'-RTTR-3' (e.g., 5'-ATTG-3' or 5'-GTTA-3'), where "N" is any nucleotide and "R" is A or G. (Item 119) 119. The method of any one of items 108 to 118, wherein the transfection is a transient transfection. (Item 120) 119. The method of any one of items 108 to 119, wherein the cells are human cells. (Item 121) (a) a CRISPR-associated protein or a nucleic acid encoding the CRISPR-associated protein, and (b) an RNA guide containing a direct repeat sequence and a spacer sequence; A composition comprising: The CRISPR-associated protein has one or more of the following amino acid sequences: (i) PX1X2X3X4F (SEQ ID NO: 216) (wherein Xi is L or M or I or C or F, X2 is Y or W or F, X3 is K or T or C or R or W or Y or H or V, and X4 is I or L or M); (ii) RX1X2X3L (SEQ ID NO: 217) (wherein Xi is I or L or M or Y or T or F, X2 is R or Q or K or E or S or T, and X3 is L or I or T or C or M or K); (iii) NX1YX2 (SEQ ID NO: 218) (wherein X1 is I or L or F and X2 is K or R or V or E); (iv) KX1X2X3FAX4X5KD (SEQ ID NO: 219) (wherein Xi is T or I or N or A or S or F or V, X2 is I or V or L or S, X3 is H or S or G or R, X4 is D or S or E, and X5 is I or V or M or T or N); (v) LX1NX2 (SEQ ID NO: 220) (wherein X1 is G or S or C or T and X2 is N or Y or K or S); (vi) PX1X2X3X4SQX5DS (SEQ ID NO: 221) (wherein Xi is S or P or A, X2 is Y or S or A or P or E or Y or Q or N, X3 is F or Y or H, X4 is T or S, and X5 is M, T, or I); (vii) KX1X2VRX3X4QEX5H (SEQ ID NO: 222) (wherein Xi is N or K or W or R or E or T or Y, X2 is M or R or L or S or K or V or E or T or I or D, X3 is L or R or H or P or T or K or P Q or S or A, X4 is G or Q or N or R or K or E or I or T or S or C, and X5 is R or W or Y or K or T or F or S or Q); and (viii) X1NGX2X3X4DX5NX6X7X8N (SEQ ID NO: 223) (wherein Xi is I or K or V or L, X2 is L or M, X3 is N or H or P, X4 is A or S or C, X5 is V or Y or I or F or T or N, X6 is A or S, X7 is S or A or P, and X8 is M or C or L or R or N or S or K or L) Including, The composition, wherein the CRISPR-associated protein binds to the RNA guide and the spacer binds to a target nucleic acid.
Claims
[Claim 1] The invention described in the present specification.