Type II CAS proteins and their applications

Shorter type II Cas proteins from unclassified bacteria address the packaging limitations of AAVs, enabling efficient genome editing with novel PAM specificities and broader targetability, suitable for treating genetic diseases.

JP2026502531APending Publication Date: 2026-01-23アリア セラピューティクス エスアールエル
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025540745
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-01-10
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

The limited packaging capacity of adeno-associated viral vectors (AAVs) poses a challenge for delivering large Cas proteins, such as SpCas9, along with guide RNAs, necessitating the development of smaller type II Cas nucleases that can be packaged together with gRNAs for more flexible and efficient genome editing.

Method used

The discovery and utilization of type II Cas proteins from unclassified bacteria, such as AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET, which are significantly shorter than SpCas9, allowing for efficient packaging and genome editing with novel PAM specificities.

Benefits of technology

These shorter type II Cas proteins enable efficient genome editing by expanding the range of targetable sites and improving the flexibility of CRISPR-Cas genome editing systems, facilitating their use in treating genetic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502531000119
    Figure 2026502531000119
  • Figure 2026502531000120
    Figure 2026502531000120
  • Figure 2026502531000121
    Figure 2026502531000121
Patent Text Reader

Abstract

Type II Cas proteins, e.g., type II Cas proteins designated as AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins; gRNAs for type II Cas proteins; systems comprising type II Cas proteins and gRNAs; nucleic acids encoding type II Cas proteins, gRNAs, and systems; particles comprising the foregoing; pharmaceutical compositions of the foregoing; and uses of the foregoing, for example, to modify genomic DNA of a cell.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1. CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 479,461, filed January 11, 2023, U.S. Provisional Patent Application No. 63 / 488,260, filed March 3, 2023, and U.S. Provisional Patent Application No. 63 / 601,409, filed November 21, 2023, the contents of which are incorporated herein by reference in their entireties.

[0002] 2. Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, which is incorporated herein by reference in its entirety. The XML Sequence Listing, created on January 9, 2024, is named ALA-011WO_SL.xml and is 944,048 bytes in size. [Background technology]

[0003] 3.Background CRISPR-Cas genome editing using type II Cas proteins and associated guide RNAs (gRNAs) is a powerful tool with the potential to treat various genetic diseases. Adeno-associated viral vectors (AAVs) are commonly used to deliver Cas proteins, such as Streptococcus pyogenes Cas9 (SpCas9), and their associated guide RNAs (gRNAs). However, packaging large Cas proteins, such as SpCas9, together with guide RNAs into a single AAV vector can be challenging due to the limited packaging capacity of AAVs. Therefore, there is a need for smaller type II Cas nucleases that can be packaged together with gRNAs into a single AAV. In addition, the discovery of novel nucleases with novel PAM specificities could broaden the range of targetable sites in the cellular genome, making genome editing more flexible and efficient. Summary of the Invention

[0004] 4. Overview The present disclosure relates, in part, to a type II Cas protein from an unclassified bacterium in the genus Acidaminococcaceae (referred to herein as "wild-type AEQH type II Cas"), a type II Cas protein from an unclassified bacterium in the family Ruminococcaceae (referred to herein as "wild-type AAOF type II Cas"), a type II Cas protein from an unclassified bacterium in the class Clostridia (referred to herein as "wild-type ACEE type II Cas"), a type II Cas protein from an unclassified bacterium in the family Ruminococcaceae (referred to herein as "wild-type AQSL type II Cas"), a type II Cas protein from an unclassified bacterium in the family Proteobacterium (referred to herein as "wild-type ASWC type II Cas"), a type II Cas protein from an unclassified bacterium in the family Ruminococcaceae (referred to herein as "wild-type ASWC type II Cas"), a type II Cas protein from an unclassified bacterium in the family Proteobacterium (referred to herein as "wild-type ASWC type II Cas"). type II Cas"), a type II Cas protein from an unclassified bacterium in the family Clostridiaceae (referred to herein as "wild-type AVFG type II Cas"), a type II Cas protein from an unclassified bacterium in the genus Gemmiger (referred to herein as "wild-type AWIT type II Cas"), a type II Cas protein from an unclassified bacterium in the genus Gemmiger (referred to herein as "wild-type AWMF type II Cas"), a type II Cas protein from an unclassified bacterium in the phylum Firmicutes (referred to herein as "wild-type BUMO type II Cas"), a type II Cas protein from an unclassified bacterium in the genus Gemmiger (referred to herein as "wild-type COIA type II Cas"), a type II Cas protein from an unclassified bacterium in the genus Gemmiger (referred to herein as "wild-type DQJA This work is based on the discovery of a type II Cas protein from an unclassified bacterium in the family Ruminococcaceae (referred to herein as "wild-type DWET type II Cas"). The wild-type AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins are each approximately 1000 amino acids in length, significantly shorter than SpCas9.

[0005] In one aspect, the disclosure provides a type II Cas protein (such a protein is referred to herein as an "AEQH type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 1. Exemplary AEQH type II Cas protein sequences are set forth in SEQ ID NO: 1, SEQ ID NO: 2, and SEQ ID NO: 3.

[0006] In one aspect, the disclosure provides a Type II Cas protein (such a protein is referred to herein as an "AAOF Type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 7. Exemplary AAOF Type II Cas protein sequences are set forth in SEQ ID NO: 7, SEQ ID NO: 8, and SEQ ID NO: 9.

[0007] In one aspect, the present disclosure provides a type II Cas protein (such a protein is referred to herein as an "ACEE type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 13. Exemplary ACEE type II Cas protein sequences are set forth in SEQ ID NO: 13, SEQ ID NO: 14, and SEQ ID NO: 15.

[0008] In one aspect, the disclosure provides a Type II Cas protein (such a protein is referred to herein as an "AQSL Type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 19. Exemplary AQSL Type II Cas protein sequences are set forth in SEQ ID NO: 19, SEQ ID NO: 20, and SEQ ID NO: 21.

[0009] In one aspect, the disclosure provides a Type II Cas protein (such a protein is referred to herein as an "ASWC Type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 25. Exemplary ASWC Type II Cas protein sequences are set forth in SEQ ID NO: 25, SEQ ID NO: 26, and SEQ ID NO: 27.

[0010] In one aspect, the disclosure provides a type II Cas protein (such a protein is referred to herein as an "AVFG type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 31. Exemplary AVFG type II Cas protein sequences are set forth in SEQ ID NO: 31, SEQ ID NO: 32, and SEQ ID NO: 33.

[0011] In one aspect, the disclosure provides a type II Cas protein (such a protein is referred to herein as an "AWIT type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 37. Exemplary AWIT type II Cas protein sequences are set forth in SEQ ID NO: 37, SEQ ID NO: 38, and SEQ ID NO: 39.

[0012] In one aspect, the disclosure provides a Type II Cas protein (such a protein is referred to herein as an "AWMF Type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 43. Exemplary AWMF Type II Cas protein sequences are set forth in SEQ ID NO: 43, SEQ ID NO: 44, and SEQ ID NO: 45.

[0013] In one aspect, the disclosure provides a Type II Cas protein (such a protein is referred to herein as a "BUMO Type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 49. Exemplary BUMO Type II Cas protein sequences are set forth in SEQ ID NO: 49, SEQ ID NO: 50, and SEQ ID NO: 51.

[0014] In one aspect, the disclosure provides a type II Cas protein (such a protein is referred to herein as a "COIA type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 55. Exemplary COIA type II Cas protein sequences are set forth in SEQ ID NO: 55, SEQ ID NO: 56, and SEQ ID NO: 57.

[0015] In one aspect, the disclosure provides a type II Cas protein (such a protein is referred to herein as a "DJQA type II Cas protein") comprising an amino acid sequence that is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 61. Exemplary DJQA type II Cas protein sequences are set forth in SEQ ID NO: 61, SEQ ID NO: 62, and SEQ ID NO: 63.

[0016] In one aspect, the disclosure provides a type II Cas protein (such a protein is referred to herein as a "DWET type II Cas protein"), the amino acid sequence of which is at least 50% identical (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% identical, or more) to SEQ ID NO: 68. Exemplary DWET type II Cas protein sequences are set forth in SEQ ID NO: 67, SEQ ID NO: 68, and SEQ ID NO: 69.

[0017] In another aspect, the disclosure provides a type II Cas protein comprising an amino acid sequence having at least 50% (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95%, or more) sequence identity to the RuvC-I domain, RuvC-II domain, RuvC-III domain, BH domain, REC domain, HNH domain, WED domain, or PID domain of an AEQH type II Cas protein, an AAOF type II Cas protein, an ACEE type II Cas protein, an AQSL type II Cas protein, an ASWC type II Cas protein, an AVFG type II Cas protein, an AWIT type II Cas protein, an AWMF type II Cas protein, a BUMO type II Cas protein, a COIA type II Cas protein, a DJQA type II Cas protein, or a DWET type II Cas protein. In some embodiments, a Type II Cas protein of the present disclosure is a chimeric Type II Cas protein that includes one or more domains from, for example, an AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and / or DWET Type II Cas protein and one or more domains from a different Type II Cas protein, such as SpCas9.

[0018] In some embodiments, the type II Cas proteins of the present disclosure are in the form of a fusion protein, including, for example, an AEQH type II Cas protein, an AAOF type II Cas protein, an ACEE type II Cas protein, an AQSL type II Cas protein, an ASWC type II Cas protein, an AVFG type II Cas protein, an AWIT type II Cas protein, an AWMF type II Cas protein, a BUMO type II Cas protein, a COIA type II Cas protein, a DJQA type II Cas protein, or a DWET type II Cas protein sequence fused to one or more additional amino acid sequences, for example, one or more nuclear localization signals and / or one or more tags. Other exemplary fusion partners can enable base editing (e.g., when the fusion partner is a nucleoside deaminase) or prime editing (e.g., when the fusion partner is a reverse transcriptase).

[0019] Exemplary features of the Type II Cas proteins of the present disclosure are described in Section 6.2 below, and in specific embodiments 1-295 and 1132-1138.

[0020] In further aspects, the present disclosure provides guide (gRNA) molecules, such as single guide RNAs (sgRNAs), and combinations of two or more gRNA molecules (e.g., combinations of sgRNA molecules). In various embodiments, the present disclosure provides a gRNA that can be used with an AEQH type II Cas protein of the present disclosure, a gRNA that can be used with an AAOF type II Cas protein of the present disclosure, a gRNA that can be used with an ACEE type II Cas protein of the present disclosure, a gRNA that can be used with an AQSL type II Cas protein of the present disclosure, a gRNA that can be used with an ASWC type II Cas protein of the present disclosure, a gRNA that can be used with an AVFG type II Cas protein of the present disclosure, a gRNA that can be used with an AWIT type II Cas protein of the present disclosure, a gRNA that can be used with an AWMF type II Cas protein of the present disclosure, a gRNA that can be used with a BUMO type II Cas protein of the present disclosure, a gRNA that can be used with a COIA type II Cas protein of the present disclosure, a gRNA that can be used with a DJQA type II Cas protein of the present disclosure, and a gRNA that can be used with a DWET type II Cas protein of the present disclosure. Exemplary features of gRNAs and gRNA combinations of the present disclosure are described in Section 6.3 below, and in specific embodiments 296-1007.

[0021] In a further aspect, the present disclosure provides a system comprising a Type II Cas protein of the present disclosure and one or more gRNAs, e.g., sgRNAs. For example, the system can comprise a ribonucleoprotein (RNP) comprising the Type II Cas protein complexed with a gRNA, e.g., sgRNA, or separate crRNA and tracrRNA. Exemplary features of the system are described in Section 6.4 below, and in specific embodiments 1008-1072.

[0022] In another aspect, the present disclosure provides a nucleic acid or nucleic acids encoding a Type II Cas protein of the present disclosure, and optionally a guide RNA, e.g., an sgRNA. In some embodiments, the nucleic acid comprises a Type II Cas protein of the present disclosure operably linked to a heterologous promoter, e.g., a mammalian promoter, e.g., a human promoter.

[0023] In another aspect, the disclosure provides a nucleic acid encoding a gRNA, e.g., an sgRNA, of the disclosure, and optionally a Type II Cas protein, e.g., an AEQH Type II Cas protein, an AAOF Type II Cas protein, an ACEE Type II Cas protein, an AQSL Type II Cas protein, an ASWC Type II Cas protein, an AVFG Type II Cas protein, an AWIT Type II Cas protein, an AWMF Type II Cas protein, a BUMO Type II Cas protein, a COIA Type II Cas protein, a DJQA Type II Cas protein, or a DWET Type II Cas protein.

[0024] In another aspect, the present disclosure provides a combination of gRNAs of the present disclosure, for example, a combination of two gRNAs, and optionally, a nucleic acid encoding a type II Cas protein.

[0025] Exemplary features of the nucleic acids and nucleic acids of the present disclosure are described in Section 6.5 below, and in specific embodiments 1073-1131.

[0026] In a further aspect, the present disclosure provides particles comprising the Type II Cas proteins, gRNAs, nucleic acids, and systems of the present disclosure. Exemplary features of the particles of the present disclosure are described in Section 6.6, below, and in specific embodiments 1139-1154.

[0027] In another aspect, the present disclosure provides cells and cell populations comprising or contacted with a Type II Cas protein, gRNA, nucleic acid, nucleic acid, system, or particle of the present disclosure. Exemplary characteristics of such cells and cell populations are described in Section 6.6, below, and in specific embodiments 1156-1165 and 1205.

[0028] In another aspect, the present disclosure provides pharmaceutical compositions comprising a Type II Cas protein, a gRNA, a nucleic acid, a plurality of nucleic acids, a system, a particle, a cell, or a population of cells, together with one or more excipients. Exemplary features of the pharmaceutical compositions are described in Section 6.7 below, and in specific embodiment 1155.

[0029] In another aspect, the present disclosure provides methods for modifying cells (e.g., editing the genome of a cell) using the disclosed type II Cas proteins, gRNAs, nucleic acids, systems, particles, and pharmaceutical compositions. Cells modified according to the disclosed methods can be used, for example, to treat subjects with a disease or disorder, such as a genetic disease or disorder, such as retinitis pigmentosa caused by an RHO mutation. Features of exemplary methods for modifying cells are described in Section 6.8 below, and in specific embodiments 1166-1204. [Brief explanation of the drawings]

[0030] 5. Brief description of the drawings [Figure 1] 1 shows the PAM logo for exemplary type II Cas proteins of the present disclosure. [Figure 2A]Schematic representations of hairpin structures generated for visualization after in silico folding using RNA Folding Form v2.3 (www.unafold.org) of sgRNA scaffolds (without spacer sequences) designed from the identified crRNAs and tracrRNAs for the AAOF type II Cas protein (Figure 2A), ACEE type II Cas protein (Figure 2B), AEQH type II Cas protein (Figure 2C), and AQSL type II Cas protein (Figure 2D). Figures 2A-2D disclose SEQ ID NOs: 124, 126, 122, and 128, respectively, in order of appearance. [Figure 2B] (As mentioned above.) [Figure 2C] (As mentioned above.) [Figure 2D] (As mentioned above.) [Figure 3A] Schematic representations of hairpin structures generated for visualization after in silico folding using RNA Folding Form v2.3 (www.unafold.org) of sgRNA scaffolds (without spacer sequences) designed from the identified crRNAs and tracrRNAs for the ASWC type II Cas protein (Figure 3A), AVFG type II Cas protein (Figure 3B), AWIT, AWMF, and COIA type II Cas proteins (Figure 3C), and BUMO type II Cas protein (Figure 3D). Figures 3A-3D disclose SEQ ID NOs: 130, 132, 134, and 138, respectively, in order of appearance. [Figure 3B] (As mentioned above.) [Figure 3C] (As mentioned above.) [Figure 3D] (As mentioned above.) [Figure 4A]Schematic representations of hairpin structures generated for visualization after in silico folding using RNA Folding Form v2.3 (www.unafold.org) of sgRNA scaffolds (without spacer sequences) designed from the identified crRNA and tracrRNA for the DJQA type II Cas protein (Figure 4A) and the DWET type II Cas protein (Figure 4B). Figures 4A-4B disclose SEQ ID NOs: 142 and 144, respectively, in order of appearance. [Figure 4B] (As mentioned above.) [Figure 5A] Figure 5A shows PAM sequence logos for AAOF type II Cas (Figure 5A), ACEE type II Cas (Figure 5C), and AEQH type II Cas (Figure 5E) type II Cas proteins from in vitro PAM probing assays, as well as PAM enrichment heat maps calculated from the same in vitro PAM probing assays showing nucleotide preferences at different positions along the PAM for AAOF type II Cas (positions 5, 6, and 7, 8) (Figure 5B), ACEE type II Cas (positions 5, 6, and 7, 8) (Figure 5D), and AEQH type II Cas (positions 3, 4, and 5, 6) (Figure 5F). [Figure 5B] (As mentioned above.) [Figure 5C] (As mentioned above.) [Figure 5D] (As mentioned above.) [Figure 5E] (As mentioned above.) [Figure 5F] (As mentioned above.) [Figure 6A] Figure 6A shows PAM sequence logos for AQSL type II Cas (Figure 6A), ASWC type II Cas (Figure 6C), and AVFG type II Cas (Figure 6E) from in vitro PAM probing assays, as well as PAM enrichment heat maps calculated from the same in vitro PAM probing assays showing nucleotide preferences at different positions along the PAM for AQSL type II Cas (positions 5, 6, and 7, 8) (Figure 6B), ASWC type II Cas (positions 2, 3, and 5, 6) (Figure 6D), and AVFG type II Cas (positions 5, 6, and 7, 8) (Figure 6F). [Figure 6B] (As mentioned above.) [Figure 6C] (As mentioned above.) [Figure 6D] (As mentioned above.) [Figure 6E] (As mentioned above.) [Figure 6F] (As mentioned above.) [Figure 7A] Figure 7A shows PAM sequence logos for AWIT type II Cas (Figure 7A), AWMF type II Cas (Figure 7C), and COIA type II Cas (Figure 7E) from in vitro PAM probing assays, as well as PAM enrichment heat maps calculated from the same in vitro PAM probing assays showing nucleotide preferences at different positions along the PAM for AWIT type II Cas (positions 5, 6, and 7, 8) (Figure 7B), AWMF type II Cas (positions 5, 6, and 7, 8) (Figure 7D), and COIA type II Cas (positions 5, 6, and 7, 8) (Figure 7F). [Figure 7B] (As mentioned above.) [Figure 7C] (As mentioned above.) [Figure 7D] (As mentioned above.) [Figure 7E] (As mentioned above.) [Figure 7F] (As mentioned above.) [Figure 8A] Figure 8A shows PAM sequence logos for BUMO type II Cas (Figure 8A), DJQA type II Cas (Figure 8C), and DWET type II Cas (Figure 8E) from in vitro PAM probing assays, as well as PAM enrichment heat maps calculated from the same in vitro PAM probing assays showing nucleotide preferences at different positions along the PAM for BUMO type II Cas (positions 5, 6, and 7, 8) (Figure 8B), DJQA type II Cas (positions 5, 6, and 7, 8) (Figure 8D), and DWET type II Cas (positions 5, 6, and 7, 8) (Figure 8F). [Figure 8B] (As mentioned above.) [Figure 8C] (As mentioned above.) [Figure 8D] (As mentioned above.) [Figure 8E] (As mentioned above.) [Figure 8F] (As mentioned above.) [Figure 9] Figure 1 shows the activity of selected type II Cas proteins assessed after transient electroporation of plasmids encoding each nuclease together with the indicated guide RNAs in U2OS cells stably expressing EGFP. [Figure 10] (Includes subsection Figure 10-1) A schematic diagram of the rs7984 SNP locus with exemplary sgRNA locations for the AAOF, ACEE, AEQH, AVFG, BUMO, DJQA, and DWET Type II Cas proteins (Example 3). Figure 10 discloses, in order of appearance, SEQ ID NOs: 494, 491, 491, 485, 493, 493, 487, 487, 490, 490, 492, 495, 495, 859, 486, 486, 484, 496, 497, 488, and 488, respectively. [Figure 10-1] (As mentioned above.) [Figure 11] Example 3 shows the reported editing activity of AAOF, ACEE, AEQH, AVFG, BUMO, DJQA, and DWET type II Cas proteins in combination with selected sgRNAs against the rs7984A RHO SNP allele after transient plasmid transfection in HEK293T cells. Data are reported as the mean ± SEM for n=2 independent trials. [Figure 12A] The ACEE (Figure 12A), DJQA (Figure 12B), and AVFG (Figure 12C) type II Cas proteins show editing activity for the rs7984A RHO SNP allele using guides with different spacer lengths (20-24 nucleotides) (Example 3). Data are reported as mean ± SEM for n=2 independent trials. [Figure 12B] (As mentioned above.) [Figure 12C] (As mentioned above.) [Figure 13A]The editing activity and allele specificity of selected gRNAs for the RHO SNP locus are shown (Example 3). Figure 13A shows the activity and allele specificity for the rs7984A SNP allele after transient transfection of HEK293T cells with ACEE, DJQA, or AVFG Type II Cas in combination with the illustrated guides that perfectly match either the rs7984A locus (on-target activity) or the rs7984G locus (off-target activity - allele specificity). Figure 13B shows the same experiment performed on HEK293T-rs7984G, which is homozygous for the rs7984G allele; guides targeting the rs7984G locus were used to measure on-target activity, and guides targeting the rs7984A locus were used to assess allele specificity. Data are reported as mean ± SEM for n=2 independent trials. [Figure 13B] (As mentioned above.) [Figure 14A] 14A-14B (including subsections Figures 14A-1 to 14A-7 and Figures 14B-1 to 14B-7) show schematic diagrams of RHO intron 1 sequences and report the locations of guide RNAs evaluated for ACEE Type II Cas (Figure 14A) and DJQA Type II Cas (Figure 14B) (Example 3). Figures 14A-14B disclose, in order of appearance, SEQ ID NOs: 521, 860, 518, 520, 517, 519, 862-863, 865, 866, 868, 532, 528, 533, 869-873, 535-536, and 875, respectively. [Figure 14A-1] (As mentioned above.) [Figure 14A-2] (As mentioned above.) [Figure 14A-3] (As mentioned above.) [Figure 14A-4] (As mentioned above.) [Figure 14A-5] (As mentioned above.) [Figure 14A-6] (As mentioned above.) [Figure 14A-7] (As mentioned above.) [Figure 14B] (As mentioned above.) [Figure 14B-1] (As mentioned above.) [Figure 14B-2] (As mentioned above.) [Figure 14B-3] (As mentioned above.) [Figure 14B-4] (As mentioned above.) [Figure 14B-5] (As mentioned above.) [Figure 14B-6] (As mentioned above.) [Figure 14B-7] (As mentioned above.) [Figure 15A] 15A and 15B show evaluation of the editing activity of guide RNAs for ACEE (FIG. 15A) and DJQA (FIG. 15B) type II Cas targeting RHO intron 1 after transient transfection in HEK293T cells (Example 3). [Figure 15B] (As mentioned above.) [Figure 16A] Evaluation of large-scale editing events at the target RHO locus (Example 3) is shown. Figure 16A shows a representative image of agarose gel electrophoresis of end-point PCR products generated with primers spanning the deleted RHO fragment using genomic DNA extracted from HEK293T cells transfected with the indicated Type II Cas proteins and associated guides. The high-molecular-weight band corresponds to the non-deleted / inverted product, while the low-molecular-weight band corresponds to the deleted RHO allele. Figure 16B reports the results of a qPCR assay utilizing primers specifically designed to assess the relative abundance of RHO alleles without large edits (deletions / inversions) in populations of HEK293T cells transiently transfected with expression plasmids for ACEE or DJQA Type II Cas together with selected sgRNAs, as indicated on the graph. NT: Untreated cells included as unmodified controls. Data are reported as mean ± SEM for n=2 independent trials. [Figure 16B] (As mentioned above.) [Figure 17]Figure 3 shows the levels of deletions and inversions measured by a custom-designed ddPCR assay generated in HEK293T cells after transient transfection of either ACEE or DJQA Type II Cas in combination with the indicated sgRNA targeting the rs7984A RHO SNP allele and RHO intron 1 (Example 3). NT: Untreated cells included as an unmodified control. Data are reported as the mean ± SEM for n=2 independent trials, except for the sample containing ACEE SNP g1, where n=1. [Figure 18A] Figure 18 shows the editing activity of AVFG, ACEE, AEQH, DJQA, and DWET type II Cas in combination with a panel of sgRNAs targeting TRAC (Figure 18A), B2M (Figure 18B), and PD1 (Figure 18C) after transient plasmid transfection in HEK293T cells (Example 4). Data are shown as the mean ± SEM for n=2 independent trials. [Figure 18B] (As mentioned above.) [Figure 18C] (As mentioned above.) DETAILED DESCRIPTION OF THE INVENTION

[0031] 6. Detailed Description In one aspect, the present disclosure provides type II Cas proteins (e.g., AEQH type II Cas protein, AAOF type II Cas protein, ACEE type II Cas protein, AQSL type II Cas protein, ASWC type II Cas protein, AVFG type II Cas protein, AWIT type II Cas protein, AWMF type II Cas protein, BUMO type II Cas protein, COIA type II Cas protein, DJQA type II Cas protein, and DWET type II Cas protein). The type II Cas proteins of the present disclosure can be in the form of a fusion protein. Unless otherwise required by context, the disclosure regarding type II Cas proteins encompasses type II Cas proteins that are not fusion proteins as well as type II Cas proteins in the form of fusion proteins (e.g., type II Cas proteins that include one or more nuclear localization signals and / or one or more tags).

[0032] In some embodiments, a type II Cas protein of the present disclosure comprises an amino acid sequence having at least 50% (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, at least 95% or more) sequence identity to the RuvC-I domain, RuvC-II domain, RuvC-III domain, BH domain, REC domain, HNH domain, WED domain, or PID domain of an AEQH type II Cas protein, an AAOF type II Cas protein, an ACEE type II Cas protein, an AQSL type II Cas protein, an ASWC type II Cas protein, an AVFG type II Cas protein, an AWIT type II Cas protein, an AWMF type II Cas protein, a BUMO type II Cas protein, a COIA type II Cas protein, a DJQA type II Cas protein, or a DWET type II Cas protein. In some embodiments, a type II Cas protein of the disclosure is a chimeric type II Cas protein comprising one or more domains from, for example, an AEQH type II Cas protein, and / or an AAOF type II Cas protein, and / or an ACEE type II Cas protein, and / or an AQSL type II Cas protein, and / or an ASWC type II Cas protein, and / or an AVFG type II Cas protein, and / or an AWIT type II Cas protein, and / or an AWMF type II Cas protein, and / or a BUMO type II Cas protein, and / or a COIA type II Cas protein, and / or a DJQA type II Cas protein, and / or a DWET type II Cas protein, and one or more domains from a different type II Cas protein, e.g., SpCas9.

[0033] Exemplary features of the Type II Cas proteins of the present disclosure are described in Section 6.2.

[0034] In further aspects, the present disclosure provides guide (gRNA) molecules, e.g., single guide RNAs (sgRNAs), and combinations of guide RNA molecules, e.g., combinations of two or more sgRNAs. A gRNA combination can include, for example, a gRNA targeting the RHO rs7984 SNP and a second gRNA targeting RHO intron 1. A combination of gRNAs targeting the RHO rs7984 SNP and RHO intron 1 can be used to selectively edit RHO alleles carrying pathogenic mutations. This dual-targeting approach is further described in Section 6.8 and Example 3. Exemplary features of the gRNAs and gRNA combinations of the present disclosure are further described in Section 6.3.

[0035] In a further aspect, the present disclosure provides a system comprising a Type II Cas protein of the present disclosure and one or more gRNAs, e.g., sgRNAs. Exemplary features of the system are described in Section 6.4.

[0036] In further aspects, the present disclosure provides nucleic acids and nucleic acids encoding a Type II Cas protein of the present disclosure and optionally a guide RNA, e.g., an sgRNA, and provides nucleic acids encoding a gRNA, e.g., an sgRNA, and optionally a Type II Cas protein of the present disclosure. Exemplary features of the nucleic acids and nucleic acids of the present disclosure are described in Section 6.5.

[0037] In a further aspect, the present disclosure provides particles comprising the Type II Cas proteins, gRNAs, nucleic acids, and systems of the present disclosure. Exemplary features of the particles of the present disclosure are described in Section 6.6.

[0038] In another aspect, the present disclosure provides cells and cell populations comprising or contacted with a Type II Cas protein, gRNA, nucleic acid, nucleic acid, system, or particle of the present disclosure. Exemplary characteristics of such cells and cell populations are described in Section 6.6.

[0039] In another aspect, the present disclosure provides a pharmaceutical composition comprising a Type II Cas protein, a gRNA, a nucleic acid, a plurality of nucleic acids, a system, a particle, a cell, or a population of cells, together with one or more excipients. Exemplary features of the pharmaceutical composition are described in Section 6.7.

[0040] In another aspect, the present disclosure provides methods for modifying cells (e.g., editing the genome of a cell) using the Type II Cas proteins, gRNAs, nucleic acids, systems, particles, and pharmaceutical compositions of the present disclosure. Features of exemplary methods for modifying cells are described in Section 6.8.

[0041] Those skilled in the art will recognize and appreciate that many changes can be made to the various embodiments described herein while still obtaining the beneficial results of the present disclosure. It will also be apparent that some of the desired advantages of the present disclosure can be obtained by selecting some of the features of the present disclosure without utilizing other features. Accordingly, those skilled in the art will recognize that many modifications and adaptations to the present disclosure are possible and may be desirable in particular circumstances and are a part of this disclosure. Accordingly, the following description is provided as an illustration of the principles of the present disclosure, not as a limitation thereof.

[0042] 6.1.Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The following definitions are provided for a full understanding of the terms used herein.

[0043] As used in this specification and claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "an agent" includes multiple agents, including mixtures thereof.

[0044] Unless otherwise indicated, the conjunction "or" is intended to be used in its precise sense as a Boolean logic operator, encompassing both the selection of features in an alternative (A or B, where the selection of A is mutually exclusive with B) and the selection of features concomitantly (A or B, where A and B are both selected). In some places in the text, the term "and / or" is used interchangeably and should not be construed to imply that "or" is being used in reference to mutually exclusive alternatives.

[0045] Type II Cas protein refers to wild-type or engineered type II Cas protein. Engineered type II Cas protein can also be referred to as type II Cas variant. For the avoidance of doubt, any disclosure of "type II Cas" or "type II Cas protein" relates to wild-type type II Cas protein and type II Cas variant, unless the context dictates otherwise. Type II Cas protein can have nuclease activity or be catalytically inactive (e.g., as in dCas).

[0046] As used herein, the percent identity between two nucleotide sequences or two amino acid sequences is calculated by multiplying the number of matches between aligned sequence pairs by 100 and dividing by the length of the aligned region. Identity scoring only counts perfect matches, and does not consider the degree of similarity between amino acids, nor does it consider substitutions or deletions as matches. Alignment for determining percent sequence identity can be achieved by various methods within the skill of the art, for example, by manual alignment or by using publicly available computer software such as BLAST, BLAST-2, ALIGN, ALIGN-2 or Megalign (DNASTAR) software. Those skilled in the art can determine the appropriate parameters to achieve maximum alignment.

[0047] A guide RNA molecule (gRNA) refers to an RNA that has the ability to form a complex with a type II Cas protein and direct the type II Cas protein to a target DNA. A gRNA typically includes a spacer of 15 to 30 nucleotides in length. In some embodiments, the gRNA of the present disclosure is a single guide RNA (sgRNA), which typically includes a spacer at the 5' end of the molecule and a 3' sgRNA scaffold. Various non-limiting examples of 3' sgRNA scaffolds are described in Section 6.3.

[0048] In some embodiments, the sgRNA may not include a uracil base at the 3' end of the sgRNA sequence. Alternatively, the sgRNA may include one or more uracil bases at the 3' end of the sgRNA sequence. For example, the sgRNA may include one uracil (U) at the 3' end of the sgRNA sequence, two uracils (UU) at the 3' end of the sgRNA sequence, three uracils (UUU) at the 3' end of the sgRNA sequence, four uracils (UUUU) at the 3' end of the sgRNA sequence, five uracils (UUUUU) at the 3' end of the sgRNA sequence, six uracils (UUUUUU) at the 3' end of the sgRNA sequence, seven uracils (UUUUUUU) at the 3' end of the sgRNA sequence, or eight uracils (UUUUUUUU) at the 3' end of the sgRNA sequence. Uracil stretches of various lengths can be added to the 3' end of the sgRNA as terminators. Thus, for example, the 3' sgRNA scaffold described in Section 6.3 can be modified by adding or removing one or more uracils to the end of the sequence.

[0049] Peptide, protein, and polypeptide are used interchangeably to refer to natural or synthetic molecules containing two or more amino acids linked by the carboxyl group of one amino acid to the alpha-amino group of another. The amino acids may be natural or synthetic and may include chemical modifications such as disulfide bridges, radioisotope substitution, phosphorylation, substrate chelation (e.g., iron or copper chelation), glycosylation, acetylation, formylation, amidation, biotinylation, and a wide range of other modifications. Polypeptides can bind to other molecules, such as those required for function. Examples of molecules to which polypeptides can bind include, but are not limited to, cofactors, polynucleotides, lipids, metal ions, phosphate groups, etc. Non-limiting examples of polypeptides include peptide fragments, denatured / unstructured polypeptides, polypeptides with quaternary or aggregated structures, etc. There is no requirement that a polypeptide have an intended function; a polypeptide may be functional, nonfunctional, function for an unexpected / unintended purpose, or have an unknown function. Polypeptides are composed of approximately the 20 standard naturally occurring amino acids, although natural and synthetic amino acids that are not members of the standard 20 amino acids may also be used. The standard 20 amino acids include alanine (Ala, A), arginine (Arg, R), asparagine (Asn, N), aspartic acid (Asp, D), cysteine ​​(Cys, C), glutamine (Gln, Q), glutamic acid (Glu, E), glycine (Gly, G), histidine (His, H), isoleucine (Ile, I), leucine (Leu, L), lysine (Lys, K), methionine (Met, M), phenylalanine (Phe, F), proline (Pro, P), serine (Ser, S), threonine (Thr, T), tryptophan (Trp, W), tyrosine (Tyr, Y), and valine (Val, V). The term "polypeptide sequence" or "amino acid sequence" is the alphabetical representation of a polypeptide molecule.

[0050] Polynucleotide and oligonucleotide are used interchangeably and refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotides may have any three-dimensional structure and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: genes or gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, primers, and gRNA. Polynucleotides may contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by conjugation with a labeling component. A polynucleotide is composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and, if the polynucleotide is RNA, uracil (U) is substituted for thymine (T). Thus, the term "nucleotide sequence" is the alphabetical representation of a polynucleotide molecule. The letters used in the polynucleotide sequences described herein correspond to the IUPAC notation. For example, the letter "N" in a nucleotide sequence represents a nucleotide that can be A, T, C, or G in a DNA sequence, or A, U, C, or G in an RNA sequence; the letter "R" in a nucleotide sequence represents a nucleotide that can be A or G; and the letter "V" in a nucleotide sequence represents a nucleotide that can be A, C, or G.

[0051] A protospacer adjacent motif (PAM) refers to a DNA sequence downstream (e.g., immediately downstream) of a target sequence on the non-target strand that is recognized by a Type II Cas protein. The PAM sequence is located 3' to the target sequence on the non-target strand.

[0052] The spacer refers to a region of a gRNA molecule that is partially or completely complementary to a target sequence found on the plus or minus strand of genomic DNA. When complexed with a Type II Cas protein, the gRNA directs Type II Cas to the target sequence in the genomic DNA. The spacer of a Type II Cas gRNA is typically 15-30 nucleotides in length (e.g., 20-25 nucleotides). The nucleotide sequence of the spacer may be, but is not necessarily, completely complementary to the target sequence. For example, the spacer can contain one or more mismatches to the target sequence; for example, the spacer can contain one, two, or three mismatches to the target sequence.

[0053] 6.2.Type II Cas Proteins 6.2.1. Type IIA Cas Proteins AEQH type II Cas protein In one aspect, the present disclosure provides an AEQH type II Cas protein. An AEQH type II Cas protein can be further classified as a type IIA Cas protein. An AEQH type II Cas protein typically comprises an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO:1. In some embodiments, an AEQH type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:1. In some embodiments, an AEQH type II Cas protein comprises an amino acid sequence identical to SEQ ID NO:1.

[0054] Exemplary AEQH type II Cas protein sequences and nucleotide sequences encoding exemplary AEQH type II Cas proteins are set forth in Table 1A.

[0055] [Table 1-1]

[0056] [Table 1-2]

[0057] [Table 1-3]

[0058] [Table 1-4]

[0059] [Table 1-5]

[0060] In some embodiments, the AEQH type II Cas protein comprises the amino acid sequence of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3. In some embodiments, the AEQH type II Cas protein has nickase activity resulting from one or more amino acid substitutions, for example, compared to the sequence of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3. In some embodiments, the one or more amino acid substitutions that confer nickase activity include a D10A substitution, where the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO:2. In some embodiments, the one or more amino acid substitutions that confer nickase activity include an N627A substitution, where the position of the N627A substitution is defined with respect to the amino acid numbering of SEQ ID NO:2. In some embodiments, the one or more amino acid substitutions that confer nickase activity include an H604A substitution, where the position of the H604A substitution is defined with respect to the amino acid numbering of SEQ ID NO:2. In some embodiments, the AEQH type II Cas protein is catalytically inactive, for example, resulting from a combination of the D10A substitution and the N627A or H604A substitution.

[0061] 6.2.2. Type IIC Cas Proteins 6.2.2.1.AAOF Type II Cas Protein In one aspect, the present disclosure provides an AAOF type II Cas protein. An AAOF type II Cas protein can be further classified as a type IIC Cas protein. An AAOF type II Cas protein typically comprises an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO:7. In some embodiments, an AAOF type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:7. In some embodiments, an AAOF type II Cas protein comprises an amino acid sequence identical to SEQ ID NO:7.

[0062] Exemplary AAOF Type II Cas protein sequences and nucleotide sequences encoding exemplary AAOF Type II Cas proteins are set forth in Table 2A.

[0063] [Table 2-1]

[0064] [Table 2-2]

[0065] [Table 2-3]

[0066] [Table 2-4]

[0067] [Table 2-5]

[0068] [Table 2-6]

[0069] In some embodiments, the AAOF Type II Cas protein comprises the amino acid sequence of SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. In some embodiments, the AAOF Type II Cas protein has nickase activity resulting from one or more amino acid substitutions, e.g., compared to the sequence of SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D9A substitution, wherein the position of the D9A substitution is defined with respect to the amino acid numbering of SEQ ID NO:8. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N610A substitution, wherein the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO:8. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H587A substitution, wherein the position of the H587A substitution is defined with respect to the amino acid numbering of SEQ ID NO:8. In some embodiments, the AAOF Type II Cas protein is catalytically inactive, e.g., resulting from a combination of the D9A substitution and the N610A or H587A substitution.

[0070] 6.2.2.2. ACEE Type II Cas Protein In one aspect, the present disclosure provides an ACEE type II Cas protein. ACEE type II Cas proteins can be further classified as type IIC Cas proteins. ACEE type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 13. In some embodiments, ACEE type II Cas proteins comprise an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 13. In some embodiments, ACEE type II Cas proteins comprise an amino acid sequence identical to SEQ ID NO: 13.

[0071] Exemplary ACEE type II Cas protein sequences and nucleotide sequences encoding exemplary ACEE type II Cas proteins are set forth in Table 2B.

[0072] [Table 3-1]

[0073] [Table 3-2]

[0074] [Table 3-3]

[0075] [Table 3-4]

[0076] [Table 3-5]

[0077] [Table 3-6]

[0078] In some embodiments, the ACEE type II Cas protein comprises the amino acid sequence of SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15. In some embodiments, the ACEE type II Cas protein has nickase activity resulting from one or more amino acid substitutions, for example, compared to the sequence of SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D8A substitution, where the position of the D8A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 14. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N613A substitution, where the position of the N613A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 14. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H590A substitution, where the position of the H590A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 14. In some embodiments, the ACEE type II Cas protein is catalytically inactive, for example, resulting from a combination of the D8A substitution with the N613A substitution or the H590A substitution.

[0079] AQSL Type II Cas Protein In one aspect, the present disclosure provides an AQSL type II Cas protein. AQSL type II Cas proteins can be further classified as type IIC Cas proteins. AQSL type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 19. In some embodiments, an AQSL type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 19. In some embodiments, an AQSL type II Cas protein comprises an amino acid sequence identical to SEQ ID NO: 19.

[0080] Exemplary AQSL Type II Cas protein sequences and nucleotide sequences encoding exemplary AQSL Type II Cas proteins are set forth in Table 2C.

[0081] [Table 4-1]

[0082] [Table 4-2]

[0083] [Table 4-3]

[0084] [Table 4-4]

[0085] [Table 4-5]

[0086] [Table 4-6]

[0087] In some embodiments, the AQSL type II Cas protein comprises the amino acid sequence of SEQ ID NO: 19, SEQ ID NO: 20, or SEQ ID NO: 21. In some embodiments, the AQSL type II Cas protein has nickase activity resulting from one or more amino acid substitutions, e.g., compared to the sequence of SEQ ID NO: 19, SEQ ID NO: 20, or SEQ ID NO: 21. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D9A substitution, where the position of the D9A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 20. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N610A substitution, where the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 20. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H587A substitution, where the position of the H587A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 20. In some embodiments, the AQSL type II Cas protein is catalytically inactive, e.g., resulting from a combination of the D9A substitution and the N610A or H587A substitution.

[0088] 6.2.2.4. ASWC Type II Cas Protein In one aspect, the present disclosure provides an ASWC type II Cas protein. ASWC type II Cas proteins can be further classified as type IIC Cas proteins. ASWC type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO:25. In some embodiments, ASWC type II Cas proteins comprise an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:25. In some embodiments, ASWC type II Cas proteins comprise an amino acid sequence identical to SEQ ID NO:25.

[0089] Exemplary ASWC Type II Cas protein sequences and nucleotide sequences encoding exemplary ASWC Type II Cas proteins are listed in Table 2D.

[0090] [Table 5-1]

[0091] [Table 5-2]

[0092] [Table 5-3]

[0093] [Table 5-4]

[0094] [Table 5-5]

[0095] [Table 5-6]

[0096] In some embodiments, the type II Cas protein of ASWC comprises the amino acid sequence of SEQ ID NO:25, SEQ ID NO:26, or SEQ ID NO:27. In some embodiments, the type II Cas protein of ASWC has nickase activity resulting from one or more amino acid substitutions, for example, compared to the sequence of SEQ ID NO:25, SEQ ID NO:26, or SEQ ID NO:27. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D15A substitution, where the position of the D15A substitution is defined with respect to the amino acid numbering of SEQ ID NO:26. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise an N633A substitution, where the position of the N633A substitution is defined with respect to the amino acid numbering of SEQ ID NO:26. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise an H610A substitution, where the position of the H610A substitution is defined with respect to the amino acid numbering of SEQ ID NO:26. In some embodiments, the type II Cas protein of ASWC is catalytically inactive, for example, resulting from a combination of the D15A substitution and the N633A or H610A substitution.

[0097] 6.2.2.5.AVFG type II Cas protein In one aspect, the present disclosure provides an AVFG type II Cas protein. AVFG type II Cas proteins can be further classified as type IIC Cas proteins. AVFG type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 31. In some embodiments, AVFG type II Cas proteins comprise an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 31. In some embodiments, AVFG type II Cas proteins comprise an amino acid sequence identical to SEQ ID NO: 31.

[0098] Exemplary AVFG type II Cas protein sequences and nucleotide sequences encoding exemplary AVFG type II Cas proteins are listed in Table 2E.

[0099] [Table 6-1]

[0100] [Table 6-2]

[0101] [Table 6-3]

[0102] [Table 6-4]

[0103] [Table 6-5]

[0104] In some embodiments, the AVFG type II Cas protein comprises the amino acid sequence of SEQ ID NO:31, SEQ ID NO:32, or SEQ ID NO:33. In some embodiments, the AVFG type II Cas protein has nickase activity resulting from one or more amino acid substitutions compared to, for example, the sequence of SEQ ID NO:31, SEQ ID NO:32, or SEQ ID NO:33. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D12A substitution, where the position of the D12A substitution is defined with respect to the amino acid numbering of SEQ ID NO:32. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N607A substitution, where the position of the N607A substitution is defined with respect to the amino acid numbering of SEQ ID NO:32. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H584A substitution, where the position of the H584A substitution is defined with respect to the amino acid numbering of SEQ ID NO:32. In some embodiments, the AVFG type II Cas protein is catalytically inactive, e.g., resulting from a combination of the D12A substitution and the N607A or H584A substitution.

[0105] 6.2.2.6.AWIT Type II Cas Protein In one aspect, the present disclosure provides an AWIT type II Cas protein. AWIT type II Cas proteins can be further classified as type IIC Cas proteins. AWIT type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 37. In some embodiments, an AWIT type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 37. In some embodiments, an AWIT type II Cas protein comprises an amino acid sequence identical to SEQ ID NO: 37.

[0106] Exemplary AWIT type II Cas protein sequences and nucleotide sequences encoding exemplary AWIT type II Cas proteins are listed in Table 2F.

[0107] [Table 7-1]

[0108] [Table 7-2]

[0109] [Table 7-3]

[0110] [Table 7-4]

[0111] [Table 7-5]

[0112] [Table 7-6]

[0113] In some embodiments, the AWIT type II Cas protein comprises the amino acid sequence of SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39. In some embodiments, the AWIT type II Cas protein has nickase activity resulting from one or more amino acid substitutions, e.g., compared to the sequence of SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D10A substitution, where the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO:38. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N615A substitution, where the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO:38. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H592A substitution, where the position of the H592A substitution is defined with respect to the amino acid numbering of SEQ ID NO:38. In some embodiments, the AWIT type II Cas protein is catalytically inactive, e.g., resulting from a combination of a D10A substitution and an N615A or H592A substitution.

[0114] 6.2.2.7.AWMF Type II Cas Protein In one aspect, the present disclosure provides an AWMF type II Cas protein. AWMF type II Cas proteins can be further classified as type IIC Cas proteins. AWMF type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO:43. In some embodiments, an AWMF type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:43. In some embodiments, an AWMF type II Cas protein comprises an amino acid sequence identical to SEQ ID NO:43.

[0115] Exemplary AWMF Type II Cas protein sequences and nucleotide sequences encoding exemplary AWMF Type II Cas proteins are set forth in Table 2G.

[0116] [Table 8-1]

[0117] [Table 8-2]

[0118] [Table 8-3]

[0119] [Table 8-4]

[0120] [Table 8-5]

[0121] [Table 8-6]

[0122] In some embodiments, the AWMF Type II Cas protein comprises the amino acid sequence of SEQ ID NO:43, SEQ ID NO:44, or SEQ ID NO:45. In some embodiments, the AWMF Type II Cas protein has nickase activity resulting from one or more amino acid substitutions, e.g., compared to the sequence of SEQ ID NO:43, SEQ ID NO:44, or SEQ ID NO:45. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D10A substitution, where the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO:44. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N615A substitution, where the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO:44. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H592A substitution, where the position of the H592A substitution is defined with respect to the amino acid numbering of SEQ ID NO:44. In some embodiments, the AWMF Type II Cas protein is catalytically inactive, e.g., resulting from a combination of the D10A substitution and the N615A or H592A substitution.

[0123] 6.2.2.8.BUMO Type II Cas Protein In one aspect, the disclosure provides a BUMO type II Cas protein. BUMO type II Cas proteins can be further classified as type IIC Cas proteins. BUMO type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO:49. In some embodiments, a BUMO type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:49. In some embodiments, a BUMO type II Cas protein comprises an amino acid sequence identical to SEQ ID NO:49.

[0124] Exemplary BUMO type II Cas protein sequences and nucleotide sequences encoding exemplary BUMO type II Cas proteins are set forth in Table 2H.

[0125] [Table 9-1]

[0126] [Table 9-2]

[0127] [Table 9-3]

[0128] [Table 9-4]

[0129] [Table 9-5]

[0130] In some embodiments, the BUMO type II Cas protein comprises the amino acid sequence of SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51. In some embodiments, the BUMO type II Cas protein has nickase activity resulting from one or more amino acid substitutions, for example, compared to the sequence of SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D8A substitution, where the position of the D8A substitution is defined with respect to the amino acid numbering of SEQ ID NO:50. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N610A substitution, where the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO:50. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H587A substitution, where the position of the H587A substitution is defined with respect to the amino acid numbering of SEQ ID NO:50. In some embodiments, the BUMO type II Cas protein is catalytically inactive, for example, resulting from a combination of the D8A substitution with the N610A or H587A substitution.

[0131] COIA type II Cas protein In one aspect, the present disclosure provides a COIA type II Cas protein. COIA type II Cas proteins can be further classified as type IIC Cas proteins. COIA type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 55. In some embodiments, a COIA type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 55. In some embodiments, a COIA type II Cas protein comprises an amino acid sequence identical to SEQ ID NO: 55.

[0132] Exemplary COIA type II Cas protein sequences and nucleotide sequences encoding exemplary COIA type II Cas proteins are listed in Table 2I.

[0133] [Table 10-1]

[0134] [Table 10-2]

[0135] [Table 10-3]

[0136] [Table 10-4]

[0137] [Table 10-5]

[0138] [Table 10-6]

[0139] In some embodiments, the COIA type II Cas protein comprises the amino acid sequence of SEQ ID NO:55, SEQ ID NO:56, or SEQ ID NO:57. In some embodiments, the COIA type II Cas protein has nickase activity resulting from one or more amino acid substitutions, e.g., compared to the sequence of SEQ ID NO:55, SEQ ID NO:56, or SEQ ID NO:57. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D10A substitution, where the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO:56. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a N615A substitution, where the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO:56. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a H592A substitution, where the position of the H592A substitution is defined with respect to the amino acid numbering of SEQ ID NO:56. In some embodiments, the COIA type II Cas protein is catalytically inactive, e.g., resulting from a combination of the D10A substitution with the N615A substitution or the H592A substitution.

[0140] 6.2.2.10.DJQA type II Cas protein In one aspect, the present disclosure provides a DJQA type II Cas protein. DJQA type II Cas proteins can be further classified as type IIC Cas proteins. DJQA type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 61. In some embodiments, a DJQA type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 61. In some embodiments, a DJQA type II Cas protein comprises an amino acid sequence identical to SEQ ID NO: 61.

[0141] Exemplary DJQA type II Cas protein sequences and nucleotide sequences encoding exemplary DJQA type II Cas proteins are shown in Table 2J.

[0142] [Table 11-1]

[0143] [Table 11-2]

[0144] [Table 11-3]

[0145] [Table 11-4]

[0146] [Table 11-5]

[0147] [Table 11-6]

[0148] In some embodiments, the DJQA type II Cas protein comprises the amino acid sequence of SEQ ID NO:61, SEQ ID NO:62, or SEQ ID NO:63. In some embodiments, the DJQA type II Cas protein has nickase activity resulting from one or more amino acid substitutions, e.g., compared to the sequence of SEQ ID NO:61, SEQ ID NO:62, or SEQ ID NO:63. In some embodiments, the one or more amino acid substitutions that confer nickase activity include a D10A substitution, where the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO:62. In some embodiments, the one or more amino acid substitutions that confer nickase activity include an N615A substitution, where the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO:62. In some embodiments, the one or more amino acid substitutions that confer nickase activity include an H592A substitution, where the position of the H592A substitution is defined with respect to the amino acid numbering of SEQ ID NO:62. In some embodiments, the DJQA type II Cas protein is catalytically inactive, e.g., resulting from a combination of a D10A substitution and an N615A or H592A substitution.

[0149] DWET Type II Cas Protein In one aspect, the present disclosure provides a DWET type II Cas protein. DWET type II Cas proteins can be further classified as type IIC Cas proteins. DWET type II Cas proteins typically comprise an amino acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to SEQ ID NO: 67. In some embodiments, a DWET type II Cas protein comprises an amino acid sequence at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 67. In some embodiments, a DWET type II Cas protein comprises an amino acid sequence identical to SEQ ID NO: 67.

[0150] Exemplary DWET type II Cas protein sequences and nucleotide sequences encoding exemplary DWET type II Cas proteins are listed in Table 2K.

[0151] [Table 12-1]

[0152] [Table 12-2]

[0153] [Table 12-3]

[0154] [Table 12-4]

[0155] [Table 12-5]

[0156] [Table 12-6]

[0157] In some embodiments, the DWET type II Cas protein comprises the amino acid sequence of SEQ ID NO:67, SEQ ID NO:68, or SEQ ID NO:69. In some embodiments, the DWET type II Cas protein has nickase activity resulting from one or more amino acid substitutions compared to the sequence of SEQ ID NO:67, SEQ ID NO:68, or SEQ ID NO:69, for example. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise a D9A substitution, where the position of the D9A substitution is defined with respect to the amino acid numbering of SEQ ID NO:68. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise an N610A substitution, where the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO:68. In some embodiments, the one or more amino acid substitutions that confer nickase activity comprise an H587A substitution, where the position of the H587A substitution is defined with respect to the amino acid numbering of SEQ ID NO:68. In some embodiments, the DWET type II Cas protein is catalytically inactive, for example, resulting from a combination of the D9A substitution and the N610A or H587A substitution.

[0158] 6.2.3. Fusion and Chimeric Proteins The present disclosure provides for Type II Cas proteins in the form of fusion proteins comprising a Type II Cas protein sequence fused to one or more additional amino acid sequences, e.g., one or more nuclear localization signals and / or one or more non-native tags (e.g., the AEQH Type II Cas protein described in Section 6.2.1.1, the AAOF Type II Cas protein described in Section 6.2.2.1, the ACEE Type II Cas protein described in Section 6.2.2.2, the AQSL Type II Cas protein described in Section 6.2.2.3, the AWSC Type II Cas protein described in Section 6.2.2.4, the AVFG Type II Cas protein described in Section 6.2.2.5, the AWIT Type II Cas protein described in Section 6.2.2.6, the AWMF Type II Cas protein described in Section 6.2.2.7, the BUMO Type II Cas protein described in Section 6.2.2.8, the COIA Type II Cas protein described in Section 6.2.2.9, the DJQA Type II Cas protein described in Section 6.2.2.10, and the like). Fusion proteins include, for example, a type II Cas protein, or a DWET type II Cas protein as described in Section 6.2.2.11. Fusion proteins can also include amino acid sequences of, for example, a nucleoside deaminase, a reverse transcriptase, a transcriptional activator (e.g., VP64), a transcriptional repressor (e.g., Krüppel-associated box (KRAB)), a histone-modifying protein, an integrase, or a recombinase.

[0159] In some embodiments, the fusion proteins of the present disclosure comprise a means for localizing a Type II Cas protein to the nucleus, e.g., a nuclear localization signal.

[0160] Non-limiting examples of nuclear localization signals include KRTADGSEFESPKKKRKV (SEQ ID NO: 145), PKKKRKV (SEQ ID NO: 146), PKKKRRV (SEQ ID NO: 147), KRPAATKKAGQAKKKK (SEQ ID NO: 148), YGRKKRRQRRR (SEQ ID NO: 149), RKKRRQRRR (SEQ ID NO: 150), PAAKRVKLD (SEQ ID NO: 151), RQRRNELKRSP (SEQ ID NO: 152), VSRKRPRP (SEQ ID NO: 153), PPKKARED (SEQ ID NO: 154), PQPKKKPL (SEQ ID NO: 155), SA These include LIKKKKKMAP (SEQ ID NO: 156), PKQKKRK (SEQ ID NO: 157), RKLKKKIKKL (SEQ ID NO: 158), REKKKFLKRR (SEQ ID NO: 159), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 160), RKCLQAGMNLEARKTKK (SEQ ID NO: 161), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 162), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 163).

[0161] Exemplary fusion partners include protein tags (e.g., a V5 tag (e.g., having the sequence GKPIPNPLLGLDST (SEQ ID NO: 164) or IPNPLLGLD (SEQ ID NO: 165)), a FLAG tag, a myc tag, an HA tag, a GST tag, a polyHis tag, an MBP tag), protein domains, transcriptional modulators, enzymes that act on small molecule substrates, DNA-modifying enzymes, RNA-modifying enzymes, and protein-modifying enzymes (e.g., adenosine deaminase, cytidine deaminase, guanosyltransferase, DNA methyltransferase, RNA methyltransferase, DNA demethylase, RNA demethylase, dioxygenase, polyadenylate polymerase, pseudouridine synthase, These include acetyltransferases, deacetylases, ubiquitin ligases, deubiquitinases, kinases, phosphatases, NEDD8 ligases, de-NEDD8 ligases, SUMO ligases, de-SUMOylases, histone deacetylases, reverse transcriptases, histone acetyltransferases, histone methyltransferases, histone demethylases), protein DNA binding domains, RNA binding proteins, polypeptide sequences with specific biological functions (e.g., nuclear localization signals, mitochondrial localization signals, plastid localization signals, cellular localization signals, destabilization signals, geminin destruction box motifs), and biological anchoring domains (e.g., MS2, Csy4, and lambda N proteins). Various type II Cas fusion proteins are described in Ribeiro et al., 2018, In. J. Genomics, Article ID: 1652567; Jayavaradhan, et al., 2019, Nat Commun 10:2866; Xiao et al., 2019, The CRISPR Journal, 2(1):51-63; Mali et al., 2013, Nat Methods. 10(10):957-63; U.S. Patent Nos. 9,322,037 and 9,388,430. In some embodiments, the fusion partner is adenosine deaminase.An exemplary adenosine deaminase is the tRNA adenosine deaminase (TadA) portion contained in the adenine base editor ABE8e (Richter, 2020, Nature Biotechnology 38:883-891). The TadA portion of ABE8e comprises the following amino acid sequence:

[0162] [ka]

[0163] In some embodiments, an adenosine deaminase fusion partner comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO: 166.

[0164] The disclosed type II Cas proteins in the form of fusion proteins with adenosine deaminase can be used as adenine base editors to change "A" to "G" in DNA. The disclosed type II Cas proteins in the form of fusion proteins with cytidine deaminase can be used as cytosine base editors to change "C" to "T" in DNA.

[0165] In some embodiments, the fusion protein of the present disclosure comprises a means for deaminating adenosine, e.g., an adenosine deaminase, e.g., a TadA variant. In some embodiments, the fusion protein of the present disclosure comprises a means for deaminating cytidine, e.g., a cytidine deaminase, e.g., cytidine deaminase 1 (CDA1) or an apolipoprotein B mRNA editing complex (APOBEC) family deaminase (Cheng et al., 2019, Nat Commun. 10(1):3612; Gehrke et al., 2018, Nat Biotechnol. 36(10):977-982).

[0166] In some embodiments, the fusion proteins of the present disclosure comprise a means for synthesizing DNA from a single-stranded template, such as a reverse transcriptase. The Type II Cas proteins of the present disclosure in the form of a fusion protein with reverse transcriptase (RT) can be used as a prime editor to perform precise base editing without double-stranded DNA cleavage.

[0167] In some embodiments, a fusion protein of the present disclosure is a prime editor, e.g., a type II Cas protein fused to a suitable RT (e.g., Moloney murine leukemia virus (M-MLV) RT or other RT enzyme). Such fusion proteins can be used with a prime editing guide RNA (pegRNA) that identifies the target site and encodes the desired editing (Anzalone et al., 2019, Nature, 576(7785):149-157).

[0168] In some embodiments, fusion proteins of the present disclosure comprise one or more nuclear localization signals located at the N-terminus and / or C-terminus of a Type II Cas protein sequence (e.g., an AEQH Type II Cas protein having the sequence of SEQ ID NO: 1). In some embodiments, fusion proteins of the present disclosure comprise N-terminal and C-terminal nuclear localization signals, each having the sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 145), for example.

[0169] The present disclosure provides chimeric type II Cas proteins comprising one or more domains of an AEQH type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), a chimeric type II Cas protein comprising one or more domains of an AAOF type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), a chimeric type II Cas protein comprising one or more domains of an ACEE type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), a chimeric type II Cas protein comprising one or more domains of an AQSL type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), a chimeric type II Cas protein comprising one or more domains of an ASWC type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), an AVFG type II Cas protein comprising one or more domains of an AQSL type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), a chimeric type II Cas protein comprising one or more domains of an ASWC ... chimeric type II Cas proteins comprising one or more domains of a type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins); AWIT chimeric type II Cas proteins comprising one or more domains of a type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins); AWMF chimeric type II Cas proteins comprising one or more domains of a type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins); BUMO chimeric type II Cas proteins comprising one or more domains of a type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins); COIAProvided are chimeric type II Cas proteins comprising one or more domains of a type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), chimeric type II Cas proteins comprising one or more domains of a DJQA type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins), and chimeric type II Cas proteins comprising one or more domains of a DWET type II Cas protein and one or more domains of one or more different proteins (e.g., one or more different type II Cas proteins).

[0170] The domain structures of wild-type AIK, BNK, HPLH, and ANAB type II Cas proteins were inferred by multiple alignment with the amino acid sequences of type II Cas proteins whose crystal structures are known, thus allowing the definition of the boundaries of each functional domain. The domains identified in type II Cas proteins are the RuvC catalytic domain (which is discontinuous and represented by the RuvC-I, RuvC-II, and RuvC-III domains), the bridge helix (BH), the recognition (REC) domain, the HNH catalytic domain, the wedge (WED) domain, and the PAM-interacting domain (PID).

[0171] Tables 3A-3B below report the amino acid positions corresponding to the boundaries between different functional domains in the wild-type AEQH (SEQ ID NO: 2), AAOF (SEQ ID NO: 8), ACEE (SEQ ID NO: 14), AQSL (SEQ ID NO: 20), ASWC (SEQ ID NO: 26), AVFG (SEQ ID NO: 32), AWIT (SEQ ID NO: 38), AWMF (SEQ ID NO: 44), BUMO (SEQ ID NO: 50), COIA (SEQ ID NO: 56), DJQA (SEQ ID NO: 62), and DWET (SEQ ID NO: 68) Type II Cas proteins.

[0172] [Table 13]

[0173] [Table 14]

[0174] The chimeric type II Cas protein may be an AAOF type II Cas protein, an ACEE type II Cas protein, an AEQH type II Cas protein, an AQSL type II Cas protein, an ASWC type II Cas protein, an AVFG type II Cas protein, an AWIT type II Cas protein, an AWMF type II Cas protein, a BUMO type II Cas protein, a COIA type II Cas protein, a DJQA type II Cas protein, and / or a DWET type II Cas protein. The chimeric protein may include one or more of the following domains (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more) from a type II Cas protein, as well as one or more domains from one or more other proteins, such as SaCas9, SpCas9, or a type II Cas protein described in U.S. Patent Application Publication Nos. 2020 / 0332273, 2019 / 0169648, or 2015 / 0247150 (each of which is incorporated by reference in its entirety): RuvC-I, BH, REC, RuvC-II, HNH, RuvC-III, WED, PID. For example, the PID domain can be swapped between different type II Cas proteins to alter the PAM specificity (conferred by the donor PID domain) of the resulting chimeric protein. Swapping other domains or portions thereof (e.g., by protein shuffling) is also within the scope of this disclosure.

[0175] In some embodiments, a type II Cas protein of the present disclosure comprises one, two, three, four, five, six, seven, or eight of a RuvC-I domain, a BH domain, a REC domain, a RuvC-II domain, an HNH domain, a RuvC-III domain, a WED domain, and a PID domain, arranged N-terminally to C-terminally. In some embodiments, all of the domains are derived from an AEQH type II Cas protein (e.g., an AEQH type II Cas protein whose amino acid sequence comprises SEQ ID NO: 1, 2, or 3). In some embodiments, all of the domains are derived from an AAOF type II Cas protein (e.g., an AAOF type II Cas protein whose amino acid sequence comprises SEQ ID NO: 7, 8, or 9). In some embodiments, all of the domains are derived from an ACEE type II Cas protein (e.g., an ACEE type II Cas protein whose amino acid sequence comprises SEQ ID NO: 13, 14, or 15). In some embodiments, all of the domains are derived from an AQSL type II Cas protein (e.g., an AQSL type II Cas protein whose amino acid sequence comprises SEQ ID NO: 19, 20, or 21). In some embodiments, all of the domains are derived from an AWSC type II Cas protein (e.g., an AWSC type II Cas protein whose amino acid sequence comprises SEQ ID NO: 25, 26, or 27). In some embodiments, all of the domains are derived from an AVFG type II Cas protein (e.g., an AVFG type II Cas protein whose amino acid sequence comprises SEQ ID NO: 31, 32, or 33). In some embodiments, all of the domains are derived from an AWIT type II Cas protein (e.g., an AWIT type II Cas protein whose amino acid sequence comprises SEQ ID NO: 37, 38, or 39). In some embodiments, all of the domains are derived from an AWMF type II Cas protein (e.g., an AWMF type II Cas protein whose amino acid sequence comprises SEQ ID NO: 43, 44, or 45). In some embodiments, all of the domains are derived from a BUMO type II Cas protein (e.g., a BUMO type II Cas protein whose amino acid sequence comprises SEQ ID NO: 49, 50, or 51).In some embodiments, all of the domains are derived from a COIA type II Cas protein (e.g., a COIA type II Cas protein whose amino acid sequence comprises SEQ ID NO: 55, 56, or 57). In some embodiments, all of the domains are derived from a DJQA type II Cas protein (e.g., a DJQA type II Cas protein whose amino acid sequence comprises SEQ ID NO: 61, 62, or 63). In some embodiments, all of the domains are derived from a DWET type II Cas protein (e.g., a DWET type II Cas protein whose amino acid sequence comprises SEQ ID NO: 67, 68, or 69). In other embodiments, one or more domains (e.g., a domain), e.g., the PID domain, are derived from another type II Cas protein.

[0176] In addition, one or more amino acid substitutions can be introduced into one or more domains to modify the properties of the resulting nuclease in terms of editing activity, targeting specificity, or PAM recognition specificity. For example, one or more amino acid substitutions can be introduced to confer nickase activity. Exemplary amino acid substitutions in SaCas9 that confer nickase activity are the D10A substitution in the RuvC domain and the N580A substitution in the HNH domain. Combining both the D10A and N580A substitutions in SaCas9 results in a catalytically inactive nuclease. Corresponding substitutions can be introduced into the type II Cas nuclease of the present disclosure to obtain a nickase and a catalytically inactive Cas protein. For example, an AEQH type II Cas protein can include a D10A substitution (corresponding to D10A in SaCas9) or an N627A substitution (corresponding to N580A in SaCas9) to confer a nickase, or a D10A and N627A substitution to confer a catalytically inactive Cas protein, where the positions of the D13A and N589A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 2. Positions corresponding to D10 and N580 in SaCas9 in type II Cas proteins of the present disclosure are shown in Table 4. An exemplary amino acid substitution in CjCas9 that confers nickase activity is the H559A substitution. Corresponding substitutions can be introduced into type II Cas nucleases of the present disclosure to obtain nickases. A substitution at the position corresponding to H559A in CjCas9 can be combined with a substitution at the position corresponding to D10 in SaCas9 to generate a catalytically inactive type II Cas protein. The disclosed nickases and catalytically inactive Type II Cas proteins can be used, for example, in base editors that include cytosine deaminase or adenosine deaminase fusion partners. Catalytically inactive Type II Cas proteins can also be used, for example, as fusion partners for transcriptional activators or repressors.

[0177] [Table 15]

[0178] Guide RNA The present disclosure provides gRNA molecules that can be used with the type II Cas proteins of the present disclosure to edit genomic DNA, e.g., mammalian DNA, e.g., human DNA. The gRNAs of the present disclosure typically include a spacer, 15-30 nucleotides in length, that can be located 5' of the crRNA scaffold to form a complete crRNA. The crRNA can be used with a tracrRNA to induce cleavage of a target genomic sequence.

[0179] An exemplary crRNA scaffold sequence that can be used for AEQH Type II Cas gRNA includes GUUUUAGUACUCUGUUGGAUAUUGAUAAACUUACAC (SEQ ID NO: 73), and an exemplary tracrRNA sequence that can be used for AEQH Type II Cas gRNA includes UGUGAGUUUAUCAAUAUCCAACAAUAGUUCUAAGAUAAGGCUAUUUAUGCCGUAGGGUAUGGCGGUAUCCCGUUAAUCCGCCUUUAAGCCAUUGCUUUGCAAUGGCUUA (SEQ ID NO: 74).

[0180] An exemplary crRNA scaffold sequence that can be used for AAOF type II Cas gRNA includes GUUGUAGUUCCCUGGUAGUUCUUGGUAUGGUAUAAU (SEQ ID NO: 75), and an exemplary tracrRNA sequence that can be used for AAOF type II Cas gRNA includes UUAUACCAUACCAAGAACUAUGCAGGUUACUAUGAUAAGGUAGUACACCGCAGAGCUCUAACGCCUCGCGUAAGCGGGGCGUUAUCUCU (SEQ ID NO: 76).

[0181] An exemplary crRNA scaffold sequence that can be used for ACEE type II Cas gRNA includes AUUGUAGUUCCCUAAUUUUUCUUGGUAUGUUAUAAU (SEQ ID NO: 77), and an exemplary tracrRNA sequence that can be used for ACEE type II Cas gRNA includes UUAUAACAUACCAAGAACAAUUAGGUUACUACAAAAAGGUAGAAAACCGAAAAGCUCUAACGGCUCCUUUUUGGAGCCGUUAUCUUUUU (SEQ ID NO: 78).

[0182] An exemplary crRNA scaffold sequence that can be used for AQSL Type II Cas gRNAs includes GUUGUAGUUCCCUGGUAGUUCUUGGUAUGGUAUAAU (SEQ ID NO: 79), and an exemplary tracrRNA sequence that can be used for AQSL Type II Cas gRNAs includes UUAUACCAUACCAAGGACUAUGCAGGUUACUAUGAUAAGGUAGUACACCGCAGAGCACUGACGCCCCGCUUUUGCGGGGCGUUAUCUCU (SEQ ID NO: 80).

[0183] An exemplary crRNA scaffold sequence that can be used for an ASWC Type II Cas gRNA includes GUUCUGGCCUAAGCUCAUUUCCUAACUGAUACAAUC (SEQ ID NO: 81), and an exemplary tracrRNA sequence that can be used for an ASWC Type II Cas gRNA includes UCAGUUAGGAAAUGGGCUUUCUCCACUAACAAGCUGAGAGAUGCACAAGAUGCGGGGUCGCUAUAUGCGACCCUUUUUCGUAUC (SEQ ID NO: 82).

[0184] An exemplary crRNA scaffold sequence that can be used for AVFG Type II Cas gRNA includes GUUAUAGUUCCUAGUAAAUUCUCGAUAUGCUAUAAU (SEQ ID NO: 83), and an exemplary tracrRNA sequence that can be used for AVFG Type II Cas gRNA includes UAUAGCAUAUCGAGAGUUUAACUAGUUGCUAUAACAAGGCAAUAAGCCGUAAAGUAUCCCCUGUACUCAUUUCUUGAGUGUAGGGGUAUCUUU (SEQ ID NO: 84).

[0185] An exemplary crRNA scaffold sequence that can be used for an AWIT type II Cas gRNA includes GUCAUAGUUCCCUAAUAGCUCUUGGUAUGGUAUAAU (SEQ ID NO: 85), and an exemplary tracrRNA sequence that can be used for an AWIT type II Cas gRNA includes UUAUACCAUACCAAGAACUAUUAUGGUUGCUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGUCUUAUGACAGGGCGUCAUCUUU (SEQ ID NO: 86).

[0186] An exemplary crRNA scaffold sequence that can be used for AWMF Type II Cas gRNA includes GUCAUAGUUCCCUAAUAGCUCUUGGUAUGGUAUAAU (SEQ ID NO: 87), and an exemplary tracrRNA sequence that can be used for AWMF Type II Cas gRNA includes UUAUACCAUACCAAGAACUAUUAUGGUUGCUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGUCUUAUGACAGGGCGUCAUCUUU (SEQ ID NO: 88).

[0187] An exemplary crRNA scaffold sequence that can be used for a BUMO type II Cas gRNA includes GUUGUAGUUCCCUGAUGAUUCUUGGUAUGGUAUAAU (SEQ ID NO: 89), and an exemplary tracrRNA sequence that can be used for a BUMO type II Cas gRNA includes UUAUACCAUACCAAGGAUUAUCGGGUUACUAUGAUAAGGUAGUACACCGAAAAGCUCUAACGCUCUGUCGCUUUUGACAGAGCGUUAUCUUUU (SEQ ID NO: 90).

[0188] An exemplary crRNA scaffold sequence that can be used for a COIA type II Cas gRNA includes GUCAUAGUUCCCUAAUAGCUCUUGGUAUGGUAUAAU (SEQ ID NO: 91), and an exemplary tracrRNA sequence that can be used for a COIA type II Cas gRNA includes UUAUACCAUACCAAGAACUAUUAUGGUUGCUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGUCUUAUGACAGGGCGUCAUCUUU (SEQ ID NO: 92).

[0189] An exemplary crRNA scaffold sequence that can be used for DJQA type II Cas gRNA includes GUCAUAGUUCCCUAAUAGCUCUUGGUAUGGUAUAAU (SEQ ID NO: 93), and an exemplary tracrRNA sequence that can be used for DJQA type II Cas gRNA includes UUAUACCAUACCAAGAACUAUUAUGGUUACUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGCCGUUUGGCAGGGCGUCAUCUUU (SEQ ID NO: 94).

[0190] An exemplary crRNA scaffold sequence that can be used for DWET Type II Cas gRNA includes GUUGUAGUUCCCUGGUAGUUCUUGGUAUGGUAUAAU (SEQ ID NO: 95), and an exemplary tracrRNA sequence that can be used for DWET Type II Cas gRNA includes UUAUACCAUACCAAGAACUAUGCAGGUUACUAUGAUAAGGUAGUAUACCGCAGAGCUCUAACGCCCCGCGUAAGCGGGGCGUUAUCUCU (SEQ ID NO: 96).

[0191] The gRNAs of the present disclosure, in some embodiments, are single guide RNAs (sgRNAs), which typically include a spacer at the 5' end of the molecule and a 3'sgRNA scaffold. Alternatively, the gRNA can include separate crRNA and tracrRNA molecules.

[0192] Further features of exemplary gRNA spacer sequences are described in Section 6.3.1, and further features of exemplary 3's gRNA scaffolds are described in Section 6.3.2.

[0193] Spacer The spacer sequence is partially or completely complementary to the target sequence found in genomic DNA sequence, for example, human genomic DNA sequence.For example, the spacer sequence can be partially or completely complementary to the nucleotide sequence in the gene that has a mutation that causes disease.The spacer that is partially complementary to the target sequence can have, for example, one, two or three mismatches with the target sequence.

[0194] A gRNA of the present disclosure can include a spacer that is 15-30 nucleotides in length (e.g., 15-25, 16-24, 17-23, 18-22, 19-21, 18-30, 20-28, 22-26, or 23-25 ​​nucleotides in length). In some embodiments, the spacer is 15 nucleotides in length. In other embodiments, the spacer is 16 nucleotides in length. In other embodiments, the spacer is 17 nucleotides in length. In other embodiments, the spacer is 18 nucleotides in length. In other embodiments, the spacer is 19 nucleotides in length. In other embodiments, the spacer is 20 nucleotides in length. In other embodiments, the spacer is 21 nucleotides in length. In other embodiments, the spacer is 22 nucleotides in length. In other embodiments, the spacer is 23 nucleotides in length. In other embodiments, the spacer is 24 nucleotides in length. In other embodiments, the spacer is 25 nucleotides in length. In other embodiments, the spacer is 26 nucleotides in length. In other embodiments, the spacer is 27 nucleotides in length. In other embodiments, the spacer is 28 nucleotides in length. In other embodiments, the spacer is 29 nucleotides in length. In other embodiments, the spacer is 30 nucleotides in length.

[0195] Type II Cas endonucleases require a specific sequence, called a protospacer adjacent motif (PAM), downstream (e.g., immediately downstream) of the target sequence on the non-target strand. Therefore, a spacer sequence for targeting a gene of interest can be identified by scanning the gene for PAM sequences recognized by Type II Cas proteins. Exemplary PAM sequences for Type II Cas proteins are shown in Tables 5A and 5B.

[0196] [Table 16]

[0197] [Table 17]

[0198] Example 3 describes exemplary sequences that can be used to target RHO genomic sequences. Example 4 describes exemplary sequences that can be used to target TRAC, B2M, and PD1 genomic sequences. In some embodiments, a gRNA of the present disclosure comprises a spacer sequence that targets RHO. In some embodiments, a gRNA of the present disclosure comprises a spacer sequence that targets TRAC. In some embodiments, a gRNA of the present disclosure comprises a spacer sequence that targets B2M. In some embodiments, a gRNA of the present disclosure comprises a spacer sequence that targets PD1.

[0199] Additional exemplary spacer sequences that can be used in the gRNAs of the present disclosure are listed in Tables 6A, 6B, and 6C.

[0200] [Table 18-1]

[0201] [Table 18-2]

[0202] [Table 18-3]

[0203] [Table 19]

[0204] [Table 20-1]

[0205] [Table 20-2]

[0206] [Table 20-3]

[0207] The RHO spacer sequences in Table 6A are useful for targeting the RHO gene near the rs7984 SNP located in the 5' untranslated region (UTR) of the RHO gene. Allele-specific targeting can be achieved by using a gRNA that targets a SNP variant found in a cell or subject. For example, if a cell or subject has an "A" at the rs7984 SNP position, a guide in Table 6 with the name "7984A" can be used; if a cell or subject has a "G" at the rs7984 SNP position, a guide with the name "7984G" can be used. Such a guide can be used with a guide RNA targeting RHO intron 1 (e.g., with a spacer shown in Table 6B) to, for example, knock out expression of a mutant protein. Allele-specific targeting of RHO is further described in Example 3. Exemplary combinations of guides include a first guide RNA with a spacer whose sequence is selected from SEQ ID NOs: 326-334 and 336-346, and a second guide RNA with a spacer whose sequence is selected from SEQ ID NOs: 402, 403, and 406.

[0208] In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 15 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 16 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 17 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 18 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 19 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 20 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 21 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 22 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 23 or more contiguous nucleotides from a sequence shown in Table 6A. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 24 contiguous nucleotides from a sequence shown in Table 6A.

[0209] In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 15 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 16 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 17 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 18 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 19 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 20 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 21 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 22 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 23 or more contiguous nucleotides from a sequence shown in Table 6B. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 24 or more contiguous nucleotides from a sequence shown in Table 6B.

[0210] In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 15 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 16 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 17 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 18 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 19 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 20 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, and the nucleotide sequence of the spacer comprises 21 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 22 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 23 or more contiguous nucleotides from a sequence shown in Table 6C. In some embodiments, a gRNA of the present disclosure has a spacer, wherein the nucleotide sequence of the spacer comprises 24 contiguous nucleotides from a sequence shown in Table 6C.

[0211] 6.3.2.sgRNA molecules The gRNA of the present disclosure may be a single guide RNA (sgRNA) molecule. The sgRNA may comprise, from 5' to 3', an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single-molecule guide linker, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. The optional tracrRNA extension may comprise an element that confers additional functionality (e.g., stability) to the guide RNA. The single-molecule guide linker may link the minimal CRISPR repeat and the minimal tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension may comprise one or more hairpins.

[0212] The sgRNA can include a variable length spacer sequence (e.g., 15-30 nucleotides) at the 5' end of the sgRNA sequence and a 3' sgRNA segment.

[0213] Type II Cas gRNAs typically comprise a repeat-antirepeat duplex and / or one or more stem loops generated by the secondary structure of the gRNA. The length of the repeat-antirepeat duplex and / or one or more stem loops can be modified to modulate (e.g., increase) the editing efficiency of the Type II Cas nuclease and / or to reduce the size of the guide RNA for easier vectorization in situations where the cargo size of the vector (e.g., AAV vector) is limited.

[0214] For example, the repeat-antirepeat duplex (fused in the sgRNA via a synthetic linker, resulting in an additional stem-loop within the structure) can generally be trimmed to various lengths without adversely affecting nuclease function, and in some cases even without increasing enzymatic activity. If bulges are present within this duplex, they should generally be retained in the final guide RNA sequence.

[0215] Further optimization of structure can be achieved by introducing targeted base changes into the stem of gRNA to improve its stability and folding.Such base changes preferably correspond to the introduction of G:C pairs, which are known to form the strongest Watson-Crick pair.To clarify, these substitutions can consist of the introduction of G or C at specific positions of stem, and the complementary substitution at other positions of gRNA sequence, which are predicted to form base pairs with the former, for example, according to available bioinformatics tools for RNA folding, such as UNAfold or RNAfold.

[0216] Stem-loop trimming can also be used to stabilize desired secondary structures by removing portions of the guide RNA that generate unwanted secondary structures through annealing with other regions of the RNA molecule.

[0217] Exemplary 3' sgRNA scaffold sequences for Type IIA Cas sgRNAs are shown in Table 7A. Exemplary 3' sgRNA scaffold sequences for Type IIC Cas sgRNAs are shown in Table 7B.

[0218] [Table 21]

[0219] [Table 22-1]

[0220] [Table 22-2]

[0221] An sgRNA (e.g., for use with an AAOF type II Cas protein, an ACEE type II Cas protein, an AEQH type II Cas protein, an AQSL type II Cas protein, an ASWC type II Cas protein, an AVFG type II Cas protein, an AWIT type II Cas protein, an AWMF type II Cas protein, a BUMO type II Cas protein, a COIA type II Cas protein, a DJQA type II Cas protein, and / or a DWET type II Cas protein) may not include a uracil base at the 3' end of the sgRNA sequence. However, an sgRNA typically includes one or more uracil bases at the 3' end of the sgRNA sequence, e.g., to facilitate correct sgRNA folding. For example, an sgRNA can include one uracil (U) at the 3' end of the sgRNA sequence. An sgRNA can include two uracils (UU) at the 3' end of the sgRNA sequence. The sgRNA can contain three uracils (UUU) at the 3' end of the sgRNA sequence. The sgRNA can contain four uracils (UUUU) at the 3' end of the sgRNA sequence. The sgRNA can contain five uracils (UUUUU) at the 3' end of the sgRNA sequence. The sgRNA can contain six uracils (UUUUUU) at the 3' end of the sgRNA sequence. The sgRNA can contain seven uracils (UUUUUUU) at the 3' end of the sgRNA sequence. The sgRNA can contain eight uracils (UUUUUUUU) at the 3' end of the sgRNA sequence. Uracil chains of various lengths can be added to the 3' end of the sgRNA as terminators. Thus, for example, the 3'sgRNA sequences shown in Table 7A and Table 7B can be modified by adding (or removing) one or more uracils to the end of the sequence.

[0222] In some embodiments, the sgRNA scaffold for use with an AEQH Type II Cas protein has the sequence

[0223] [ka] In some embodiments, an sgRNA scaffold for use with an AEQH type II Cas protein comprises the sequence GUUUUAGUACUCUGUGAAAACAAUAGUUCUAAGAUAAGGCUAUUUAUGCCGUAGGGUAUGGCGGUAUCCCGUUAAUCCGCCUUUAAGCCAUUGCUUUGCAAUGGCUUAUUUUU (SEQ ID NO: 122).

[0224] In some embodiments, the sgRNA scaffold for use with an AAOF Type II Cas protein has the sequence

[0225] [ka] In some embodiments, an sgRNA scaffold for use with an AAOF Type II Cas protein comprises the sequence GUUGUAGUUCCCUGGUAGGAAACUAUGCAGGUUACUAUGAUAAGGUAGUACACCGCAGAGCUCUAACGCCUCGCGUAAGCGGGGCGUUAUCUCUUUUUU (SEQ ID NO: 124).

[0226] In some embodiments, the sgRNA scaffold for use with an ACEE type II Cas protein has the sequence

[0227] [ka] In some embodiments, the sgRNA scaffold for use with an ACEE type II Cas protein comprises the sequence AUUGUAGUUCCCUGAAAAGGUUACUACAAAAAGGUAGAAAACCGAAAAGCUCUAACGGCUCCGAAAGGAGCCGUUAUCUUUUUU (SEQ ID NO: 126).

[0228] In some embodiments, the sgRNA scaffold for use with an AQSL Type II Cas protein has the sequence

[0229] [ka] In some embodiments, an sgRNA scaffold for use with an AQSL type II Cas protein comprises the sequence GUUGUAGUUCCCUGGUAGGAAACUAUGCAGGUUACUAUGAUAAGGUAGUACACCGCAGAGCACUGACGCCCCGCUUUUGCGGGGCGUUAUCUCUUUUUU (SEQ ID NO: 128).

[0230] In some embodiments, the sgRNA scaffold for use with an ASWC type II Cas protein has the sequence

[0231] [ka] In some embodiments, an sgRNA scaffold for use with an ASWC type II Cas protein comprises the sequence GUUCUGGCCUAAGGAAACUUUCUCCACUAACAAGCUGAGAGAUGCACAAGAUGCGGGGUCGCUAUAUGCGACCCUUAUUCGUAUCCAAAUUUUUU (SEQ ID NO: 130).

[0232] In some embodiments, the sgRNA scaffold for use with an AVFG Type II Cas protein has the sequence

[0233] [ka] In some embodiments, the sgRNA scaffold for use with an AVFG type II Cas protein comprises the sequence Contains GUUAUAGUUCCUAGUAAGAAAUUAACUAGUUGCUAUAACAAGGCAAUAAGCCGUAAAGUAUCCCCUGUACUCAUUUCUUGAGUGUAGGGGUAUCUUUUUU (SEQ ID NO: 132).

[0234] In some embodiments, the sgRNA scaffold for use with an AWIT type II Cas protein has the sequence

[0235] [ka] In some embodiments, an sgRNA scaffold for use with an AWIT type II Cas protein comprises the sequence GUCAUAGUUCCCUAAGAAAUUAUGGUUGCUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGUCUUAUGACAGGGCGUCAUCUUUUUU (SEQ ID NO: 134).

[0236] In some embodiments, an sgRNA scaffold for use with an AWMF Type II Cas protein has the sequence

[0237] [ka] In some embodiments, an sgRNA scaffold for use with an AWMF Type II Cas protein comprises the sequence GUCAUAGUUCCCUAAGAAAUUAUGGUUGCUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGUCUUAUGACAGGGCGUCAUCUUUUUU (SEQ ID NO: 136).

[0238] In some embodiments, the sgRNA scaffold for use with a BUMO type II Cas protein has the sequence

[0239] [ka] In some embodiments, an sgRNA scaffold for use with a BUMO type II Cas protein comprises the sequence GUUGUAGUUCCCUGGAAACGGGUUACUAUGAUAAGGUAGUACACCGAAAAGCUCUAACGCUCUGUCGCUUUUGACAGAGCGUUAUCUUUUUU (SEQ ID NO: 138).

[0240] In some embodiments, the sgRNA scaffold for use with a COIA type II Cas protein has the sequence

[0241] [ka] In some embodiments, an sgRNA scaffold for use with a COIA type II Cas protein comprises the sequence GUCAUAGUUCCCUAAGAAAUUAUGGUUGCUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGUCUUAUGACAGGGCGUCAUCUUUUUU (SEQ ID NO: 140).

[0242] In some embodiments, the sgRNA scaffold for use with a DJQA type II Cas protein has the sequence

[0243] [ka] In some embodiments, an sgRNA scaffold for use with a DJQA type II Cas protein comprises the sequence GUCAUAGUUCCCUAAGAAAUUAUGGUUACUAUGAUAAGGUCAUAGGACCGUAAAGCUCUGACGCCCUGCCGUUUGGCAGGGCGUCAUCUUUUUU (SEQ ID NO: 142).

[0244] In some embodiments, the sgRNA scaffold for use with a DWET Type II Cas protein has the sequence

[0245] [ka] In some embodiments, an sgRNA scaffold for use with a DWET type II Cas protein comprises the sequence GUUGUAGUUCCCUGGUAGGAAACUAUGCAGGUUACUAUGAUAAGGUAGUAUACCGCAGAGCUCUAACGCCCCGCGUAAGCGGGGCGUUAUCUCUUUUUU (SEQ ID NO: 144).

[0246] 6.3.3. Modified gRNA Molecules Guide RNAs can be easily synthesized by chemical means, as described in the art, which allows for the easy incorporation of several modifications. The disclosed gRNA (e.g., sgRNA) molecules can be unmodified or can include any one or more of a range of chemical modifications.

[0247] Although chemical synthesis procedures are constantly evolving, purification of such RNAs by procedures such as high-performance liquid chromatography (HPLC, which avoids the use of gels such as PAGE) tends to become more difficult as the length of the polynucleotide increases significantly beyond about 100 nucleotides. One approach that can be used to create longer chemically modified RNAs is to generate two or more molecules ligated together. Much longer RNAs, such as those encoding type II Cas endonucleases, are more easily produced enzymatically. Although the types of modifications available for use in enzymatically generated RNAs are fewer, as described herein and in the art, there are still modifications that can be used to enhance stability, reduce the likelihood or severity of innate immune responses, and / or enhance other attributes.

[0248] Examples of various types of modifications, particularly those frequently used with smaller chemically synthesized RNAs, include one or more nucleotides modified at the 2' position of the sugar, such as 2'-O-alkyl, 2'-O-alkyl-O-alkyl, or 2'-fluoro modified nucleotides. In some examples, RNA modifications can include 2'-fluoro, 2'-amino, or 2'-O-methyl modifications on the ribose of pyrimidines, abasic residues, or inverted bases at the 3' end of the RNA. Such modifications can be routinely incorporated into oligonucleotides, and these oligonucleotides have been shown to have higher Tm (and therefore higher target binding affinity) for a given target than 2'-deoxyoligonucleotides.

[0249] Some nucleotide and nucleoside modifications have been shown to make the oligonucleotides they incorporate more resistant to nuclease digestion than natural oligonucleotides; these modified oligonucleotides remain unchanged for a longer period of time than unmodified oligonucleotides.Specific examples of modified oligonucleotides include those that contain modified backbones, such as phosphorothioates, phosphotriesters, methylphosphonates, short-chain alkyl or cycloalkyl intersugar bonds, or short-chain heteroatom or heterocyclic intersugar bonds. Some oligonucleotides include oligonucleotides with phosphorothioate backbones and oligonucleotides with heteroatom backbones, particularly CH2-NH-O-CH2, CH, ~N(CH3)-O-CH2 (known as the methylene(methylimino) or MMI backbone), CH2-ON(CH3)-CH2, CH2-N(CH3)-N(CH3)-CH2, and ON(CH3)-CH2-CH2 backbones, where the natural phosphodiester backbone is represented as OPO-CH; amide backbones (see De Mesmaeker et al. 1995, Ace. Chem. Res., 28:366-374); morpholino backbone structures (see U.S. Pat. No. 5,034,506); peptide nucleic acid (PNA) backbones, in which the phosphodiester backbone of the oligonucleotide is replaced by a polyamide backbone and the nucleotides are attached directly or indirectly to the aza nitrogen atoms of the polyamide backbone (Nielsen et al., 1991, Science 254:1497).Phosphorus-containing linkages include, but are not limited to, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, including 3' alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates, including 3'-aminophosphoramidates and aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, and boranophosphates with normal 3'-5' linkages, 2'-5' linked analogs thereof, and those with reverse polarity, in which adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'; U.S. Pat. Nos. 3,687,808; 4,469,866; Specification No. 3; Specification No. 4,476,301; Specification No. 5,023,243; Specification No. 5,177,196; Specification No. 5,188,897; Specification No. 5,264,423; Specification No. 5,276,0 Specification No. 19; Specification No. 5,278,302; Specification No. 5,286,717; Specification No. 5,321,131; Specification No. 5,399,676; Specification No. 5,405,939; Specification No. 5,453,4 See US Pat. Nos. 96; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050.

[0250] Morpholino-based oligomeric compounds are described in Braasch and David Corey, 2002, Biochemistry, 41(14):4503-4510; Genesis, Volume 30, Issue 3, (2001); Heasman, 2002, Dev. Biol., 243: 209-214; Nasevicius et al., 2000, Nat. Genet., 26:216-220; Lacerra et al., 2000, Proc. Natl. Acad. Sci., 97: 9591-9596; and U.S. Pat. No. 5,034,506.

[0251] Cyclohexenyl nucleic acid oligonucleotide mimetics are described in Wang et al., 2000, J. Am. Chem. Soc., 122: 8595-8602.

[0252] Modified oligonucleotide backbones that do not contain an internal phosphorus atom have backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages. These include morpholino linkages (formed in part from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; those with amide backbones; and others with mixed N, O, S, and CH moieties; U.S. Pat. Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; and 5,235,033. ; Specification No. 5,264,562; Specification No. 5,264,564; Specification No. 5,405,938; Specification No. 5,434,257; Specification No. 5,466,677; Specification No. 5,4 Specification No. 70,967; Specification No. 5,489,677; Specification No. 5,541,307; Specification No. 5,561,225; Specification No. 5,596,086; Specification No. 5,602,240 See US Pat. Nos. 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439.

[0253] One or more substituted sugar moieties can also be included, for example, one of the following at the 2' position: OH, SH, SCH, F, OCN, OCH, OCH0(CH)CH, O(CH)NH, or O(CH)CH, where n is 1 to about 10; C to C10 Lower alkyl, alkoxyalkoxy, substituted lower alkyl, alkaryl, or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S-, or bi-alkyl; O-, S-, or N-alkenyl; SOCH3; SO2CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; RNA cleaving group; reporter group; intercalator; group for improving the pharmacokinetic properties of oligonucleotides; or group for improving the pharmacodynamic properties of oligonucleotides, and other substituents with similar properties. In some embodiments, the modification includes 2'-methoxyethoxy (2'-O-CH2CHOCH3, also known as 2'-O-(2-methoxyethyl)) (Martin et al., 1995, Helv. Chim. Acta, 78, 486). Other modifications include 2'-methoxy (2'-O-CH), 2'-propoxy (2'-OCHCHCH), and 2'-fluoro (2'-F). Similar modifications can also be made at other positions on the oligonucleotide, particularly the 3' position of the sugar on the 3'-terminal nucleotide and the 5' position of the 5'-terminal nucleotide. Oligonucleotides can also have sugar mimetics, such as cyclobutyl, in place of the pentofuranosyl group.

[0254] In some cases, both the sugar and the internucleoside linkage (in the backbone) of a nucleotide unit can be replaced with novel groups. The base unit can be maintained for hybridization with an appropriate nucleic acid target compound. One such oligomeric compound, an oligonucleotide mimic that has been shown to have excellent hybridization properties, is called peptide nucleic acid (PNA). In PNA compounds, the sugar backbone of an oligonucleotide can be replaced with an amide-containing backbone (e.g., an aminoethylglycine backbone). The nucleobases can be retained and directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone. Representative U.S. patents that teach the preparation of PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262. Further teachings on PNA compounds can be found in Nielsen et al., 1991, Science, 254: 1497-1500.

[0255] RNAs, such as guide RNAs, can also additionally or alternatively include nucleobase (often referred to in the art simply as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include nucleobases that are found only rarely or occasionally in natural nucleic acids, such as hypoxanthine, 6-methyladenine, 5-Me pyrimidines, particularly 5-methylcytosine (also called 5-methyl-2'deoxycytosine, often referred to in the art as 5-Me-C), 5-hydroxymethylcytosine (HMC), glycosyl HMC, and gentobiosyl HMC, as well as synthetic nucleobases such as 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolylalkyl)adenine, 2-(aminoalkylamino)adenine or other heterosubstituted alkyl adenines, 2-thiouracil, 2-thiothymine, 5-bromouracil, 5-hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N6(6-aminohexyl)adenine, and 2,6-diaminopurine. Komberg, A., DNA Replication, WH Freeman & Co., San Francisco, pp. 75-77 (1980); Gebeyehu et al., Nucl. Acids Res. 15:4513 (1997). "Universal" bases known in the art (e.g., inosine) may also be included. 5-Me-C substitution has been described to increase the stability of nucleic acid duplexes by approximately 0.6 to 1.2°C (Sanghvi, YS, in Crooke, ST and Lebleu, B., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278), and these are examples of base substitutions.

[0256] Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azouracil, cytosine and and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.

[0257] Further examples of nucleobases include those disclosed in U.S. Pat. No. 3,687,808, those disclosed in 'The Concise Encyclopedia of Polymer Science and Engineering', pp. 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, those disclosed in Englisch et al., Angewandle Chemie, International Edition, 1991, 30, p. 613, and those disclosed in Sanghvi, YS, Chapter 15, Antisense Research and Applications, pp. 289-302, Crooke, ST and Lebleu, B. ea., CRC Press, 1993. Some of these nucleobases may be useful for increasing the binding affinity of the oligomeric compounds of the present invention. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-Methylcytosine substitutions have been shown to increase nucleic acid duplex stability by approximately 0.6-1.2°C (Sanghvi, YS, Crooke, ST, and Lebleu, B., eds., 'Antisense Research and Applications', CRC Press, Boca Raton, 1993, 276-278), and these are some of the base substitutions, especially when combined with 2'-O-methoxyethyl sugar modifications.Modified nucleobases are disclosed in U.S. Patent Nos. 3,687,808, as well as U.S. Patent Nos. 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5 ,525,711; 5,552,540; 5,587,469; 5,596,091; 5,614,617; 5,681,941; 5,750,692; 5,763,588; 5,830,653; 6,005,096; and U.S. Patent Application Publication No. 2003 / 0158403.

[0258] Thus, modified gRNAs can include, for example, one or more unnatural sugars, internucleotide linkages, and / or bases. Not all positions in a given gRNA need be uniformly modified; in fact, more than one of the foregoing modifications can be incorporated in a single oligonucleotide, or even in a single nucleoside within an oligonucleotide.

[0259] The guide RNA and / or mRNA (or DNA) encoding the endonuclease can be chemically linked to one or more moieties or conjugates that enhance the activity, cellular distribution, or cellular uptake of the oligonucleotide. Such moieties include lipid moieties such as cholesterol moieties (Letsinger et al. 1989, Proc. Natl. Acad. Sci. USA, 86: 6553-6556); cholic acid (Manoharan et al., 1994, Bioorg. Med. Chem. Let., 4: 1053-1060); thioethers, e.g., hexyl-S-tritylthiol (Manoharan et al., 1992, Ann. NY Acad. Sci., 660: 306-309; Manoharan et al., 1993, Bioorg. Med. Chem. Let., 3: 2765-2770); thiocholesterols (Oberhauser et al., 1992, Nucl. Acids Res., 20: 533-538); aliphatic chains, such as dodecanediol or undecyl residues (Kabanov et al, 1990, FEBS Lett., 259: 327-330; Svinarchuk et al, 1993, Biochimie, 75: 49-54); phospholipids, such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., 1995, Tetrahedron Lett., 36: 3651-3654; and Shea et al, 1990, Nucl. Acids Res., 18: 3777-3783); polyamine or polyethylene glycol chains (Mancharan et al, 1995, Nucleosides & Nucleotides, 14: 969-973); adamantane acetic acid (Manoharan et al, 1995, Tetrahedron Lett., 36: 3651-3654); palmityl moiety (Mishra et al., 1995, Biochim. Biophys.Acta, 1264: 229-237); or an octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al, 1996, J. Pharmacol. Exp. Ther., 277: 923-937). Also, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,046; Specification No. 77; Specification No. 5,486,603; Specification No. 5,512,439; Specification No. 5,578,718; Specification No. 5,608,046; Specification No. 4,587,044; Specification No. 4,605,735; Specification No. 4,667,025; Specification No. 4,762,779; Specification No. 4,789,737; Specification No. 4,824,941; Specification No. 4,835,263; Specification No. 4,876,335; Specification No. 4,904,582; Specification No. 4,958,013; Specification No. 5,082,8 Specification No. 30; Specification No. 5,112,963; Specification No. 5,214,136; Specification No. 5,082,830; Specification No. 5,112,963; Specification No. 5,214,136; Specification No. 5,245,022; Specification No. 5,254,469; Specification No. 5,258,506; Specification No. 5,262,536; Specification No. 5,272,250; Specification No. 5,292,873; Specification No. 5,317,098; Specification No. 5,371,241; Specification No. 5,391,723; Specification No. 5,416, See also US Pat. Nos. 203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0260] Sugar and other moieties can be used to target proteins and complexes containing nucleotides, such as cationic polysomes and liposomes, to specific sites.For example, hepatocyte-directed import can be mediated through asialoglycoprotein receptor (ASGPR); see, for example, Hu, et al., 2014, Protein Pept Lett. 21(10):1025-30.Other systems known in the art and frequently developed can be used to target the biomolecules and / or their complexes useful in the present invention to specific target cells of interest.

[0261] The targeting moiety or conjugate can include a conjugate group covalently bound to a functional group such as a primary or secondary hydroxyl group. Conjugate groups of the present disclosure include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacodynamic properties of oligomers, and groups that enhance the pharmacokinetic properties of oligomers. Typical conjugate groups include cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. In the context of the present disclosure, groups that enhance the pharmacodynamic properties include groups that improve uptake, enhance resistance to degradation, and / or strengthen sequence-specific hybridization with target nucleic acids. In the context of the present disclosure, groups that enhance the pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of the compounds of the present disclosure. Representative conjugate groups are disclosed in International Patent Application Publication No. 1993007883 and U.S. Patent No. 6,287,860. Conjugate moieties include, but are not limited to, lipid moieties such as cholesterol moieties, cholic acid, thioethers such as hexyl-5-tritylthiol, thiocholesterol, aliphatic chains such as dodecanediol or undecyl residues, phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate, polyamines or polyethylene glycol chains, or adamantane acetic acid, palmityl moieties, or octadecylamine or hexylamino-carbonyl-oxycholesterol moieties.See, e.g., U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077 specification; Specification No. 4,762,779; Specification No. 4,789,737; Specification No. 4,824,941; Specification No. 4,835,263; Specification No. 4,876,335; Specification No. 4,904,582; Specification No. 4,958,013; Specification No. 5,082, Specification No. 830; Specification No. 5,112,963; Specification No. 5,214,136; Specification No. 5,082,830; Specification No. 5,112,963; Specification No. 5,214,136; Specification No. 5,245,022; Specification No. 5,254,469 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241; 5,391,723; 5,4 See Nos. 16,203, 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0262] As described herein, a wide variety of modifications have been developed and applied to enhance RNA stability, mitigate innate immune responses, and / or achieve other benefits that may be useful in connection with the introduction of polynucleotides into human cells; see, e.g., Whitehead KA et al., 2011, Annual Review of Chemical and Biomolecular Engineering, 2: 77-96; Gaglione and Messere, 2010, Mini Rev Med Chem, 10(7):578-95; Chernolovskaya et al., 2010, Curr Opin Mol Ther., 12(2): 158-67; Deleaviey et al., 2009, Curr Protoc Nucleic Acid Chem Chapter 16:Unit 16.3; Behlke, 2008, Oligonucleotides 18(4):305-19; Fucini et al., 2012, Nucleic Acid Ther 22(3): 205-210; see review by Bremsen et al, 2012, Front Genet 3: 154.

[0263] 6.4.System The present disclosure provides a system comprising a Type II Cas protein of the present disclosure (e.g., as described in Section 6.2) and a means for targeting the Type II Cas protein to a target genomic sequence. The means for targeting the Type II Cas protein to a target genomic sequence can be a guide RNA (gRNA) (e.g., as described in Section 6.3).

[0264] The present disclosure also provides systems comprising a Type II Cas protein (e.g., as described in Section 6.2) and a gRNA (e.g., as described in Section 6.3) of the present disclosure. The systems can comprise a ribonucleoprotein particle (RNP) in which the Type II Cas protein is complexed with a gRNA, e.g., an sgRNA or separate crRNA and tracrRNA. The systems of the present disclosure can, in some embodiments, further comprise genomic DNA complexed with the Type II Cas protein and the gRNA. Thus, the present disclosure provides systems comprising a Type II Cas protein, genomic DNA, and a gRNA, all complexed together.

[0265] The systems of the present disclosure can be present intracellularly (whether the cell is in vivo, ex vivo, or in vitro) or extracellularly (eg, within or outside a particle).

[0266] 6.5. Nucleic acids The present disclosure provides nucleic acids (e.g., DNA or RNA) encoding type II Cas proteins (e.g., AAOF type II Cas protein, ACEE type II Cas protein, AEQH type II Cas protein, AQSL type II Cas protein, ASWC type II Cas protein, AVFG type II Cas protein, AWIT type II Cas protein, AWMF type II Cas protein, BUMO type II Cas protein, COIA type II Cas protein, DJQA type II Cas protein, and DWET type II Cas protein), nucleic acids encoding gRNAs of the disclosure (e.g., a single gRNA or a combination of gRNAs), nucleic acids encoding both a type II Cas protein and a gRNA, and multiple nucleic acids, e.g., those comprising a nucleic acid encoding a type II Cas protein and a gRNA.

[0267] The nucleic acid encoding the Type II Cas protein and / or gRNA can be, for example, a plasmid or a viral genome (e.g., a lentivirus, retrovirus, adenovirus, or adeno-associated virus genome). The plasmid can be, for example, a plasmid for producing a viral particle, e.g., a lentivirus particle, or a plasmid for propagating the Type II Cas and gRNA coding sequence in a bacterial (e.g., E. coli) or eukaryotic (e.g., yeast) cell.

[0268] The nucleic acid encoding the type II Cas protein can, in some embodiments, further encode a gRNA, or the gRNA may be encoded by a separate nucleic acid (e.g., DNA or mRNA).

[0269] Nucleic acids encoding type II Cas proteins can be codon-optimized, e.g., by replacing at least one non-common or low-common codon with a codon that is common in the host cell. For example, the codon-optimized nucleic acid can direct the synthesis of an optimized messenger mRNA (e.g., optimized for expression in a mammalian expression system). As an example, if the intended target nucleic acid is in a human cell, a human codon-optimized polynucleotide encoding a type II Cas can be used to produce a type II Cas polypeptide. Exemplary codon-optimized sequences are shown in Tables 1A-1G and 2A-2C.

[0270] The nucleic acids of the present disclosure, e.g., plasmids and viral vectors, can contain one or more regulatory elements, such as promoters, enhancers, and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Such regulatory elements are described, for example, in Goeddel, 1990, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest or in a specific cell type. Regulatory elements can also direct expression in a time-dependent manner, e.g., cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell-type-specific. In some embodiments, the nucleic acids of the present disclosure comprise one or more pol III promoters (e.g., one, two, three, four, five, or more pol III promoters), one or more pol II promoters (e.g., one, two, three, four, five, or more pol II promoters), one or more pol I promoters (e.g., one, two, three, four, five, or more pol I promoters), or a combination thereof, e.g., to separately express a type II Cas protein and a gRNA. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters.Examples of Pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) (see, e.g., Boshart et al., 1985, Cell 41:521-530), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter (e.g., the full-length EF1α promoter and the EFS promoter, which is an intronless, shortened form of the full-length EF1α promoter). Exemplary enhancer elements include the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I; the SV40 enhancer; and the intron sequence between exon 2 and exon 3 of rabbit β-globin. Those skilled in the art will understand that the design of the expression vector can depend on factors such as the selection of the host cell and the desired expression level.

[0271] The term "vector" refers to a polynucleotide molecule capable of transporting another nucleic acid linked thereto. One type of polynucleotide vector includes a "plasmid," which refers to a circular double-stranded DNA loop into which additional nucleic acid segments are or may be ligated. Another type of polynucleotide vector is a viral vector; in this case, additional nucleic acid segments can be ligated into the viral genome. Certain vectors have the ability to autonomously replicate in host cells into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell, and thereby replicated along with the host genome.

[0272] In some cases, vectors are capable of directing the expression of nucleic acids to which they are operatively linked. Such vectors may be referred to herein as "recombinant expression vectors" or, more simply, "expression vectors," which serve equivalent functions.

[0273] The term "operably linked" means that the nucleotide sequence of interest is linked to a regulatory sequence in a manner that allows expression of the nucleotide sequence. The term "regulatory sequence" is intended to include, for example, promoters, enhancers, and other expression control elements (e.g., polyadenylation signals). Such regulatory sequences are well known in the art and are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990). Regulatory sequences include those that direct the constitutive expression of a nucleotide sequence in many types of host cells and those that direct the expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). Those skilled in the art will understand that the design of the expression vector can depend on factors such as the selection of the target cell, the desired expression level, etc.

[0274] Vectors may include, but are not limited to, viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus (e.g., AAV2, AAV5, AAV7m8, AAV8, AAV9, AAVrh8r, AAVrh10), SV40, herpes simplex virus, human immunodeficiency virus, retrovirus (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), and other recombinant vectors. Other vectors contemplated for eukaryotic target cells include, but are not limited to, vectors pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). Additional vectors contemplated for eukaryotic target cells include, but are not limited to, vectors pCTx-1, pCTx-2, and pCTx-3. Other vectors can also be used so long as they are compatible with the host cell.

[0275] In some examples, vectors can comprise one or more transcription and / or translation control elements.Depending on the host / vector system used, any of several suitable transcription and translation control elements can be used in expression vectors, including constitutive promoters and inducible promoters, transcription enhancer elements, transcription terminators, etc.The vector can be a self-inactivating vector that inactivates any of the components of viral sequences or CRISPR mechanisms or other elements.

[0276] Non-limiting examples of suitable eukaryotic promoters (promoters functional in eukaryotic cells) include those derived from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) from retroviruses, the human elongation factor-1 promoter (e.g., the complete EF1α promoter and the EFS promoter), a hybrid construct comprising the cytomegalovirus (CMV) enhancer fused to the chicken β-actin promoter (CAG), the murine stem cell virus promoter (MSCV), the phosphoglycerate kinase-1 locus promoter (PGK), and mouse metallothionein-I.

[0277] The expression vector may also contain a ribosome binding site for translation initiation and a transcription terminator. The expression vector may also contain an appropriate sequence for amplifying expression. The expression vector may also contain a nucleotide sequence encoding a non-natural tag (e.g., histidine tag, hemagglutinin tag, green fluorescent protein, etc.) that is fused with the site-specific polypeptide, thereby resulting in a fusion protein.

[0278] The promoter may be an inducible promoter (e.g., a heat shock promoter, a tetracycline-regulated promoter, a steroid-regulated promoter, a metal-regulated promoter, an estrogen receptor-regulated promoter, etc.). The promoter may be a constitutive promoter (e.g., a CMV promoter, a UBC promoter). In some cases, the promoter may be a spatially and / or temporally constrained promoter (e.g., a tissue-specific promoter, e.g., a human RHO promoter or a human rhodopsin kinase promoter (hGRK), a cell-type specific promoter, etc.).

[0279] 6.6. Particles and Cells The present disclosure further provides particles comprising a type II Cas protein of the present disclosure (e.g., an AAOF type II Cas protein, an ACEE type II Cas protein, an AEQH type II Cas protein, an AQSL type II Cas protein, an ASWC type II Cas protein, an AVFG type II Cas protein, an AWIT type II Cas protein, an AWMF type II Cas protein, a BUMO type II Cas protein, a COIA type II Cas protein, a DJQA type II Cas protein, or a DWET type II Cas protein), a gRNA of the present disclosure, a system of the present disclosure, and a nucleic acid or nucleic acids of the present disclosure. In some embodiments, the particle can comprise or further comprise a gRNA or a nucleic acid (e.g., DNA or mRNA) encoding the gRNA. For example, the particle can comprise an RNP of the present disclosure. Exemplary particles include lipid nanoparticles, vesicles, virus-like particles (VLPs), and gold nanoparticles. See, for example, International Publication No. WO 2020 / 012335, the entire contents of which are incorporated herein by reference, which describes vesicles that can be used to deliver gRNA molecules and type II Cas proteins (e.g., complexed together as RNPs) to cells.

[0280] The present disclosure provides a particle (e.g., a viral particle) comprising a nucleic acid encoding a type II Cas protein of the present disclosure. The particle can further comprise a nucleic acid encoding a gRNA. Alternatively, the nucleic acid encoding the type II Cas protein can further encode a gRNA.

[0281] The present disclosure further provides a plurality of particles (e.g., a plurality of viral particles). Such a plurality can include a particle encoding a type II Cas protein and a different particle encoding a gRNA. For example, the plurality of particles can include a viral particle encoding a type II Cas protein (e.g., an AAV2, AAV5, AAV7m8, AAV8, AAV9, AAVrh8r, or AAVrh10 viral particle) and a second viral particle encoding a gRNA (e.g., an AAV2, AAV5, AAV7m8, AAV8, AAV9, AAVrh8r, or AAVrh10 viral particle). Alternatively, the plurality of particles can include multiple viral particles, each particle encoding a type II Cas protein and a gRNA.

[0282] The present disclosure further provides cells and populations of cells (e.g., ex vivo cells and populations of cells) that may contain a type II Cas protein (e.g., introduced into a cell as an RNP) or a nucleic acid (e.g., DNA or mRNA) encoding a type II Cas protein (optionally also encoding a gRNA). The present disclosure further provides cells and populations of cells that contain a gRNA (optionally complexed with a type II Cas protein) or a nucleic acid (e.g., DNA or mRNA) encoding a gRNA (optionally also encoding a type II Cas protein). The cells and populations of cells can be, for example, human cells such as stem cells, e.g., hematopoietic stem cells (HSCs), pluripotent stem cells, induced pluripotent stem cells (iPS), or embryonic stem cells. In some embodiments, the cells and populations of cells are T cells. Methods for introducing proteins and nucleic acids into cells are known in the art. For example, RNPs can be generated by mixing a type II Cas protein and one or more guide RNAs in an appropriate buffer. RNPs can be introduced into cells, for example, via electroporation and other methods known in the art.

[0283] The cell populations of the present disclosure can be cells that have been gene-edited using the system of the present disclosure, or cells that have had components of the system of the present disclosure introduced or expressed but have not been gene-edited, or a combination thereof. The cell population can include, for example, a population in which at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70% of the cells have been gene-edited using the system of the present disclosure.

[0284] Pharmaceutical Compositions Also disclosed herein are pharmaceutical formulations and medicaments comprising a Type II Cas protein, gRNA, nucleic acid or nucleic acids, system, particle, or particle(s) of the present disclosure, together with a pharmaceutically acceptable excipient.

[0285] Suitable excipients include, but are not limited to, salts, diluents (e.g., Tris-HCl, acetate, phosphate), preservatives (e.g., thimerosal, benzyl alcohol, parabens), binders, fillers, solubilizers, disintegrants, adsorbents, solvents, pH adjusters, antioxidants, anti-infective agents, suspending agents, wetting agents, viscosity adjusters, isotonicity agents, stabilizers, and other ingredients, and combinations thereof. Suitable pharmaceutically acceptable excipients can be selected from materials generally recognized as safe (GRAS) and can be administered to individuals without causing undesirable biological side effects or undesirable interactions. Suitable excipients and their formulations are described in Remington's Pharmaceutical Sciences, 16th ed. 1980, Mack Publishing Co. In addition, such compositions can be complexed with polyethylene glycol (PEG), metal ions, or incorporated into polymeric compounds such as polylactic acid, polyglycolic acid, hydrogels, or incorporated into liposomes, microemulsions, micelles, unilamellar or multilamellar vesicles, erythrocyte ghosts, or spheroplasts. Suitable dosage forms for administration, e.g., parenteral administration, include solutions, suspensions, and emulsions.

[0286] The components of the pharmaceutical formulation can be dissolved or suspended in a suitable solvent, such as, for example, water, Ringer's solution, phosphate-buffered saline (PBS), or isotonic sodium chloride. The formulation may also be a sterile solution, suspension, or emulsion in a non-toxic parenterally acceptable diluent or solvent (e.g., 1,3-butanediol).

[0287] In some cases, the formulation can contain one or more isotonicity agents to adjust the isotonic range of the formulation.Suitable isotonicity agents are well known in the art, and include glycerin, mannitol, sorbitol, sodium chloride and other electrolytes.In some cases, the formulation can be buffered with an effective amount of buffer solution necessary to maintain the pH suitable for parenteral administration.Suitable buffer solutions are well known to those skilled in the art, and some examples of useful buffer solutions are acetate buffer solution, borate buffer solution, carbonate buffer solution, citrate buffer solution and phosphate buffer solution.

[0288] In some embodiments, the formulation can be distributed or packaged in liquid form or as a solid, for example, obtained by lyophilization of a suitable liquid formulation that can be reconstituted with a suitable carrier or diluent prior to administration. In some embodiments, the formulation can include a pharmaceutically effective amount of guide RNA and type II Cas protein sufficient to edit a gene in a cell. The pharmaceutical composition can be formulated for medical and / or veterinary use.

[0289] 6.8. Methods for Modifying Cells The present disclosure further provides methods of using the Type II Cas proteins, gRNAs, nucleic acids (including multiple nucleic acids), systems, and particles (including multiple particles) of the present disclosure to modify cells.

[0290] In one aspect, a method of modifying a cell comprises contacting a eukaryotic cell (e.g., a human cell) with a nucleic acid, particle, system, or pharmaceutical composition described herein.

[0291] Contacting cells with the disclosed nucleic acid, particle, system or pharmaceutical composition can be achieved by any method known in the art, and can be carried out in vivo, ex vivo or in vitro.In some embodiments, the method can include obtaining one or more cells from a subject before contacting cells with the disclosed nucleic acid, particle, system or pharmaceutical composition.In some embodiments, the method can further include returning or transplanting the contacted cells or their progeny back into the subject.

[0292] Type II Cas and gRNA, and nucleic acids encoding Type II Cas and gRNA, can be delivered to cells by any means known in the art, for example, by viral or non-viral delivery vehicles, electroporation, or lipid nanoparticles.

[0293] Polynucleotides encoding type II Cas and gRNA can be delivered to cells (ex vivo or in vivo) via lipid nanoparticles (LNPs). LNPs can have diameters of, for example, 1000 nm, 500 nm, 250 nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, or less than 25 nm. Alternatively, nanoparticles can range in size from 1 to 1000 nm, 1 to 500 nm, 1 to 250 nm, 25 to 200 nm, 25 to 100 nm, 35 to 75 nm, or 25 to 60 nm. LNPs can be composed of cationic, anionic, neutral lipids, and combinations thereof. Neutral lipids, such as the fusogenic phospholipid DOPE or the membrane component cholesterol, can be included in LNPs as "helper lipids" to enhance transfection activity and nanoparticle stability.

[0294] LNPs can also be composed of hydrophobic lipids, hydrophilic lipids, or both hydrophobic and hydrophilic lipids. Lipids and lipid combinations known in the art can be used to generate LNPs. Examples of lipids used to generate LNPs are: DOTMA, DOSPA, DOTAP, DMRIE, DC-cholesterol, DOTAP-cholesterol, GAP-DMORIE-DPyPE, and GL67A-DOPE-DMPE-polyethylene glycol (PEG). Examples of cationic lipids are 98N12-5, C12-200, DLin-KC2-DMA (KC2), DLin-MC3-DMA (MC3), XTC, MD1, and 7C1. Examples of neutral lipids are DPSC, DPPC, POPC, DOPE, and SM. Examples of PEG-modified lipids are PEG-DMG, PEG-CerCl4, and PEG-CerC20. Lipids can be combined in any number of molar ratios to generate LNPs. In addition, polynucleotides can be combined with lipids in a wide range of molar ratios to produce LNPs.

[0295] Type II Cas and / or gRNA can be delivered to cells via an adeno-associated viral vector (e.g., of the AAV2, AAV5, AAV7m8, AAV8, AAV9, AAVrh8r, or AAVrh10 serotype) or by another viral vector. Other viral vectors include, but are not limited to, lentivirus, adenovirus, alphavirus, enterovirus, pestivirus, baculovirus, herpesvirus, Epstein-Barr virus, papovavirus, poxvirus, vaccinia virus, and herpes simplex virus. In some embodiments, type II Cas mRNA is formulated in lipid nanoparticles, while the sgRNA is delivered to cells in an AAV or other viral vector. In some embodiments, one or more AAV vectors (e.g., one or more of the AAV2, AAV5, AAV7m8, AAV8, AAV9, AAVrh8r, or AAVrh10 serotype) are used to deliver both the sgRNA and type II Cas. In some embodiments, type II Cas and sgRNA are delivered using separate vectors.In other embodiments, type II Cas and sgRNA are delivered using a single vector.The BNK type II Cas and AIK type II Cas, which are relatively small in size, can be delivered together with gRNA (e.g., sgRNA) using a single AAV vector.

[0296] Compositions and methods for delivering Type II Cas and gRNA to cells and / or subjects are further described in PCT Patent Applications WO 2019 / 102381, WO 2020 / 012335, and WO 2020 / 053224, each of which is incorporated by reference herein in its entirety.

[0297] DNA cleavage can result in single-strand breaks (SSBs) or double-strand breaks (DSBs) at specific locations within a DNA molecule. Such breaks can be, and often are, repaired by natural endogenous cellular processes such as homology-dependent repair (HDR) and non-homologous end joining (NHEJ). These repair processes can edit the target polynucleotide by introducing mutations, thereby resulting in a polynucleotide with a sequence that is different from the sequence of the polynucleotide before cleavage by type II Cas.

[0298] NHEJ and HDR DNA repair processes consist of a family of alternative pathways. Non-homologous end joining (NHEJ) refers to a natural cellular process in which double-stranded DNA breaks are repaired by the direct joining of two non-homologous DNA segments. See, for example, Cahill et al., 2006, Front. Biosci. 11:1958-1976. DNA repair by non-homologous end joining is error-prone and often results in the non-templated addition or deletion of DNA sequences at the repair site. Thus, the NHEJ repair mechanism can introduce mutations into coding sequences that can disrupt gene function. NHEJ directly joins the DNA ends resulting from the double-stranded break, sometimes accompanied by modifications to the polynucleotide sequence, such as the loss or addition of nucleotides in the polynucleotide sequence. Alterations to the polynucleotide sequence can disrupt (or possibly enhance) gene expression.

[0299] Homology-dependent repair (HDR) uses homologous sequence or donor sequence as a template for inserting a specified DNA sequence into the breakpoint.Homologous sequence can be present in endogenous genome, such as sister chromatid.Alternatively, donor can be exogenous nucleic acid, such as plasmid, single-stranded oligonucleotide, double-stranded oligonucleotide, double-stranded oligonucleotide or virus, which has a highly homologous region with the locus cut by nuclease, but also contains additional sequence or sequence change (including deletion) that can be integrated into the cut target locus.

[0300] The third repair mechanism involves microhomology-mediated end joining (MMEJ), also known as "alternative NHEJ (ANHEJ)." This mechanism has a genetic outcome similar to NHEJ in that small deletions and insertions can occur at the break site. MMEJ can utilize a small number of base pairs of homologous sequences adjacent to the DNA break site to produce a more favorable DNA end joining repair outcome. In some cases, it is possible to predict the likely repair outcome based on an analysis of potential microhomologies at the site of the DNA break.

[0301] Modification of the cleaved polynucleotide by HDR, NHEJ, and / or ANHEJ can result in, for example, mutation, deletion, alteration, integration, gene correction, gene replacement, gene tagging, transgene insertion, nucleotide deletion, gene disruption, translocation, and / or gene mutation. The results of the aforementioned processes are examples of editing a polynucleotide.

[0302] Advantages of the ex vivo cell therapy approach include the ability to perform comprehensive analysis of the therapy before administration. Nuclease-based therapies may have some level of off-target effects. Performing gene correction ex vivo allows the user of the method to characterize the corrected cell population before transplantation, including identifying any undesired off-target effects. If undesired effects are observed, the user of the method may choose not to transplant the cells or their progeny, further edit the cells, or select new cells for editing and analysis. Another advantage is the ease of gene correction in iPSCs compared to other primary cell sources. iPSCs have excellent proliferation properties, making it easy to obtain the large numbers of cells required for cell-based therapies. Furthermore, iPSCs are an ideal cell type for clonal isolation. This allows for screening for correct genome correction without the risk of reduced viability.

[0303] Certain cells are attractive targets for ex vivo treatment and therapy, and the increased efficiency of delivery allows direct in vivo delivery to such cells.Ideally, targeting and editing are directed to relevant cells.Also, the use of promoters that are only active in specific cell types and / or developmental stages can prevent the cleavage in other cells.

[0304] Additional promoters are inducible, and therefore, when nucleases are delivered as plasmids, they can be temporally controlled.In addition, the length of time that delivered proteins and RNAs remain in cells can be adjusted by adding treatments or domains to change their half-life.In vivo treatment eliminates some treatment steps, but due to the low delivery rate, higher editing rates may be required.In vivo treatment can eliminate the problems and losses caused by ex vivo treatment and engraftment.

[0305] The advantage of in vivo gene therapy may be the ease of preparation and administration of therapeutic drug.The same therapeutic approach and treatment may be used to treat more than one patient, for example, several patients with the same or similar genotype or allele.In contrast, ex vivo cell therapy typically requires the use of the subject's own cells, which are isolated, manipulated and then returned to the same patient.

[0306] Progenitor cells (also referred to herein as stem cells) are capable of both proliferation and giving rise to more progenitor cells, thereby generating a large number of cells that can give rise to differentiated or differentiable daughter cells. The daughter cells themselves can be induced to proliferate and generate progeny that subsequently differentiate into one or more mature cell types, while retaining one or more cells with the developmental potential of the parent. In this context, the term "stem cell" refers to a cell that, under certain circumstances, has the ability or potential to differentiate into a more specialized or differentiated phenotype and, under certain circumstances, retains the ability to proliferate without substantial differentiation. In one aspect, the term progeny or stem cell refers to a generalized mother cell whose progeny (descendants) often specialize in different directions through differentiation, for example, by acquiring entirely separate characteristics as occurs in the progressive diversification of embryonic cells and tissues. Cell differentiation is a complex process that typically occurs through many cell divisions. Differentiated cells can themselves be derived from multipotent cells that are derived from multipotent cells, and so on. While each of these multipotent cells can be considered a stem cell, the range of cell types each can give rise to can vary considerably. Some differentiated cells also have the ability to give rise to cells with greater developmental potential. Such ability can be natural or can be artificially induced by treatment with various factors. In many biological cases, stem cells can also be "multipotent" because they can generate progeny of two or more distinct cell types, but this is not required.

[0307] The human cells described herein may be induced pluripotent stem cells (iPSCs). The advantage of using iPSCs in the methods of the present disclosure is that the cells can be derived from the same subject to which the progenitor cells are administered. That is, somatic cells can be obtained from the subject, reprogrammed into induced pluripotent stem cells, and then differentiated into progenitor cells (e.g., autologous cells) that are administered to the subject. Because the precursors are essentially derived from an autologous source, the risk of engraftment rejection or allergic reaction can be reduced compared to the use of cells derived from another subject or group of subjects. Furthermore, the use of iPSCs eliminates the need for cells obtained from an embryonic source. Thus, in one aspect, the stem cells used in the disclosed methods are not embryonic stem cells.

[0308] Methods that can be used to generate pluripotent stem cells from somatic cells are known in the art, and pluripotent stem cells generated by such methods can be used in the methods of the present disclosure.

[0309] A reprogramming method for generating pluripotent cells using a defined combination of transcription factors has been described. Mouse somatic cells can be converted into ES cell-like cells with expanded developmental potential by direct transduction of Oct4, Sox2, Klf4, and c-Myc; see, e.g., Takahashi and Yamanaka, 2006, Cell 126(4): 663-76. iPSCs resemble ES cells because they restore much of the transcriptional circuitry and epigenetic landscape associated with pluripotency. Furthermore, mouse iPSCs fulfill all standard assays for pluripotency: specifically, in vitro differentiation into cell types of the three germ layers, teratoma formation, chimera contribution, germline transmission (see, e.g., Maherali and Hochedlinger, 2008, Cell Stem Cell. 3(6): 595-605), and tetraploid complementation.

[0310] Human iPSCs can be obtained using similar transduction methods, and the transcription factor triad OCT4, SOX2, and NANOG has been established as a core set of transcription factors governing pluripotency; see, e.g., 2014, Budniatzky and Gepstein, Stem Cells Transl Med. 3(4):448-57; Barrett et al., 2014, Stem Cells Trans Med 3: 1-6 sctm.2014-0121; Focosi et al., 2014, Blood Cancer Journal 4: e211. The generation of iPSCs has historically been achieved by using viral vectors to introduce nucleic acid sequences encoding stem cell-associated genes into adult somatic cells.

[0311] iPSCs can be generated or derived from terminally differentiated somatic cells as well as from adult or somatic stem cells. That is, non-pluripotent progenitor cells can be reprogrammed to pluripotency or multipotency. In such cases, it may not be necessary to include the number of reprogramming factors required to reprogram terminally differentiated cells. Furthermore, reprogramming can be induced by non-viral introduction of reprogramming factors, for example, by introducing the protein itself, or by introducing a nucleic acid encoding the reprogramming factor, or by introducing a messenger RNA that produces the reprogramming factor upon translation (see, e.g., Warren et al., 2010, Cell Stem Cell, 7(5):618-30). Reprogramming can be achieved by introducing a combination of nucleic acids encoding stem cell-related genes, including, for example, Oct-4 (also known as Oct-3 / 4 or Pouf5l), Soxl, Sox2, Sox3, Sox15, Sox18, NANOG, Klf1, Klf2, Klf4, Klf5, NR5A2, c-Myc, l-Myc, n-Myc, Rem2, Tert, and LIN28. Reprogramming using the methods and compositions described herein may further include introducing one or more of Oct-3 / 4, a member of the Sox family, a member of the Klf family, and a member of the Myc family into somatic cells. The methods and compositions described herein may further include introducing one or more of Oct-4, Sox2, Nanog, c-MYC, and Klf4 for reprogramming. As noted above, the exact method used for reprogramming is not necessarily critical to the methods and compositions described herein. However, when cells differentiated from reprogrammed cells are used, for example, in human therapy, in one aspect, the reprogramming is not affected by methods that alter the genome, and thus, in such instances, reprogramming can be achieved without the use of, for example, viral or plasmid vectors.

[0312] The efficiency of reprogramming (the number of reprogrammed cells) from a population of starting cells can be enhanced by adding various drugs, such as small molecules, as shown by Shi et al., 2008, Cell-Stem Cell 2:525-528; Huangfu et al., 2008, Nature Biotechnology 26(7):795-797; and Marson et al., 2008, Cell-Stem Cell 3: 132-135.Therefore, drugs or drug combinations that enhance the efficiency or speed of induced pluripotent stem cell generation can be used in generating patient-specific or disease-specific iPSCs. Some non-limiting examples of agents that enhance reprogramming efficiency include soluble Wnt, Wnt-conditioned medium, BIX-01294 (G9a histone methyltransferase), PD0325901 (MEK inhibitor), DNA methyltransferase inhibitors, histone deacetylase (HDAC) inhibitors, valproic acid, 5'-azacytidine, dexamethasone, suberoylanilide, hydroxamic acid (SAHA), vitamin C, and trichostatin (TSA), among others.Other non-limiting examples of reprogramming enhancers include: suberoylanilide hydroxamic acid (SAHA (e.g., MK0683, vorinostat) and other hydroxamic acids), BML-210, depudecin (e.g., (-)-depudecin), HC toxin, Nullscript (4-(l,3-dioxo-lH,3H-benzo[de]isoquinolin-2-yl)-N-hydroxybutanamide), phenylbutyrate (e.g., sodium phenylbutyrate) and valproic acid ((VPA) and other short-chain fatty acids), Scriptaid, suramin sodium, trichostatin A (TSA), APHA compound 8, apicidin, sodium butyrate, pivaloyl Oxymethylbutyrate (Pivanex, AN-9), trapoxin B, chlamydocin, depsipeptide (also known as FR901228 or FK228), benzamides (e.g., CI-994 (e.g., N-acetyldinaline) and MS-27-275), MGCD0103, NVP-LAQ-824, CBHA (m-carboxycinnamic acid bishydroxamic acid), JNJ16241199, tubacin, A-161906, proxamide, oxamflatin, 3-C1-UCHA (e.g., 6-(3-chlorophenylureido)caproic acid hydroxamic acid), AOE (2-amino-8-oxo-9,10-epoxydecanoic acid), CHAP31, and CHAP50. Other reprogramming enhancers include, for example, dominant negative forms (e.g., catalytically inactive forms) of HDAC, siRNA inhibitors of HDAC, and antibodies that specifically bind to HDAC. Such inhibitors are available from, for example, BIOMOL International, Fukasawa, Merck Biosciences, Novartis, Gloucester Pharmaceuticals, Titan Pharmaceuticals, MethylGene, and Sigma-Aldrich.

[0313] To confirm the induction of pluripotent stem cells, isolated clones can be tested for the expression of stem cell markers.Such expression in cells derived from somatic cells identifies the cells as induced pluripotent stem cells.Stem cell markers can be selected from a non-limiting group, including SSEA3, SSEA4, CD9, Nanog, Fbxl5, Ecatl, Esgl, Eras, Gdfi, Fgf4, Cripto, Daxl, Zpf296, Slc2a3, Rexl, Utfl, and Natl.In one case, for example, cells that express Oct4 or Nanog are identified as pluripotent.Methods for detecting the expression of such markers include, for example, RT-PCR and immunological methods for detecting the presence of encoded polypeptides, such as Western blotting or flow cytometry analysis.Detection can include not only RT-PCR but also the detection of protein markers. Intracellular markers can best be identified by protein detection methods such as RT-PCR or immunocytochemistry, while cell surface markers are readily identified by, for example, immunocytochemistry.

[0314] The pluripotency of isolated cells can be confirmed by testing the ability of iPSCs to differentiate into cells of each of the three germ layers. As an example, teratoma formation in nude mice can be used to evaluate the pluripotency characteristics of isolated clones. The cells can be introduced into nude mice, and tumors arising from the cells can be subjected to histology and / or immunohistochemistry. For example, the growth of tumors containing cells from all three germ layers further indicates that the cells are pluripotent stem cells.

[0315] Patient-specific iPS cells or cell lines can be produced.For example, as described in Takahashi and Yamanaka 2006; Takahashi, Tanabe et al. 2007, many established methods for producing patient-specific iPS cells exist in the art.For example, the producing step can include a) isolating somatic cells such as skin cells or fibroblasts from patients, and b) introducing a set of pluripotency-related genes into somatic cells to induce cells to become pluripotent stem cells.The set of pluripotency-related genes can be one or more of the genes selected from the group consisting of OCT4, SOX1, SOX2, SOX3, SOX15, SOX18, NANOG, KLF1, KLF2, KLF4, KLF5, c-MYC, n-MYC, REM2, TERT and LIN28.

[0316] In some embodiments, a biopsy or aspiration of the subject's bone marrow can be performed. A biopsy or aspirate is a sample of tissue or fluid taken from the body. There are many different types of biopsies or aspirations. Nearly all of them involve using a sharp tool to remove a small amount of tissue. If the biopsy sample is on the skin or other sensitive area, an anesthetic can be applied first. The biopsy or aspiration can be performed according to any method known in the art. For example, in a bone marrow aspiration, a large needle is used to enter the pelvis and collect bone marrow.

[0317] In some embodiments, mesenchymal stem cells can be isolated from a subject. Mesenchymal stem cells can be isolated from a subject's bone marrow or peripheral blood, for example, according to any method known in the art. For example, bone marrow aspirate can be collected in a syringe containing heparin. Cells can be washed and centrifuged on a Percoll™ density gradient. Cells such as blood cells, hepatocytes, stromal cells, macrophages, mast cells, and thymocytes can be separated using Percoll™, a density gradient centrifugation medium. Subsequently, cells can be cultured in Dulbecco's modified Eagle's medium (DMEM) (low glucose) containing 10% fetal bovine serum (FBS) (Pittinger et al., 1999, Science 284: 143-147).

[0318] 6.8.1. Exemplary Genomic Targets The type II Cas protein and gRNA of the present disclosure can be used to modify various genome targets.In some embodiments, the method for modifying cells is for modifying CCR5, EMX1, Fas, FANCF, HBB, ZSCAN2, Chr6, ADAMTSL1, B2M, CXCR4, PD1, DNMT1, Match8, TRAC, TRBC, VEGFAsite2, VEGFAsite3, CACNA, HEKsite3, HEKsite4, Chr8, BCR, ATM, HBG1, HPRT, IL2RG, NF1, USH2A, RHO, BcLenh or CTFR genome sequence.In some embodiments, the method for modifying cells is for modifying TRAC, B2M, PD1 or LAG3 genome sequence.The reference sequences of RHO, TRAC, B2M, PD1 and LAG3 can be obtained from public databases, for example, those maintained by NCBI. For example, RHO has NCBI gene ID: 6010; TRAC has NCBI gene ID: 28755; B2M has NCBI gene ID: 567; PD1 has NCBI gene ID: 5133; and LAG3 has NCBI gene ID: 3902.

[0319] In some embodiments, the method for modifying cells is a method for modifying the hemoglobin subunit beta (HBB) gene. HBB mutations are associated with β-thalassemia and SCD. Dever et al., 2016 Nature 539(7629):384-389.

[0320] In some embodiments, the method for modifying cells is a method for modifying CCR5 gene.CCR5 has been demonstrated to be involved in several different disease states, including but not limited to human immunodeficiency virus (HIV) and acquired immune deficiency syndrome (AIDS).International Publication No. 2018 / 119359 describes CCR5 editing by CRISPR-Cas to eliminate the function of CCR5, so as to provide protection against HIV infection, alleviate one or more symptoms of HIV infection, stop or delay the progression from HIV to AIDS, and / or alleviate one or more symptoms of AIDS.

[0321] In some embodiments, the method for modifying cells is a method for modifying PD1, B2M, TRAC, or a combination thereof. CAR-T cells with PD1, B2M, and TRAC genes disrupted by CRISPR-II Cas showed enhanced activity in preclinical glioma models. Choi et al., 2019, Journal for ImmunoTherapy of Cancer 7:309.

[0322] In some embodiments, the method for modifying a cell is a method for modifying the USH2A gene. Mutations in the USH2A gene can cause Usher syndrome type 2A, which is characterized by progressive hearing and vision loss.

[0323] In some embodiments, the method of modifying a cell is a method for modifying the RHO gene. Mutations in the RHO gene can cause retinitis pigmentosa (RP).

[0324] Allele-specific editing of human RHO alleles carrying pathogenic mutations (e.g., P23 mutations such as P23H, or P347 mutations such as P347L, P347S, P347R, P347Q, P347T, or P347A) can be achieved using guide RNA (gRNA) molecules (e.g., having a spacer as shown in Table 6A) targeting the rs7984 SNP located in the 5' untranslated region (UTR) of the RHO gene. The SNP is highly common in the human population, and a significant proportion of subjects are heterozygous for the rs7984 SNP. For subjects who are heterozygous for the rs7984 SNP and heterozygous for a pathogenic RHO gene mutation, allele-specific editing of the RHO allele carrying the pathogenic mutation can be achieved by using gRNAs targeting the SNP variant found in the RHO allele of the subject carrying the pathogenic mutation. This allele-specific editing strategy, which does not directly target specific pathogenic RHO gene mutations, advantageously allows for editing of RHO genes with a wide variety of pathogenic mutations. A gRNA targeting the rs7984 SNP of the present disclosure can be used in combination with a second gRNA targeting a second site in the RHO gene, such as a site in intron 1 (e.g., a gRNA with a spacer as shown in Table 6B), to promote two cuts in the RHO gene with a pathogenic mutation. By cutting the RHO gene with a pathogenic mutation at two sites, deletion of the RHO gene with a pathogenic mutation can be promoted, resulting in reduced expression of mutant RHO protein.

[0325] Editing a subject's RHO allele may include editing the RHO allele in one or more cells derived from the subject (e.g., photoreceptor cells or retinal progenitor cells), or one or more cells derived from the subject's cells (e.g., induced pluripotent stem cells (iPSCs)). For example, one or more cells derived from the subject, or one or more cells derived from the subject's cells, can be contacted ex vivo with a nucleic acid, system, or particle of the present disclosure, and then transplanted into the subject a cell or its progeny having the edited RHO gene. The edited iPSCs can be differentiated, for example, into photoreceptor cells or retinal progenitor cells. In some embodiments, the resulting differentiated cells can be transplanted into the subject. If differentiated cells from the subject are edited, transplantation of the edited cells can proceed without an intervening differentiation step.

[0326] In vivo methods of RHO allele editing can include editing a pathogenic RHO allele in a subject's cells, such as photoreceptor cells or retinal progenitor cells. In some embodiments, the in vivo method includes administering one or more pharmaceutical compositions of the present disclosure to or near the subject's eye, for example, by subretinal or intravitreal injection. For example, a single pharmaceutical composition containing one or more AAV particles encoding one or more gRNAs of the present disclosure (e.g., a gRNA targeting the rs7984 SNP and a gRNA targeting RHO intron 1) and a type II Cas protein can be used; alternatively, multiple pharmaceutical compositions can be used, such as a first pharmaceutical composition containing AAV particles encoding a gRNA and a second separate pharmaceutical composition containing a second AAV particle encoding a type II Cas protein. When multiple pharmaceutical compositions are used, they are preferably administered sufficiently close in time so that the gRNA and type II Cas protein provided by the pharmaceutical compositions are present together in vivo.

[0327] Targeting of the human TRAC gene, the human B2M gene, the human PD1 gene, and the human LAG3 gene (one or more of these) can be used, for example, in engineering chimeric antigen receptor (CAR) T cells. For example, CRISPR / Cas technology has been used to deliver CAR-encoding DNA sequences to loci such as TRAC and PD1 (see, e.g., Eyquem et al., 2017, Nature 543(7643):113-117; Hu et al., 2023, eClinicalMedicine 60:102010), while TRAC, B2M, PD1, and LAG3 knockout CAR T cells have been reported (see, e.g., Dimitri et al., 2022, Molecular Cancer 21:78; Liu et al., 2016, Cell Research 27:154-157; Ren et al., 2017, Clin Cancer Res. 23(9):2255-2266; Zhang et al., 2017, Front Med. 11(4):554-562). Thus, the disclosed Type II Cas proteins and TRAC, B2M, PD1, and LAG3 guides can be used for targeted knock-in of exogenous DNA sequences to desired genomic sites in human cells, and / or knock-out of TRAC, B2M, PD1, or LAG3 in human cells, e.g., human T cells. In some embodiments, T cells are edited ex vivo to generate CAR-T cells, which are then administered to a subject in need of CAR-T cell therapy.

[0328] In some embodiments, the method for modifying cells is the method for modifying DNMT1 gene.Mutation in DNMT1 gene can cause DNMT1-related disorder, which is the degenerative disorder of central and peripheral nervous system.DNMT1-related disorder is characterized by sensory impairment, loss of sweating, dementia and hearing loss. [Example]

[0329] 7. Working Example

[0330] [Example 1] 7.1. Example 1: Identification and Characterization of Type II Cas Proteins This example describes studies performed to identify and characterize the AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins.

[0331] 7.1.1 Materials and Methods 7.1.1.1. Identification of Type II Cas Proteins from Metagenomic Data To identify novel type II Cas proteins, we screened 154,723 bacterial and archaeal metagenomic assembled genomes (MAGs) reconstructed from the human microbiome (Pasolli, et al., 2019, Cell 176(3):649-662.e20). Cas1, cas2, and cas9 genes were identified from protein annotation performed using Prokka version 1.12 (Seemann, 2014, Bioinformatics 30(14):2068-2069). CRISPR arrays were identified using MinCED version 0.4.2 (default parameters) (Bland, et al., 2007, BMC bioinformatics 8:209). Only loci with a CRISPR array and cas1-2-9 genes within a maximum distance of 10 kbp from each other were considered. Loci containing type II Cas proteins shorter than 950 aa were discarded. The resulting 17,173 CRISPR-type II Cas loci were filtered by selecting short proteins (<1,100 aa) from putatively unknown species. Type II Cas proteins from the same species with similar length but slightly different sequences were compared by multiple sequence alignment. Proteins showing deletions in nucleotide domains were discarded. The remaining proteins were compared for sequencing coverage, and the orthologs with the highest coverage were selected for each species.

[0332] 7.1.1.2.tracrRNA identification We identified tracrRNAs for CRISPR-II Cas loci of interest using a method based on the work of Chyou and Brown (Chyou and Brown, 2019, RNA biology 16(4):423-434). Starting with unique direct repeats in the CRISPR array, we used BLAST® (National Library of Medicine) version 2.2.31 (parameters: -task blastn-short -gapopen 2 -gapextend 1 -penalty -1 -reward 1 -evalue 1 -word_size 8) (Altschul, et al., 1990, Journal of Molecular Biology 215(3):403-410) to identify antirepeats within a 3000-bp window flanking the CRISPR-II Cas locus. A custom version of RNIE (Gardner, et al., 2011, Nucleic Acids Research 39(14):5845-5852) was used to predict Rho-independent transcription terminators (RITs) near the antirepeats. Putative tracrRNA sequences beginning with the antirepeat and ending with either a RIT (if found) or a polyT were combined with the directed repeats to form sgRNA scaffolds. The secondary structure of the sgRNA scaffolds was predicted using RNAsubopt version 2.4.14 (with parameters -noLP -e 5) (Lorenz, et al., 2011, Algorithms for Molecular Biology 6(1):26). sgRNAs lacking the functional modules identified by (Briner, et al., 2014 Molecular Cell 56(2):333-339), i.e., the repeat:antirepeat duplex, nexus, and 3' hairpin-like fold, were discarded.

[0333] 7.1.2.Results The AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins were identified. The amino acid sequences of the AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins and the nucleotide sequences encoding exemplary AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins are listed in Tables 1A-2K. Exemplary PAM sequences are listed in Table 5A. The PAM logo is shown in Figure 1. The crRNAs and tracrRNAs for the nucleases are described in Section 6.3. Exemplary sgRNA scaffolds are listed in Tables 7A-7B.

[0334] [Example 2] 7.2. Example 2: Further Characterization of Type II Cas Proteins This example describes studies performed to further characterize the AEQH, AAOF, ACEE, AQSL, ASWC, AVFG, AWIT, AWMF, BUMO, COIA, DJQA, and DWET type II Cas proteins.

[0335] 7.2.1 Materials and Methods Plasmids Type II Cas proteins were expressed in mammalian cells from a plasmid vector characterized by an EF1 alpha-driven cassette. Each type II Cas protein coding sequence was optimized for human codons and modified by the addition of an N-terminal SV5 tag and two bipartite nuclear localization signals (one at the N-terminus and one at the C-terminus). sgRNAs were expressed from a U6-driven cassette located on a separate plasmid construct. The human codon-optimized coding sequences of type II Cas proteins and the sgRNA scaffold were synthesized by Twist Bioscience. Spacer sequences were cloned into the sgRNA plasmid as annealed DNA oligonucleotides (Eurofins Genomics) using the double BsaI sites present in the plasmid. A list of the spacer sequences and relative cloning oligonucleotides used in this example is reported in Table 8. In all cases where the spacer did not contain a matching natural 5'-G, this nucleotide was added upstream of the targeting sequence to enable efficient transcription from the U6 promoter.

[0336] [Table 23-1]

[0337] [Table 23-2]

[0338] [Table 23-3]

[0339] 7.2.1.2. Cell lines U2OS-EGFP cells, carrying a single integrated copy of the EGFP reporter gene, were cultured in DMEM (Life Technologies) supplemented with 10% FBS (Life Technologies), 2 mM L-glutamine (Life Technologies), and penicillin / streptomycin (Life Technologies). All cells were incubated at 37°C and 5% CO in a humidified atmosphere. All cells tested were mycoplasma-negative (PlasmoTest, Invivogen).

[0340] 7.2.1.3. In vitro Cas PAM identification assay In vitro PAM evaluation of novel type II Cas proteins was performed according to the protocol from Karvelis et al., 2019, Methods in Enzymology, 616, pp. 219-240. Briefly, human codon-optimized versions of type II Cas protein genes obtained as synthetic constructs (Twist Bioscience) were cloned into an expression vector for in vitro transcription and translation (IVT) (pT7-N-His-GST, Thermo Fisher Scientific). sgRNAs for the assay were obtained by in vitro transcription of guides using a HighYield T7 RNA Synthesis Kit (Jena Bioscience), starting from PCR templates generated by amplification from each sgRNA expression construct, as commonly performed in the art. The primers used to generate the IVT templates are reported in Table 9. The in vitro transcribed gRNAs were then purified using a MEGAClear Transcription Cleanup Kit (Thermo Fisher Scientific). In vitro transcription and translation reactions for Cas expression were performed according to the manufacturer's protocol (1-Step Human High-Yield Mini IVT Kit, Thermo Fisher Scientific). A nuclease-guide RNA RNP complex was assembled by combining 20 μL of supernatant containing soluble type II Cas protein with 1 μL of RiboLock RNase Inhibitor (Thermo Fisher Scientific) and 2 μg of guide RNA (previously in vitro transcribed). This RNP complex was used to digest 1 μg of a PAM plasmid DNA library (containing a defined target sequence flanked at the 3' end by a randomized 8-nucleotide PAM sequence) at 37°C for 1 hour.

[0341] [Table 24]

[0342] Double-stranded DNA adapters (Table 10) were ligated to the DNA ends generated by targeted Cas cleavage, and the final ligation products were purified using a GeneJet PCR purification kit (Thermo Fisher Scientific).

[0343] [Table 25]

[0344] One round of two-step PCR (Phusion HF DNA polymerase, Thermo Fisher Scientific) was performed to enrich the cleaved sequences using a set of forward primers annealing on the adapter and a reverse primer designed on the plasmid backbone downstream of the PAM (Table 11). A second PCR was performed to ligate the Illumina index and adapter. The PCR product was purified using a GeneJet PCR purification kit (Thermo Fisher Scientific).

[0345] [Table 26]

[0346] Libraries were analyzed by 71 bp single-read sequencing using a flow cell v2 micro on an Illumina MiSeq® sequencer.

[0347] PAM sequences were extracted from Illumina MiSeq reads and used to generate PAM sequence logos using Logomaker version 0.8. PAM enrichment was displayed using a PAM heatmap, calculated as the frequency of a PAM sequence in the truncated library divided by the frequency of the same sequence in the control uncut library.

[0348] 7.2.1.4. Cell Line Transfection To perform the editing assay, 200,000 U2OS-EGFP cells were nucleofected with 500 ng of nuclease expression plasmid and 250 ng of sgRNA expression plasmid containing a guide designed to target EGFP using the 4D-Nucleofector™ SE Kit (Lonza) with the DN-100 program according to the manufacturer's protocol. After electroporation, cells were plated into 24-well plates. EGFP knockout was analyzed 4 days after nucleofection using a BD FACSymphony A1 (BD) flow cytometer.

[0349] 7.2.2.Results Having determined the sgRNA requirements for the selected type II Cas proteins, we were able to proceed with discovering the PAM sites recognized by each nuclease. To this end, we utilized a previously described in vitro PAM assay (see Methods). Briefly, this assay uses in vitro-translated type II Cas proteins coupled with in vitro-synthesized sgRNAs to generate functional ribonucleoprotein complexes that cleave a plasmid library characterized by a defined target sequence followed by a randomized 8-nt stretch corresponding to the putative PAM. The cleaved PAMs can then be recovered after library preparation by next-generation sequencing. Table 12, reported herein below, contains the PAM preferences determined based on the assay results. PAM logos and PAM heat maps reporting nucleotide preferences for specific positions along the PAM are reported in Figures 5A-8F.

[0350] [Table 27]

[0351] After identifying the PAM sequences and sgRNAs for selected Type II Cas proteins and obtaining preliminary information on the ability of these nucleases to cleave the desired target (the plasmid target used during the PAM assay) in vitro, we investigated their ability to cleave the selected target in mammalian cells. We used the EGFP reporter system because it allows for an easier readout of editing activity based on the loss of fluorescence in treated cells, quantitatively measured by cytofluorometry. To this end, sgRNAs targeting the EGFP coding sequence (three for each Type II Cas protein being evaluated) were designed for all Type II Cas proteins and evaluated in U2OS cells stably expressing a single copy of the EGFP reporter by transient electroporation. Loss of EGFP fluorescence, expressed as the percentage of EGFP-negative cells, was measured by cytofluorometry. Data are shown in Figure 9 as the mean ± SEM of n = 2 biologically independent experiments.

[0352] Surprisingly, as reported in Figure 9, several of the guides evaluated in combination with their respective type II Cas proteins were able to significantly downregulate EGFP expression in target cells. In particular, the AVFG, ACEE, AEQH, DJQA, and DWET type II Cas proteins showed very high activity (>90% EGFP KO) with one or more of the evaluated guides; the AAOF, BUMO, and COIA type II Cas proteins showed fairly high knockout activity (>50% EGFP KO) with at least one of the evaluated sgRNAs; the AWIT type II Cas protein showed lower editing activity (>30% with one of the three guide RNAs); and the ASWC and AQSL type II Cas proteins showed lower, but still moderate, activity (>10% EGFP KO) with at least one of the evaluated guide RNAs. The remaining AWMF type II Cas proteins did not show editing levels above the assay background for the targets evaluated in the EGFP-coding sequence. These data clearly demonstrate that some of the selected type II Cas proteins can highly efficiently modify gene targets in mammalian cells and can therefore be utilized to edit mammalian genomes.

[0353] [Example 3] 7.3. Example 3: Allele-specific RHO editing using AAOF, ACEE, AEQH, AVFG, BUMO, DJQA, and DWET Type II Cas This example describes the design and evaluation of a mutation-independent allele-specific strategy for selectively inactivating mutated RHO alleles. The RHO gene, encoding the photopigment rhodopsin, is one of the most frequently mutated genes in autosomal dominant retinitis pigmentosa, with more than 100 mutations described in the art. The large heterogeneity of mutations in affected patients and the overall low prevalence of most of these mutations make a mutation-independent approach to targeting the disease particularly desirable. Furthermore, effective knockout of affected alleles can be achieved effectively using gene editing tools such as type II Cas enzymes. The key to the success of this approach is the ability to preferentially downregulate RHO mutant alleles while leaving wild-type counterparts intact to preserve photoreceptor function.

[0354] The strategy described in this example utilizes rs7984, a non-pathogenic SNP located in the 5'-UTR of the gene, common in the general population, and commonly present in the RHO gene, to selectively target only one RHO allele containing a dominant-negative mutation, regardless of the true nature of the mutation.Only patients who are heterozygous for the rs7984 SNP may be eligible for this targeting strategy, which is based on accurate knowledge of the phase between the SNP allele and the mutation affecting each patient.Allele selectivity is achieved by selectively targeting the rs7984 allele that is in phase with the patient's mutation.

[0355] Because the rs7984 SNP is located outside the RHO coding sequence, a second break can be introduced into RHO intron 1 to remove the entire exon 1 and knock out expression of the mutant protein. This second break must occur in phase with the break on the rs7984 locus to create the desired deletion and can be biallelic, targeting a site present on both RHO alleles.

[0356] 7.3.1 Materials and Methods Plasmids The AAOF, ACEE, AEQH, AVFG, BUMO, DJQA, and DWET type II Cas proteins were expressed in mammalian cells using EF1 alpha-driven expression plasmids. Briefly, the human codon-optimized coding sequences of different type II Cas proteins were cloned into the aforementioned expression plasmids. The sgRNA scaffolds of each type II Cas protein (the trimmed scaffolds reported in Table 7A or Table 7B, with 3' uracil added) were cloned into an expression plasmid containing a human U6 promoter to drive guide RNA expression in mammalian cells. Each type II Cas coding sequence, modified with an N-terminal SV5 tag and two bipartite nuclear localization signals (one at the N-terminus and one at the C-terminus), and the sgRNA expression cassette (U6 promoter + sgRNA scaffold) were obtained as synthetic constructs by Twist Bioscience. The spacer sequences were cloned into the sgRNA expression plasmid as annealed DNA oligonucleotides using the double BsaI sites present in the plasmid. A list of the spacer sequences and relative cloning oligonucleotides used in this example is reported in Table 13.

[0357] [Table 28-1]

[0358] [Table 28-2]

[0359] [Table 28-3]

[0360] 7.3.1.2.Cell lines HEK293T cells (obtained from ATCC) and HEK293-rs7984G cells were cultured in DMEM (Life Technologies) supplemented with 10% FBS (Life Technologies), 2 mM L-glutamine (Life Technologies), and penicillin / streptomycin (Life Technologies). HEK293-rs7984G cells homozygous for the rs7984G SNP allele were obtained by base editing, and individual clones were subsequently isolated, expanded, characterized, and selected for further testing. All cells were incubated in a humidified atmosphere at 37°C and 5% CO2. All tested cells were mycoplasma-negative (PlasmoTest™, Invivogen).

[0361] 7.3.1.3. Cell Line Transfection To perform editing experiments on the target RHO locus, 100,000 HEK293T or HEK293-rs7984G cells were seeded into 24-well plates 24 hours before transfection. Subsequently, 500 ng of nuclease expression plasmid was transfected into the cells along with 250 ng of sgRNA expression vector targeting the locus of interest using TransIT®-LT1 Reagent (Mirus Bio) according to the manufacturer's protocol. Cell pellets were collected three days after transfection for analysis.

[0362] 7.3.1.4. Gene editing assessment Three days after transfection, cells were harvested and DNA was extracted using QuickExtract™ DNA Extraction Solution (Lucigen) according to the manufacturer's instructions. To amplify the target loci, PCR reactions were performed using HOT FIREPol® polymerase (Solis BioDyne) and the oligonucleotides listed in Table 14. The amplified products were purified, transported for Sanger sequencing (EasyRun service, Microsynth), and analyzed with the TIDE web tool (shinyapps.datacurators.nl / tide / ) to quantify indels. The primers used for the Sanger sequencing reactions on the amplicons are reported in Table 15 along with their respective target loci.

[0363] [Table 29]

[0364] [Table 30]

[0365] 7.3.1.5. Evaluation of large-scale editing at target loci The presence of large editing events (deletions and inversions caused by double cleavage) at the target RHO locus was assessed using multiple approaches. Deletion formation was detected by end-point PCR amplification of the target locus using the primers reported in Table 16, followed by visualization of the amplification product on an agarose gel. Alternatively, a qPCR assay (HOT FIREPol® Multiplex qPCR Mix) using the primers and probes reported in Table 17 was developed to detect the amount of unrearranged RHO alleles complementary to those containing large edits. Furthermore, a highly sensitive ddPCR assay to specifically quantify the desired deletion and inversion events was developed using the primers reported in Table 17 and Biorad ddPCR Supermix for Probes (without dUTP). Both the qPCR and ddPCR assays were probe-based and shared some of the same probes and primers (Tables 17 and 18). An internal reference was included in both assays by placing a specific primer-probe pair in a downstream exon of the RHO gene (RHO exon-intron 4) (see Tables 17 and 18).

[0366] [Table 31]

[0367] [Table 32]

[0368] [Table 33]

[0369] 7.3.2.Results 7.3.2.1. Type II Cas sgRNA for allele-specific targeting of the RHO rs7984 SNP We designed a set of sgRNAs associated with PAMs strongly recognized by the AAOF, ACEE, AEQH, AVFG, BUMO, DJQA, or DWET Type II Cas proteins spanning the rs7984 SNP (Figure 10). The editing activity of selected guides in combination with each nuclease was evaluated by transient transfection of HEK293T cells. These cells are homozygous for the rs7984A allele of the SNP, and sgRNAs targeting the rs7984A allele were used. As shown in Figure 11, many of the evaluated guides demonstrated a fairly high level of indel generation, with the ACEE, AVFG, and DJQA Type II Cas proteins demonstrating particularly high levels of modification (>40% indel formation) with at least one of the designed sgRNAs. On the other hand, AEQH type II Cas, which evaluated only one guide, showed intermediate levels of modification, while DWET, AAOF, and BUMO type II Cas proteins only resulted in moderate levels of editing at the target locus. Considering these results, the ACEE+g1 / g3, DJQA+g1, and AVFG+g2 nuclease-guide combinations were selected for further characterization.

[0370] Different spacer lengths of the selected guides were evaluated, ranging from 20 to 24 matching nucleotides (all targeting the rs7984A allele). Where necessary, a non-matching G nucleotide was added to the 5' end to enable efficient transcription of the human U6 promoter, as reported in the art. As shown in Figures 12A-12C, after transient transfection in HEK293T cells, the highest editing levels for all evaluated nucleases were obtained using the initially designed 23-nucleotide (nt) spacer. Therefore, this spacer length was used for subsequent testing.

[0371] Next, we performed two parallel approaches to verify the allele specificity of candidate Type II Cas (ACEE, AVFG, DJQA) in combination with selected guides for the rs7984 SNP target site. First, two alternative versions of each selected sgRNA targeting either the rs7984A or rs7984G allele were transiently transfected into HEK293T cells along with the corresponding Type II Cas. Because HEK293T cells are homozygous for rs7984A, on-target editing activity was measured using the rs7984A-targeting guide (Figure 13A, left bar), and allele specificity (G vs. A orientation) was assessed using the rs7984G-targeting sgRNA (Figure 13A, right bar).

[0372] In parallel, the same combination of constructs was transfected into an engineered HEK293T cell clone (HEK293T-rs7984G) modified by base editing to be homozygous for the rs7984G SNP allele to assess on-target cleavage activity against the rs7984G allele (Figure 13B, left bar) and verify allele specificity in the A vs. G orientation (Figure 13B, right bar). AVFG Type II Cas was not included in these evaluations.

[0373] This allowed for a thorough evaluation of the activity and specificity of alternative versions of candidate guides for both alleles of the rs7984 SNP, which will necessarily be utilized to cover the entire presumed eligible patient population. Overall, all evaluated sgRNAs showed allele preference for their intended target, which edited more efficiently compared to the non-targeting counter-allele. In particular, highly significant selectivity was observed in the A vs. G orientation for all guides. Based on the combined data regarding activity and specificity for the rs7984A / G allele, candidates ACEE+g1 / g3 and DJQA+g1 were selected for further testing.

[0374] 7.3.2.2. ACEE and DJQA Type II Cas sgRNAs for Targeting RHO Intron 1 Next, sgRNAs for ACEE and DJQA type II Cas targeting RHO intron 1 were used in combination with guide RNAs targeting the rs7984 SNP to screen for cleavage activity, with the ultimate goal of identifying high-performance guides for generating the desired large edits (deletions or inversions) encompassing RHO exon 1, including the ATG translation start site, to knock out mutant protein expression. A schematic representation of the sgRNA positions is reported in Figure 14A-B. HEK293T cells were transfected with either ACEE or DJQA type II Cas, along with a panel of sgRNAs targeting the first half of RHO intron 1, to generate deletions of less than 1,000 bp in size when used in combination with selected sgRNAs targeting the rs7984 SNP for each nuclease (ACEE:g1 / g3; DJQA:g1). Varying levels of indel formation were observed with different sgRNAs, and some guides failed to produce detectable editing at the target site in this study (Figures 15A-B). Notably, highly efficient guide RNAs were identified for all evaluated nucleases. Guides g481, g485, g543, g546, g794, and g948 for ACEE type II Cas and g700, g825, g884, and g973 for DJQA type II Cas performed particularly well and were therefore selected for further testing.

[0375] 7.3.2.3. Evaluating the formation of large editing events in the RHO gene using the best-performing sgRNA pairs Next, the formation of large editing events (e.g., large deletions and inversions) at the target RHO locus was evaluated after transient transfection of HEK293T cells with a combination of SNP-targeting sgRNAs (directed to the rs7984A allele) for ACEE and DJQA type II Cas with a guide selected to target RHO intron 1. As a control, a guide RNA that was not observed to result in high levels of indel formation in RHO intron 1 was included in the evaluated combination to assess the sensitivity of the readout.

[0376] Deletion formation was first assessed using a specially designed PCR assay with primers spanning the deletion region and visualized using agarose gel electrophoresis, where a low molecular weight band should be present if the desired deletion was correctly formed. As shown in Figure 16A, each of the tested combinations efficiently generated the desired deletion at the RHO locus. Notably, the control guides (g594 for ACEE and g703 for DJQA) showed reduced (g594 for ACEE) or eliminated (g703 for DJQA) deletion formation, as expected (Figure 16A).

[0377] Next, we quantitatively measured the degree of RHO gene modification and ranked the candidates to assess their efficacy by using a qPCR assay specifically designed to detect unmodified RHO alleles. This in turn allowed us to estimate the amount of large edits (especially the desired deletions and inversions) generated by each nuclease-guide RNA pair. As shown in Figure 16B, many candidates exhibited high levels of RHO gene modification (unmodified RHO exhibited low levels), and ACEE type II Cas was generally more effective than DJQA type II Cas in these tests. The control guides (g594+ACEE and g703+DJQA) again exhibited the highest levels of unmodified RHO alleles, consistent with previous data (see Figure 16B).

[0378] Among the evaluated candidates, those that were more effective at modifying the target locus were further investigated using a highly sensitive droplet digital PCR (ddPCR) assay to measure the absolute number of RHO alleles characterized by the desired deletion and inversion events after transient transfection of a cell population. To this end, HEK293T cells (homozygous rs7984A) were transiently transfected with either ACEE type II Cas in combination with g1 or g3 SNP targeting guides (rs7984A version) and intron guides g481, g543, g546, g794, or g948, or DJQA type II Cas in combination with g1 SNP targeting guides (rs7984A version) and intron guides g825 or g884. Subsequent quantification of deletion and inversion formation by ddPCR demonstrated efficient modification of the target locus for most of the evaluated candidates (Figure 17). Interestingly, different guide combinations resulted in different ratios of deletions to inversions, demonstrating specific repair preferences based on the precise patterns of DSBs generated by specific guide RNAs at the two target sites (see, for example, ACEE g1+g794 versus ACEE g3+g794 in Figure 17). Among all evaluated candidates, ACEE type II Cas generally showed higher efficacy in generating the desired edits at the RHO target locus, with combinations of SNP g1 / g3 with introns g543, g546, or g794 being among the top performers (Figure 17).

[0379] Taken together, we have identified sgRNAs that target the RHO gene and generate high-level desired deletions / inversions of the first exon of the gene, while also being selective for the desired RHO copy.

[0380] [Example 4] 7.4. Example 4: Gene Editing Using BDLP and EQSC Type II Cas Proteins To comprehensively evaluate the cleavage activity of the AVFG, ACEE, AEQH, DJQA, and DWET type II Cass, we selected a panel of endogenous loci (B2M, TRAC, and PD-1) commonly targeted to generate allogeneic CAR-T cells (chimeric antigen receptor T cells) for editing studies. For each target locus, multiple sgRNAs were designed and evaluated in parallel by transient plasmid transfection in HEK293T cells.

[0381] Materials and methods were similar to those used in Example 3. Table 19 shows the protospacer and oligo sequences used to clone the sgRNA spacers. Table 20 shows the oligos used for TIDE analysis.

[0382] [Table 34-1]

[0383] [Table 34-2]

[0384] [Table 34-3]

[0385] [Table 34-4]

[0386] [Table 34-5]

[0387] [Table 35]

[0388] As shown in Figures 18A-18C, for each of the evaluated loci, the majority of nucleases exhibited significant editing activity with at least one of the selected guide RNAs, demonstrating the ability of these novel Type II Cas proteins to effectively modify genomic targets of interest.

[0389] 8. Specific Embodiments The present disclosure is illustrated by the following specific embodiments. 1. (a) Amino acid sequence of the RuvC-I domain of the reference protein sequence; (b) Amino acid sequence of the RuvC-II domain of the reference protein sequence; (c) the amino acid sequence of the RuvC-III domain of the reference protein sequence; (d) the amino acid sequence of the BH domain of the reference protein sequence; (e) the amino acid sequence of the REC domain of the reference protein sequence; (f) the amino acid sequence of the HNH domain of the reference protein sequence; (g) the amino acid sequence of the WED domain of the reference protein sequence; (h) the amino acid sequence of the PID domain of the reference protein sequence; or (i) the full-length amino acid sequence of the reference protein sequence A type II Cas protein comprising an amino acid sequence having at least 50% sequence identity to A Type II Cas protein, wherein the reference protein sequence is SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:67, or SEQ ID NO:68. 2. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 3. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 4. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 5. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 6. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 7. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 8. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 9. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the RuvC-I domain of a reference protein sequence. 10. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the RuvC-I domain of a reference protein sequence. 11. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the RuvC-I domain of a reference protein sequence. 12. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 13. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 14. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 15. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the RuvC-I domain of a reference protein sequence. 16. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the RuvC-I domain of the reference protein sequence. 17. The type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the RuvC-II domain of the reference protein sequence. 18. The type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the RuvC-II domain of the reference protein sequence. 19. The type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the RuvC-II domain of the reference protein sequence. 20. The type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the RuvC-II domain of the reference protein sequence. 21. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 22. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 23. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 24. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 25. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 26. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 27. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 28. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 29. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the RuvC-II domain of the reference protein sequence. 30. The type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the RuvC-II domain of the reference protein sequence. 31. A type II Cas protein according to any one of embodiments 1 to 16, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the RuvC-II domain of a reference protein sequence. 32. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 33. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 34. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 35. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 36. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 37. A type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 38. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 39. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 40. A type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 41. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 42. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 43. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 44. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 45. The type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 46. ​​A type II Cas protein according to any one of embodiments 1 to 31, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the RuvC-III domain of the reference protein sequence. 47. A type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the BH domain of the reference protein sequence. 48. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the BH domain of the reference protein sequence. 49. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the BH domain of the reference protein sequence. 50. A type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the BH domain of the reference protein sequence. 51. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the BH domain of the reference protein sequence. 52. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the BH domain of the reference protein sequence. 53. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the BH domain of the reference protein sequence. 54. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the BH domain of the reference protein sequence. 55. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the BH domain of the reference protein sequence. 56. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the BH domain of the reference protein sequence. 57. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the BH domain of the reference protein sequence. 58. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the BH domain of the reference protein sequence. 59. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the BH domain of the reference protein sequence. 60. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the BH domain of the reference protein sequence. 61. The type II Cas protein according to any one of embodiments 1 to 46, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the BH domain of the reference protein sequence. 62. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the REC domain of the reference protein sequence. 63. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the REC domain of the reference protein sequence. 64. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the REC domain of the reference protein sequence. 65. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the REC domain of the reference protein sequence. 66. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the REC domain of the reference protein sequence. 67. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the REC domain of the reference protein sequence. 68. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the REC domain of the reference protein sequence. 69. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the REC domain of the reference protein sequence. 70. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the REC domain of the reference protein sequence. 71. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the REC domain of the reference protein sequence. 72. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the REC domain of the reference protein sequence. 73. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the REC domain of the reference protein sequence. 74. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the REC domain of the reference protein sequence. 75. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the REC domain of the reference protein sequence. 76. The type II Cas protein according to any one of embodiments 1 to 61, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the REC domain of the reference protein sequence. 77. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 78. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 79. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 80. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 81. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 82. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 83. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 84. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 85. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 86. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 87. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 88. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 89. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 90. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the HNH domain of the reference protein sequence. 91. The type II Cas protein according to any one of embodiments 1 to 76, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the HNH domain of the reference protein sequence. 92. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the WED domain of the reference protein sequence. 93. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the WED domain of the reference protein sequence. 94. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the WED domain of the reference protein sequence. 95. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the WED domain of the reference protein sequence. 96. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the WED domain of the reference protein sequence. 97. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the WED domain of the reference protein sequence. 98. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the WED domain of the reference protein sequence. 99. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the WED domain of the reference protein sequence. 100. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the WED domain of the reference protein sequence. 101. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the WED domain of the reference protein sequence. 102. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the WED domain of the reference protein sequence. 103. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the WED domain of the reference protein sequence. 104. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the WED domain of the reference protein sequence. 105. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the WED domain of the reference protein sequence. 106. The type II Cas protein according to any one of embodiments 1 to 91, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the WED domain of the reference protein sequence. 107. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 50% identical to the amino acid sequence of the PID domain of the reference protein sequence. 108. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical to the amino acid sequence of the PID domain of the reference protein sequence. 109. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical to the amino acid sequence of the PID domain of the reference protein sequence. 110. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical to the amino acid sequence of the PID domain of the reference protein sequence. 111. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of the PID domain of the reference protein sequence. 112. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical to the amino acid sequence of the PID domain of the reference protein sequence. 113. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the PID domain of the reference protein sequence. 114. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of the PID domain of the reference protein sequence. 115. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of the PID domain of the reference protein sequence. 116. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of the PID domain of the reference protein sequence. 117. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical to the amino acid sequence of the PID domain of the reference protein sequence. 118. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of the PID domain of the reference protein sequence. 119. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of the PID domain of the reference protein sequence. 120. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of the PID domain of the reference protein sequence. 121. The type II Cas protein according to any one of embodiments 1 to 106, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the amino acid sequence of the PID domain of a reference protein sequence. 122. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 55% identical over the entire length of a reference protein sequence. 123. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 60% identical over the entire length of a reference protein sequence. 124. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 65% identical over the entire length of a reference protein sequence. 125. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 70% identical over the entire length of a reference protein sequence. 126. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 75% identical over the entire length of a reference protein sequence. 127. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 80% identical over the entire length of a reference protein sequence. 128. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 85% identical over the entire length of a reference protein sequence. 129. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 90% identical over the entire length of a reference protein sequence. 130. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 95% identical over the entire length of a reference protein sequence. 131. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 96% identical over the entire length of a reference protein sequence. 132. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 97% identical over the entire length of a reference protein sequence. 133. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 98% identical over the entire length of a reference protein sequence. 134. The type II Cas protein of embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is at least 99% identical over the entire length of a reference protein sequence. 135. The type II Cas protein according to embodiment 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence that is identical to the entire length of a reference protein sequence. 136. The type II Cas protein according to any one of embodiments 1 to 135, which is a chimeric type II Cas protein. 137. A type II Cas protein according to any one of embodiments 1 to 136, which is a fusion protein. 138. The type II Cas protein of embodiment 137, comprising one or more nuclear localization signals. 139. The type II Cas protein of embodiment 138, comprising two or more nuclear localization signals. 140. A type II Cas protein according to embodiment 138 or embodiment 139, comprising an N-terminal nuclear localization signal. 141. A type II Cas protein according to any one of embodiments 138 to 140, comprising a C-terminal nuclear localization signal. 142. The type II Cas protein of any one of embodiments 138-141, comprising an N-terminal nuclear localization signal and a C-terminal nuclear localization signal. 143. The amino acid sequence of one or more of the nuclear localization signals is selected from the group consisting of the amino acid sequences KRTADGSEFESPKKKRKV (SEQ ID NO: 145), PKKKRKV (SEQ ID NO: 146), PKKKRRV (SEQ ID NO: 147), KRPAATKKAGQAKKKK (SEQ ID NO: 148), YGRKKRRQRRR (SEQ ID NO: 149), RKKRRQRRR (SEQ ID NO: 150), PAAKRVKLD (SEQ ID NO: 151), RQRRNELKRSP (SEQ ID NO: 152), VSRKRPRP (SEQ ID NO: 153), PPKKARED (SEQ ID NO: 154), PQPKKKPL (SEQ ID NO: 155), SALIKKK 143. The type II Cas protein of any one of embodiments 138 to 142, comprising KKMAP (SEQ ID NO: 156), PKQKKRK (SEQ ID NO: 157), RKLKKKIKKL (SEQ ID NO: 158), REKKKFLKRR (SEQ ID NO: 159), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 160), RKCLQAGMNLEARKTKK (SEQ ID NO: 161), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 162), or RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 163). 144. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 145). 145. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence PKKKRKV (SEQ ID NO: 146). 146. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence PKKKRRV (SEQ ID NO: 147). 147. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence KRPAATKKAGQAKKKK (SEQ ID NO: 148). 148. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence YGRKKRRQRRR (SEQ ID NO: 149). 149. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence RKKRRQRRR (SEQ ID NO: 150). 150. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence PAAKRVKLD (sequence number 151). 151. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence RQRRNELKRSP (sequence number 152). 152. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence VSRKRPRP (SEQ ID NO: 153). 153. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence PPKKARED (SEQ ID NO: 154). 154. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence PQPKKKPL (SEQ ID NO: 155). 155. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence SALIKKKKKMAP (SEQ ID NO: 156). 156. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence PKQKKRK (SEQ ID NO: 157). 157. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence RKLKKKIKKL (SEQ ID NO: 158). 158. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence REKKKFLKRR (sequence number 159). 159. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 160). 160. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 161). 161. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (sequence number 162). 162. The type II Cas protein of embodiment 143, wherein the amino acid sequence of one or more of the nuclear localization signals comprises the amino acid sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 163). 163. A type II Cas protein according to any one of embodiments 138 to 162, wherein the amino acid sequence of each nuclear localization signal is the same. 164. The type II Cas protein according to any one of embodiments 136-163, comprising a fusion partner which is a DNA-modifying enzyme, an RNA-modifying enzyme or a protein-modifying enzyme, optionally wherein the DNA-modifying enzyme, the RNA-modifying enzyme or the protein-modifying enzyme is an adenosine deaminase, a cytidine deaminase, a reverse transcriptase, a guanosyltransferase, a DNA methyltransferase, an RNA methyltransferase, a DNA demethylase, an RNA demethylase, a dioxygenase, a polyadenylate polymerase, a pseudouridine synthase, an acetyltransferase, a deacetylase, a ubiquitin ligase, a deubiquitinase, a kinase, a phosphatase, a NEDD8 ligase, a NEDD8 deconjugator, a SUMO ligase, a SUMO deconjugator, a histone deacetylase, a histone acetyltransferase, a histone methyltransferase or a histone demethylase. 165. A type II Cas protein according to any one of embodiments 136 to 164, comprising a means for deaminating adenosine, optionally wherein the means for deaminating adenosine is adenosine deaminase. 166. The type II Cas protein of any one of embodiments 136-164, comprising a fusion partner that is an adenosine deaminase, optionally wherein the amino acid sequence of the adenosine deaminase comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 166, and optionally wherein the adenosine deaminase is an adenosine deaminase moiety comprised in the adenine base editor ABE8e. 167. A type II Cas protein according to any one of embodiments 136 to 164, comprising a means for deaminating cytidine, optionally wherein the means for deaminating cytidine is a cytidine deaminase. 168. A type II Cas protein according to any one of embodiments 136 to 164, comprising a fusion partner that is a cytidine deaminase. 169. A type II Cas protein according to any one of embodiments 136 to 164, comprising a means for synthesizing DNA from a single-stranded template, optionally wherein the means for synthesizing DNA from a single-stranded template is a reverse transcriptase. 170. A type II Cas protein according to any one of embodiments 136 to 164, comprising a fusion partner that is a reverse transcriptase. 171. A type II Cas protein according to any one of embodiments 136 to 170, comprising a tag. 172. The type II Cas protein of embodiment 171, wherein the tag is an SV5 tag, and optionally the SV5 tag comprises the amino acid sequence GKPIPNPLLGLDST (SEQ ID NO: 164) or IPNPLLGLD (SEQ ID NO: 165). 173. A type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 1. 174. The type II Cas protein of embodiment 173, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 1. 175. A type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 2. 176. A type II Cas protein according to any one of embodiments 173 to 175, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 2. 177. The type II Cas protein of embodiment 173 or embodiment 174, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 3. 178. A type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 7. 179. The type II Cas protein of embodiment 178, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 7. 180. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 8. 181. The type II Cas protein according to any one of embodiments 178 to 180, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 8. 182. The type II Cas protein of embodiment 178 or embodiment 179, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 9. 183. The type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 13. 184. The type II Cas protein of embodiment 183, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 13. 185. The type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 14. 186. The type II Cas protein according to any one of embodiments 183 to 185, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 14. 187. The type II Cas protein of embodiment 183 or embodiment 184, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 15. 188. A type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 19. 189. The type II Cas protein of embodiment 188, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 19. 190. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 20. 191. A type II Cas protein according to any one of embodiments 188 to 190, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 20. 192. The type II Cas protein of embodiment 188 or embodiment 189, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 21. 193. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 25. 194. The type II Cas protein of embodiment 193, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 25. 195. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 26. 196. The type II Cas protein according to any one of embodiments 193 to 195, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 26. 197. The type II Cas protein of embodiment 194 or embodiment 195, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 27. 198. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 31. 199. The type II Cas protein of embodiment 198, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 31. 200. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 32. 201. The type II Cas protein of any one of embodiments 199-200, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 32. 202. The type II Cas protein of embodiment 198 or embodiment 199, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 33. 203. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 37. 204. The type II Cas protein of embodiment 203, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 37. 205. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 38. 206. The type II Cas protein according to any one of embodiments 203 to 205, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 38. 207. The type II Cas protein of embodiment 203 or embodiment 204, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 39. 208. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 43. 209. The type II Cas protein of embodiment 208, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 43. 210. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 44. 211. The type II Cas protein according to any one of embodiments 208 to 210, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 44. 212. The type II Cas protein of embodiment 208 or embodiment 209, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 45. 213. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 49. 214. The type II Cas protein of embodiment 213, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 49. 215. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 50. 216. A type II Cas protein according to any one of embodiments 213 to 215, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 50. 217. The type II Cas protein of embodiment 213 or embodiment 214, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 51. 218. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 55. 219. The type II Cas protein of embodiment 218, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 55. 220. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 56. 221. A type II Cas protein according to any one of embodiments 218 to 220, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 56. 222. The type II Cas protein of embodiment 218 or embodiment 219, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 57. 223. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 61. 224. The type II Cas protein of embodiment 223, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 61. 225. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 62. 226. The type II Cas protein according to any one of embodiments 223 to 225, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 62. 227. The type II Cas protein of embodiment 223 or embodiment 224, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 63. 228. The type II Cas protein of any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 67. 229. The type II Cas protein of embodiment 228, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 67. 230. The type II Cas protein according to any one of embodiments 1 to 172, wherein the reference protein sequence is SEQ ID NO: 68. 231. The type II Cas protein according to any one of embodiments 228 to 230, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 68. 232. The type II Cas protein of embodiment 228 or embodiment 229, wherein the amino acid sequence comprises the amino acid sequence of SEQ ID NO: 69. 233. A type II Cas protein whose amino acid sequence is identical to the type II Cas protein of any one of embodiments 1 to 232, except for one or more amino acid substitutions compared to a reference sequence that confer nickase activity, optionally wherein the one or more amino acid substitutions include a substitution (e.g., an alanine substitution) at a position corresponding to position D10 of SaCas9, position N580 of SaCas9, or position H559 of CjCas9 (e.g., as shown in Table 4). 234. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions that confer nickase activity are in the RuvC or HNH domain. 235. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include a D10A substitution, and the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 2. 236. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N627A substitution, the position of the N627A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 2. 237. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H604A substitution, the position of the H604A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 2. 238. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include a D9A substitution, and the position of the D9A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 8. 239. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N610A substitution, and the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 8. 240. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an H587A substitution, wherein the position of the H587A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 8. 241. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include a D8A substitution, and the position of the D8A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 14. 242. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an N613A substitution, wherein the position of the N613A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 14. 243. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an H590 substitution, the position of the H590 substitution being defined with respect to the amino acid numbering of SEQ ID NO: 14. 244. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include a D9A substitution, and the position of the D9A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 20. 245. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N610A substitution, wherein the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 20. 246. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H587A substitution, the position of the H587A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 20. 247. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include a D15A substitution, and the position of the D15A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 26. 248. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N633A substitution, wherein the position of the N633A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 26. 249. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H610A substitution, the position of the H610A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 26. 250. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include a D12A substitution, the position of the D12A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 32. 251. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N607A substitution, wherein the position of the N607A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 32. 252. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an H584A substitution, wherein the position of the H584A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 32. 253. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include a D10A substitution, and the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 38. 254. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N615A substitution, wherein the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 38. 255. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H592A substitution, the position of the H592A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 38. 256. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include a D10A substitution, and the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 44. 257. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N615A substitution, wherein the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 44. 258. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an H592A substitution, the position of the H592A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 44. 259. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include a D8A substitution, and the position of the D8A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 50. 260. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an N610A substitution, wherein the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 50. 261. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H587A substitution, the position of the H587A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 50. 262. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include a D10A substitution, and the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 56. 263. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N615A substitution, wherein the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 56. 264. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H592A substitution, the position of the H592A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 56. 265. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include a D10A substitution, and the position of the D10A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 62. 266. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N615A substitution, wherein the position of the N615A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 62. 267. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an H592A substitution, the position of the H592A substitution being defined with respect to the amino acid numbering of SEQ ID NO: 62. 268. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include a D9A substitution, and the position of the D9A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 68. 269. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions compared to a reference sequence that confer nickase activity include an N610A substitution, wherein the position of the N610A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 68. 270. The type II Cas of embodiment 233, wherein the one or more amino acid substitutions relative to a reference sequence that confer nickase activity include an H587A substitution, wherein the position of the H587A substitution is defined with respect to the amino acid numbering of SEQ ID NO: 68. 271. A type II Cas protein whose amino acid sequence is identical to the type II Cas protein of any one of embodiments 1-232, except for one or more amino acid substitutions compared to a reference sequence that render the type II Cas protein catalytically inactive, optionally wherein the one or more amino acid substitutions comprise a substitution (e.g., an alanine substitution) at a position corresponding to position D10 of SaCas9, position N580 of SaCas9, or position H559 of CjCas9 (e.g., as shown in Table 4), or a combination thereof. 272. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and N627A substitutions, wherein the positions of the D10A and N627A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 2. 273. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and H604A substitutions, wherein the positions of the D10A and H604A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 2. 274. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D9A and N610A substitutions, wherein the positions of the D9A and N610A substitutions are defined with respect to the amino acid numbering of SEQ ID NO:8. 275. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D9A and H587A substitutions, wherein the positions of the D9A and H587A substitutions are defined with respect to the amino acid numbering of SEQ ID NO:8. 276. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D8A and N613A substitutions, wherein the positions of the D8A and N613A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 14. 277. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D8A and H590A substitutions, wherein the positions of the D8A and H590A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 14. 278. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D9A and N610A substitutions, wherein the positions of the D9A and N610A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 20. 279. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D9A and H587A substitutions, wherein the positions of the D9A and H587A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 20. 280. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D15A and N633A substitutions, wherein the positions of the D15A and N633A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 26. 281. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D15A and H610A substitutions, wherein the positions of the D15A and H610A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 26. 282. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D12A and N607A substitutions, wherein the positions of the D12A and N607A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 32. 283. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D12A and H584A substitutions, wherein the positions of the D12A and H584A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 32. 284. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and N615A substitutions, wherein the positions of the D10A and N615A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 38. 285. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and H592A substitutions, wherein the positions of the D10A and H592A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 38. 286. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and N615A substitutions, wherein the positions of the D10A and N615A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 44. 287. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and H592A substitutions, wherein the positions of the D10A and H592A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 44. 288. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D8A and N610A substitutions, wherein the positions of the D8A and N610A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 50. 289. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D8A and H587A substitutions, wherein the positions of the D8A and H587A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 50. 290. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and N615A substitutions, wherein the positions of the D10A and N615A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 56. 291. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and H592A substitutions, wherein the positions of the D10A and H592A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 56. 292. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and N615A substitutions, wherein the positions of the D10A and N615A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 62. 293. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D10A and H592A substitutions, wherein the positions of the D10A and H592A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 62. 294. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D9A and N615A substitutions, wherein the positions of the D9A and N615A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 68. 295. The type II Cas protein of embodiment 271, wherein the one or more amino acid substitutions that render the type II Cas protein catalytically inactive comprise D9A and H587A substitutions, wherein the positions of the D9A and H587A substitutions are defined with respect to the amino acid numbering of SEQ ID NO: 68. 296.AEQH type II Cas guide RNA (gRNA) molecule. 297. AAOF type II Cas guide RNA (gRNA) molecule. 298. ACEE type II Cas guide RNA (gRNA) molecule. 299.AQSL type II Cas guide RNA (gRNA) molecule. 300. ASWC type II Cas guide RNA (gRNA) molecule. 301.AVFG type II Cas guide RNA (gRNA) molecule. 302.AWIT type II Cas guide RNA (gRNA) molecule. 303.AWMF type II Cas guide RNA (gRNA) molecule. 304. BUMO type II Cas guide RNA (gRNA) molecule. 305. COIA type II Cas guide RNA (gRNA) molecule. 306.DJQA Type II Cas guide RNA (gRNA) molecule. 307. DWET type II Cas guide RNA (gRNA) molecule. 308. A gRNA according to any one of embodiments 296 to 307, which is a gRNA for editing the human RHO gene. 309. A gRNA according to any one of embodiments 296 to 307, which is a gRNA for editing the human B2M gene. 310. A gRNA according to any one of embodiments 296 to 307, which is a gRNA for editing the human TRAC gene. 311. A gRNA according to any one of embodiments 296 to 307, which is a gRNA for editing the human LAG3 gene. 312. A gRNA according to any one of embodiments 296 to 307, which is a gRNA for editing the human PD1 gene. 313. A guide RNA (gRNA) molecule for editing a human RHO gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 314 to 399. 314. A guide RNA (gRNA) molecule for editing the human RHO gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 400 to 419. 315. A guide RNA (gRNA) molecule for editing the human B2M gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 420 to 441. 316. A guide RNA (gRNA) molecule for editing the human TRAC gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 442 to 461. 317. A guide RNA (gRNA) molecule for editing the human PD1 gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 462 to 482. 318. The gRNA of any one of embodiments 313 to 317, comprising a spacer that is 15 to 30 nucleotides in length. 319. The gRNA of embodiment 318, wherein the spacer is 18 to 30 nucleotides in length. 320. The gRNA of embodiment 318, wherein the spacer is 20 to 28 nucleotides in length. 321. The gRNA of embodiment 318, wherein the spacer is 22 to 26 nucleotides in length. 322. The gRNA of embodiment 318, wherein the spacer is 23 to 25 nucleotides in length. 323. The gRNA of embodiment 318, wherein the spacer is 22 to 25 nucleotides in length. 324. The gRNA of embodiment 318, wherein the spacer is 15 to 25 nucleotides in length. 325. The gRNA of embodiment 318, wherein the spacer is 16 to 24 nucleotides in length. 326. The gRNA of embodiment 318, wherein the spacer is 17 to 23 nucleotides in length. 327. The gRNA of embodiment 318, wherein the spacer is 18 to 22 nucleotides in length. 328. The gRNA of embodiment 318, wherein the spacer is 19 to 21 nucleotides in length. 329. The gRNA of embodiment 318, wherein the spacer is 25 nucleotides in length. 330. The gRNA of embodiment 318, wherein the spacer is 24 nucleotides in length. 331. The gRNA of embodiment 318, wherein the spacer is 23 nucleotides in length. 332. The gRNA of embodiment 318, wherein the spacer is 22 nucleotides in length. 333. The gRNA of embodiment 318, wherein the spacer is 21 nucleotides in length. 334. The gRNA of embodiment 318, wherein the spacer is 20 nucleotides in length. 335. The gRNA of any one of embodiments 313 to 334, wherein the spacer comprises 16 or more consecutive nucleotides of the reference sequence. 336. The gRNA of any one of embodiments 313 to 334, wherein the spacer comprises 17 or more consecutive nucleotides of the reference sequence. 337. The gRNA of any one of embodiments 313 to 334, wherein the spacer comprises 18 or more consecutive nucleotides of the reference sequence. 338. The gRNA of any one of embodiments 313 to 334, wherein the spacer comprises 19 or more consecutive nucleotides of the reference sequence. 339. A gRNA according to any one of embodiments 313 to 334, wherein the spacer comprises 20 consecutive nucleotides of the reference sequence. 340. A gRNA according to any one of embodiments 313 to 333, wherein the reference sequence is a reference sequence having at least 21 nucleotides and the spacer comprises 21 consecutive nucleotides of the reference sequence. 341. A gRNA according to any one of embodiments 313 to 332, wherein the reference sequence is a reference sequence having at least 22 nucleotides and the spacer comprises 22 consecutive nucleotides of the reference sequence. 342. A gRNA according to any one of embodiments 313 to 331, wherein the reference sequence is a reference sequence having at least 23 nucleotides and the spacer comprises 23 consecutive nucleotides of the reference sequence. 343. A gRNA according to any one of embodiments 313 to 330, wherein the reference sequence is a reference sequence having at least 24 nucleotides and the spacer comprises 24 consecutive nucleotides of the reference sequence. 344. The gRNA of any one of embodiments 313 to 334, wherein the spacer comprises a nucleotide sequence that is at least 90% identical to a reference sequence. 345. The gRNA of embodiment 344, wherein the spacer comprises a nucleotide sequence that is at least 95% identical to a reference sequence. 346. A gRNA according to any one of embodiments 313 to 334, wherein the spacer comprises a nucleotide sequence having one mismatch compared to the reference sequence. 347. A gRNA according to any one of embodiments 313 to 334, wherein the spacer comprises a nucleotide sequence having two mismatches compared to the reference sequence. 348. A gRNA according to any one of embodiments 313 to 334, wherein the spacer comprises a reference sequence. 349. A gRNA according to any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCAUUCUUGGGUGGGAGCAGCCR (SEQ ID NO: 314). 350. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCAUUCUUGGGUGGGAGCAGCCA (SEQ ID NO: 315). 351. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCAUUCUUGGGUGGGAGCAGCCG (SEQ ID NO: 316). 352. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GUGGCUGACCCGYGGCUGCUCCCA (SEQ ID NO: 317). 353. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GUGGCUGACCCGUGGCUGCUCCCA (SEQ ID NO: 318). 354. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GUGGCUGACCCGCGGCUGCUCCCA (SEQ ID NO: 319). 355. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCRCGGGUC (SEQ ID NO: 320). 356. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCACGGGUC (SEQ ID NO: 321). 357. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCGCGGGUC (SEQ ID NO: 322). 358. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGCUGACCCGYGGCUGCUCCCAC (SEQ ID NO: 323). 359. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGCUGACCCGUGGCUGCUCCCAC (SEQ ID NO: 324). 360. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGCUGACCCGCGGCUGCUCCCAC (SEQ ID NO: 325). 361. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCCCUUGUGGCUGACCCGYGGCU (SEQ ID NO: 326). 362. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCCCUUGUGGCUGACCCGUGGCU (SEQ ID NO: 327). 363. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCCCUUGUGGCUGACCCGCGGCU (SEQ ID NO: 328). 364. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGUGGCUGACCCGUGGCU (SEQ ID NO: 329). 365. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CCUUGUGGCUGACCCGUGGCU (SEQ ID NO: 330). 366. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGCCCUUGUGGCUGACCCGUGGCU (SEQ ID NO: 331). 367. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGUGGCUGACCCGCGGCU (SEQ ID NO: 332). 368. A gRNA according to any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CCUUGUGGCUGACCCGCGGCU (SEQ ID NO: 333). 369. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGCCCUUGUGGCUGACCCGCGGCU (SEQ ID NO: 334). 370. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUCGCAGCAUUCUUGGGUGGGA (sequence number 335). 371. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UCUUGGGUGGGAGCAGCCRCGGG (SEQ ID NO: 336). 372. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UCUUGGGUGGGAGCAGCCACGGG (SEQ ID NO: 337). 373. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UCUUGGGUGGGAGCAGCCGCGGG (SEQ ID NO: 338). 374. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCACGGG (SEQ ID NO: 339). 375. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCACGGG (SEQ ID NO: 340). 376. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCACGGG (SEQ ID NO: 341). 377. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUCUUGGGUGGGAGCAGCCACGGG (SEQ ID NO: 342). 378. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCGCGGG (SEQ ID NO: 343). 379. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCGCGGG (SEQ ID NO: 344). 380. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCGCGGG (SEQ ID NO: 345). 381. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUCUUGGGUGGGAGCAGCCGCGGG (SEQ ID NO: 346). 382. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGAGCAGCCRCGGGUCAGCCACA (SEQ ID NO: 347). 383. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGAGCAGCCACGGGUCAGCCACA (SEQ ID NO: 348). 384. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGAGCAGCCGCGGGUCAGCCACA (SEQ ID NO: 349). 385. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGUGGCUGACCCGYGGCUGCUCC (SEQ ID NO: 350). 386. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGUGGCUGACCCGUGGCUGCUCC (SEQ ID NO: 351). 387. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGUGGCUGACCCGCGGCUGCUCC (SEQ ID NO: 352). 388. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCRCGGGUCA (SEQ ID NO: 353). 389. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCACGGGUCA (SEQ ID NO: 354). 390. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCGCGGGUCA (SEQ ID NO: 355). 391. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GUGGGAGCAGCCACGGGUCA (SEQ ID NO: 356). 392. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGUGGGAGCAGCCACGGGUCA (SEQ ID NO: 357). 393. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGGUGGGAGCAGCCACGGGUCA (SEQ ID NO: 358). 394. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCACGGGUCA (SEQ ID NO: 359). 395. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GUGGGAGCAGCCGCGGGUCA (SEQ ID NO: 360). 396. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGUGGGAGCAGCCGCGGGUCA (SEQ ID NO: 361). 397. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGGUGGGAGCAGCCGCGGGUCA (SEQ ID NO: 362). 398. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCGCGGGUCA (SEQ ID NO: 363). 399. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGUGGCUGACCCGYGGCUGCUC (SEQ ID NO: 364). 400. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGUGGCUGACCCGUGGCUGCUC (SEQ ID NO: 365). 401. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGUGGCUGACCCGCGGCUGCUC (SEQ ID NO: 366). 402. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCRCGGGU (SEQ ID NO: 367). 403. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCACGGGU (SEQ ID NO: 368). 404. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCGCGGGU (SEQ ID NO: 369). 405. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCCCUUGUGGCUGACCCGYGGCU (SEQ ID NO: 370). 406. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCCCUUGUGGCUGACCCGUGGCU (SEQ ID NO: 371). 407. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GCCCUUGUGGCUGACCCGCGGCU (SEQ ID NO: 372). 408. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUCACAGCAUUCUUGGGUGGGA (sequence number 373). 409. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UCUUGGGUGGGAGCAGCCRCGGG (SEQ ID NO: 374). 410. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UCUUGGGUGGGAGCAGCCACGGG (SEQ ID NO: 375). 411. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UCUUGGGUGGGAGCAGCCGCGGG (SEQ ID NO: 376). 412. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGAGCAGCCRCGGGUCAGCCACA (SEQ ID NO: 377). 413. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGAGCAGCCACGGGUCAGCCACA (SEQ ID NO: 378). 414. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGAGCAGCCGCGGGUCAGCCACA (SEQ ID NO: 379). 415. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCRCGGGUC (SEQ ID NO: 380). 416. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCACGGGUC (SEQ ID NO: 381). 417. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGGGUGGGAGCAGCCGCGGGUC (SEQ ID NO: 382). 418. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGUGGGAGCAGCCACGGGUC (SEQ ID NO: 383). 419. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGGUGGGAGCAGCCACGGGUC (SEQ ID NO: 384). 420. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCACGGGUC (SEQ ID NO: 385). 421. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCACGGGUC (SEQ ID NO: 386). 422. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGUGGGAGCAGCCGCGGGUC (SEQ ID NO: 387). 423. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GGGUGGGAGCAGCCGCGGGUC (SEQ ID NO: 388). 424. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UGGGUGGGAGCAGCCGCGGGUC (SEQ ID NO: 389). 425. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCGCGGGUC (SEQ ID NO: 390). 426. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGUGGCUGACCCGYGGCUGCUC (SEQ ID NO: 391). 427. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGUGGCUGACCCGUGGCUGCUC (SEQ ID NO: 392). 428. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is UUGUGGCUGACCCGCGGCUGCUC (SEQ ID NO: 393). 429. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCRCGGGU (SEQ ID NO: 394). 430. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCACGGGU (SEQ ID NO: 395). 431. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is CUUGGGUGGGAGCAGCCGCGGGU (SEQ ID NO: 396). 432. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GAGCAGCCRCGGGUCAGCCACAA (SEQ ID NO: 397). 433. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GAGCAGCCACGGGUCAGCCACAA (SEQ ID NO: 398). 434. The gRNA of any one of embodiments 313 and 318 to 348 when dependent on embodiment 313, wherein the reference sequence is GAGCAGCCGCGGGUCAGCCACAA (SEQ ID NO: 399). 435. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is AUGCUCCCGGGCUCCUGCACACC (SEQ ID NO: 400). 436. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is CUCCAUGCUCCCGGGCUCCUGCA (SEQ ID NO: 401). 437. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is AGGAGAAGGGAGAAGGCCUCUCA (SEQ ID NO: 402). 438. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is GACAGGAGAAGGGAGAAGGCCUC (SEQ ID NO: 403). 439. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is GUUAUCCAAAGCCCUCAUAUAUU (SEQ ID NO: 404). 440. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is GCCCGAGAUAGAUGCGGGCUUCC (SEQ ID NO: 405). 441. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is ACAUGGCCCGAGAUAGAUGCGGG (SEQ ID NO: 406). 442. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is AUGACUGGAGAAUGGAAAAUCCA (SEQ ID NO: 407). 443. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is AGACCCAAUGACUGGAGAAUGGA (SEQ ID NO: 408). 444. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is UGCUUCAGAGCGGCUGCUUGCGG (SEQ ID NO: 409). 445. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is UGUUGACUGAAUAUAUGAGGGCU (SEQ ID NO: 410). 446. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is UAUCCAAAGCCCUCAUAUAUUCA (SEQ ID NO: 411). 447. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is CAAGGCAGUGUUCAGUGCCAGCC (SEQ ID NO: 412). 448. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is AAAUUAGACAAGCGCAUAUUGCU (SEQ ID NO: 413). 449. The gRNA of any one of embodiments 314 and 318 to 348 when dependent on embodiment 314, wherein the reference sequence is AGCUCAGUUUUCUUGCUGUGAAA (sequence number 414...

Claims

1. (a) the amino acid sequence of the RuvC-I domain of the reference protein sequence; (b) the amino acid sequence of the RuvC-II domain of the reference protein sequence; (c) the amino acid sequence of the RuvC-III domain of the reference protein sequence; (d) the amino acid sequence of the BH domain of the reference protein sequence; (e) the amino acid sequence of the REC domain of the reference protein sequence; (f) the amino acid sequence of the HNH domain of the reference protein sequence; (g) the amino acid sequence of the WED domain of the reference protein sequence; (h) the amino acid sequence of the PID domain of the reference protein sequence; or (i) the full-length amino acid sequence of the reference protein sequence A type II Cas protein comprising an amino acid sequence having at least 50% sequence identity to The reference protein sequence is SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:67, or SEQ ID NO:

68.

2. 2. The type II Cas protein of claim 1, wherein the amino acid sequence of the type II Cas protein is at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or comprises an amino acid sequence that is identical to the entire length of the reference protein sequence.

3. 2. The type II Cas protein of claim 1, wherein the amino acid sequence of the type II Cas protein comprises an amino acid sequence identical to the entire length of the reference protein sequence.

4. The type II Cas protein of any one of claims 1 to 3, which is a fusion protein.

5. 5. The type II Cas protein of claim 4, comprising one or more nuclear localization signals, such as two or more nuclear localization signals, optionally comprising an N-terminal nuclear localization signal and / or a C-terminal nuclear localization signal.

6. The amino acid sequence of one or more of the nuclear localization signals may be selected from the group consisting of the amino acid sequences KRTADGSEFESPKKKRKV (SEQ ID NO: 145), PKKKRKV (SEQ ID NO: 146), PKKKRRV (SEQ ID NO: 147), KRPAATKKAGQAKKKK (SEQ ID NO: 148), YGRKKRRQRRR (SEQ ID NO: 149), RKKRRQRRR (SEQ ID NO: 150), PAAKRVKLD (SEQ ID NO: 151), RQRRNELKRSP (SEQ ID NO: 152), VSRKRPRP (SEQ ID NO: 153), PPKKARED (SEQ ID NO: 154), PQPKKKPL (SEQ ID NO: 155), S 6. The type II Cas protein of claim 5, comprising ALIKKKKKKMAP (SEQ ID NO: 156), PKQKKRK (SEQ ID NO: 157), RKLKKKIKKL (SEQ ID NO: 158), REKKKFLKRR (SEQ ID NO: 159), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 160), RKCLQAGMNLEARKTKK (SEQ ID NO: 161), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 162), or RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 163).

7. The type II Cas protein of claim 5 or claim 6, wherein the amino acid sequences of each nuclear localization signal are the same.

8. 8. The Type II Cas protein of any one of claims 4 to 7, comprising a fusion partner that is a DNA-modifying enzyme, an RNA-modifying enzyme, or a protein-modifying enzyme, optionally wherein the DNA-modifying enzyme, the RNA-modifying enzyme, or the protein-modifying enzyme is an adenosine deaminase, a cytidine deaminase, a reverse transcriptase, a guanosyltransferase, a DNA methyltransferase, an RNA methyltransferase, a DNA demethylase, an RNA demethylase, a dioxygenase, a polyadenylate polymerase, a pseudouridine synthase, an acetyltransferase, a deacetylase, a ubiquitin ligase, a deubiquitinase, a kinase, a phosphatase, a NEDD8 ligase, a de-NEDD8 deconjugating enzyme, a SUMO ligase, a de-SUMO deconjugating enzyme, a histone deacetylase, a histone acetyltransferase, a histone methyltransferase, or a histone demethylase.

9. 9. The Type II Cas protein of any one of claims 4 to 8, comprising: (a) a fusion partner that is an adenosine deaminase, optionally wherein the amino acid sequence of the adenosine deaminase comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 166, and optionally wherein the adenosine deaminase is the adenosine deaminase moiety contained in the adenine base editor ABE8e; (b) a fusion partner that is a cytosine deaminase; or (c) a fusion partner that is a reverse transcriptase.

10. 10. The type II Cas protein of any one of claims 4 to 9, comprising a tag, for example an SV5 tag, optionally wherein the SV5 tag comprises the amino acid sequence GKPIPNPLLGLDST (SEQ ID NO: 164).

11. (a) the reference protein sequence is SEQ ID NO: 1 or SEQ ID NO: 2, and optionally the amino acid sequence of the Type II Cas protein comprises the amino acid sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3; (b) the reference protein sequence is SEQ ID NO: 7 or SEQ ID NO: 8, and optionally the amino acid sequence of the Type II Cas protein comprises the amino acid sequence of SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9; (c) the reference protein sequence is SEQ ID NO: 13 or SEQ ID NO: 14, and optionally the amino acid sequence of the Type II Cas protein comprises the amino acid sequence of SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15; (d) the reference protein sequence is SEQ ID NO: 19 or SEQ ID NO: 20, and optionally the amino acid sequence of the Type II Cas protein comprises the amino acid sequence of SEQ ID NO: 19, SEQ ID NO: 20, or SEQ ID NO: 21; (e) the reference protein sequence is SEQ ID NO: 25 or SEQ ID NO: 26, and optionally the amino acid sequence of the Type II Cas protein comprises the amino acid sequence of SEQ ID NO: 25, SEQ ID NO: 26, or SEQ ID NO: 27; (f (g) the reference protein sequence is SEQ ID NO:37 or SEQ ID NO:38, optionally wherein the amino acid sequence of the type II Cas protein comprises the amino acid sequence of SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39; (h) the reference protein sequence is SEQ ID NO:43 or SEQ ID NO:44, optionally wherein the amino acid sequence of the type II Cas protein comprises the amino acid sequence of SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45; (i) the reference protein sequence is SEQ ID NO:49 or SEQ ID NO:50, optionally wherein the amino acid sequence of the type II Cas protein comprises the amino acid sequence of SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51; (j) the reference protein sequence is SEQ ID NO:55 or SEQ ID NO:56, optionally wherein the amino acid sequence of the type II Cas protein comprises the amino acid sequence of SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57;(k) the reference protein sequence is SEQ ID NO:61 or SEQ ID NO:62, and optionally the amino acid sequence of the type II Cas protein comprises the amino acid sequence of SEQ ID NO:61, SEQ ID NO:62, or SEQ ID NO:63; (l) the reference protein sequence is SEQ ID NO:67 or SEQ ID NO:68, and optionally the amino acid sequence of the type II Cas protein comprises the amino acid sequence of SEQ ID NO:67, SEQ ID NO:68, or SEQ ID NO:

69. The type II Cas protein of any one of claims 1 to 10.

12. 12. A type II Cas protein whose amino acid sequence is identical to the type II Cas protein of any one of claims 1 to 11, except for one or more amino acid substitutions compared to the reference sequence that confer nickase activity, optionally wherein the one or more amino acid substitutions comprise a substitution (e.g., an alanine substitution) at a position corresponding to position D10 of SaCas9, position N580 of SaCas9, or position H559 of CjCas9 (e.g., as shown in Table 4).

13. 12. A type II Cas protein whose amino acid sequence is identical to the type II Cas protein of any one of claims 1 to 11, except for one or more amino acid substitutions compared to the reference sequence that render the type II Cas protein catalytically inactive, optionally wherein the one or more amino acid substitutions comprise a substitution (e.g., an alanine substitution) at a position corresponding to position D10 of SaCas9, position N580 of SaCas9, or position H559 of CjCas9 (e.g., as shown in Table 4), or a combination thereof.

14. A guide RNA (gRNA) molecule for editing a human RHO gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 314-399.

15. A guide RNA (gRNA) molecule for editing a human RHO gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 400-419.

16. 1. A guide RNA (gRNA) molecule for editing a human B2M gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 420-441.

17. A guide RNA (gRNA) molecule for editing a human TRAC gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 442-461.

18. A guide RNA (gRNA) molecule for editing the human PD1 gene, comprising a spacer, wherein the nucleotide sequence of the spacer comprises 15 or more consecutive nucleotides of a reference sequence or comprises a nucleotide sequence that is at least 85% identical to the reference sequence, and the reference sequence is selected from SEQ ID NOs: 462-482.

19. 19. The gRNA of any one of claims 14-18, comprising a spacer that is 15-30 nucleotides in length, 18-30 nucleotides in length, 20-28 nucleotides in length, 22-26 nucleotides in length, 23-25 ​​nucleotides in length, 22-25 nucleotides in length, 15-25 nucleotides in length, 16-24 nucleotides in length, 17-23 nucleotides in length, 18-22 nucleotides in length, 19-21 nucleotides in length, 25 nucleotides in length, 24 nucleotides in length, 23 nucleotides in length, 22 nucleotides in length, 21 nucleotides in length, or 20 nucleotides in length.

20. The gRNA of any one of claims 14 to 19, wherein the spacer comprises the reference sequence.

21. The gRNA of any one of claims 14 to 20, which is a single guide RNA (sgRNA).

22. A gRNA comprising a spacer and an sgRNA scaffold, (a) the spacer is located 5′ to the sgRNA scaffold; and (b) the nucleotide sequence of the sgRNA scaffold comprises a nucleotide sequence that is at least 50% identical to a reference scaffold sequence, and the reference scaffold sequence is any one of SEQ ID NOs: 97-120; gRNA.

23. 23. The gRNA of Claim 22, wherein the sgRNA scaffold comprises a nucleotide sequence that is at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to the reference scaffold sequence; a nucleotide sequence that has no more than 5 nucleotide mismatches, no more than 4 nucleotide mismatches, no more than 3 nucleotide mismatches, no more than 2 nucleotide mismatches, or no more than 1 nucleotide mismatch with respect to the reference scaffold sequence; or a nucleotide sequence that is 100% identical to the reference scaffold sequence.

24. 24. The gRNA of claim 22 or claim 23, wherein the sgRNA scaffold comprises 1 to 8 uracils at its 3' end.

25. 25. The gRNA of any one of claims 22 to 24, wherein the nucleotide sequence of the spacer is partially or completely complementary to a target mammalian genomic sequence.

26. 26. The gRNA of any one of claims 22 to 25, wherein the nucleotide sequence of the spacer comprises a sequence selected from SEQ ID NOs: 314 to 482.

27. (a) a combination of gRNAs comprising a first gRNA comprising a spacer, wherein the sequence of the spacer comprises any one of SEQ ID NOs: 326 to 348, and a second gRNA comprising a spacer, wherein the sequence of the spacer comprises any one of SEQ ID NOs: 400 to 409, optionally wherein the spacer of the first and / or second gRNA is located 5' to an sgRNA scaffold, and the sequence of the sgRNA scaffold comprises the sequence of SEQ ID NO: 101, 102, 125, or 126; or (b) a combination of gRNAs comprising a first gRNA comprising a spacer, wherein the sequence of the spacer comprises any one of SEQ ID NOs: 380 to 390, and a second gRNA comprising a spacer, wherein the sequence of the spacer comprises any one of SEQ ID NOs: 410 to 419, optionally wherein the spacer of the first and / or second gRNA is located 5' to an sgRNA scaffold, and the sequence of the sgRNA scaffold comprises the sequence of SEQ ID NO: 117, 118, 141, or 142.

28. 27. A system comprising a Type II Cas protein according to any one of claims 1 to 13 and a guide RNA (gRNA) comprising a spacer sequence, optionally wherein the gRNA is a gRNA according to any one of claims 14 to 26.

29. 14. A nucleic acid encoding a type II Cas protein according to any one of claims 1 to 13, optionally wherein the nucleotide sequence encoding said type II Cas protein is operably linked to a promoter heterologous to said type II Cas protein.

30. 30. The nucleic acid of claim 29, wherein the nucleotide sequence encoding the Type II Cas protein is codon-optimized for expression in a human cell.

31. 40. The nucleic acid of any one of claims 37 to 39, which is a plasmid or a viral genome, optionally an adeno-associated virus (AAV) genome, such as an AAV2, AAV5, AAV7m8, AAV8, AAV9, AAVrh8r or AAVrhlO genome.

32. The nucleic acid of any one of claims 29 to 31, further encoding a gRNA.

33. A nucleic acid encoding a gRNA according to any one of claims 14 to 26 or a combination of gRNAs according to claim 27.

34. 29. A nucleic acid encoding the type II Cas protein and gRNA of the system of claim 28.

35. 29. A plurality of nucleic acids comprising separate nucleic acids encoding a Type II Cas protein and a gRNA of the system of claim 28.

36. A particle comprising a type II Cas protein according to any one of claims 1 to 13, a gRNA according to any one of claims 14 to 26, a combination of gRNAs according to claim 27, a system according to claim 28, a nucleic acid according to any one of claims 29 to 34, or a plurality of nucleic acids according to claim 35.

37. A pharmaceutical composition comprising a type II Cas protein according to any one of claims 1 to 13, a gRNA according to any one of claims 14 to 26, a combination of gRNAs according to claim 27, a system according to claim 28, a nucleic acid according to any one of claims 29 to 34, a plurality of nucleic acids according to claim 35, or a particle according to claim 36, and at least one pharmaceutically acceptable excipient.

38. 38. A type II Cas protein according to any one of claims 1 to 13, a gRNA according to any one of claims 14 to 26, a combination of gRNAs according to claim 27, a system according to claim 28, a nucleic acid according to any one of claims 29 to 34, a plurality of nucleic acids according to claim 35, a particle according to claim 36, or a pharmaceutical composition according to claim 37, for use in a method for editing a human genome sequence.

39. 39. The type II Cas protein, gRNA, gRNA combination, system, nucleic acid, multiple nucleic acids, particle, or pharmaceutical composition for use according to claim 38, wherein the human genomic sequence is a RHO genomic sequence, and optionally the RHO genomic sequence has a pathogenic mutation.

40. 39. The Type II Cas protein, gRNA, gRNA combination, system, nucleic acid, plurality of nucleic acids, particle, or pharmaceutical composition for use according to claim 38, wherein the human genomic sequence is a TRAC, B2M, PD1, or LAG3 genomic sequence, and optionally the human genomic sequence is in a T cell.

41. 36. An ex vivo human cell comprising a type II Cas protein according to any one of claims 1 to 13, a gRNA according to any one of claims 14 to 26, a combination of gRNAs according to claim 27, a system according to claim 28, a nucleic acid according to any one of claims 29 to 34, a plurality of nucleic acids according to claim 35, or a particle according to claim 36.