CAS9 effector proteins with improved stability

JP2024518793A5Pending Publication Date: 2025-07-22ASTRAZENECA AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023571310
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-27
Filing Date
2022-05-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Cas9 proteins are susceptible to degradation in cellular environments, limiting their stability and effectiveness in genome editing and other applications.

Method used

Attaching a first nuclear localization signal to the N-terminus and a second nuclear localization signal to the C-terminus of the Cas9 effector protein to enhance its nuclear transport and stability.

Benefits of technology

The enhanced nuclear localization signals improve the stability of Cas9 effector proteins, reducing degradation and maintaining significant activity, allowing for more effective site-specific modification of target sequences in eukaryotic cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Cas9 effector proteins with improved stability are provided. Cas9 effector protein embodiments have a first nuclear localization signal attached to the N-terminus and a second nuclear localization signal attached to the C-terminus. Cas9 systems are also provided that include Cas9 effector proteins with improved stability and guide polynucleotides that form complexes with the Cas9 effector proteins. Additionally, methods are provided for providing site-specific modification of target sequences in eukaryotic cells using Cas9 effector proteins.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure provides a Cas9 effector protein with improved stability. An embodiment of the Cas9 effector protein has a first nuclear localization signal attached to the N-terminus and a second nuclear localization signal attached to the C-terminus. The present disclosure also provides a Cas9 system comprising such a Cas9 effector protein and a guide polynucleotide complexed with the Cas9 effector protein. The present disclosure further provides a method of providing site-specific modification of a target sequence in a eukaryotic cell using the Cas9 effector protein. [Background technology]

[0002] The use of CRIPR / Cas gene editing technology has revolutionized biotechnology. The CRISPR-Cas9 gene editing system has been successfully used in a wide range of organisms and cell systems, both to induce double-strand break (DSB) formation in DNA using wild-type Cas9 protein, or to nick single DNA strands using a mutant protein called Cas9n / Cas9 D10A (see, for example, Non-Patent Document 1 and Non-Patent Document 2, each of which is incorporated herein by reference in its entirety). While DSB formation leads to the creation of small insertions and deletions (indels) that can disrupt gene function, Cas9n / Cas9 D10A nickase avoids indel creation (a result of repair via non-homologous end joining) while stimulating endogenous homologous recombination machinery. Thus, Cas9n / Cas9 D10A nickase can be used to insert DNA regions into genomes with high fidelity.

[0003] In addition to genome editing, the CRISPR system has multiple other applications, including regulating gene expression, genetic circuit construction, and functional genomics, among others (reviewed in Non-Patent Document 2).

[0004] Although the Cas9 protein has been shown to be effective in a variety of in vivo and in vitro applications, as a protein it is potentially prone to degradation, especially in the cellular environment. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Mali et al.,Science,339(6121):823-826(2013) [Non-Patent Document 2] Sander and Joung, Nature Biotechnology 32(4):347-355(2014) Summary of the Invention [Means for solving the problem]

[0006] The present disclosure relates to a Cas9 effector protein comprising: a) a first nuclear localization signal linked to the N-terminus of the Cas9 effector protein; and b) a second nuclear localization signal linked to the C-terminus of the Cas9 effector protein.

[0007] In some embodiments of the protein, the first nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal. In some embodiments, the second nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the second nuclear localization signal is a bisegmental nuclear localization signal. In some embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal.

[0008] In some embodiments, the monopartite nuclear localization signal is a nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. In some embodiments, the bipartite nuclear localization signal is a classical bipartite nuclear localization signal. In some embodiments, the first nuclear localization signal is a classical bipartite nuclear localization signal and the second nuclear localization signal is a nuclear localization signal of SV40 large T antigen.

[0009] In some embodiments of the protein, the first nuclear localization signal is directly linked to the Cas9 effector protein. In some embodiments, the first nuclear localization signal is linked to the Cas9 effector protein via a linker. In some embodiments, the second nuclear localization signal is directly linked to the Cas9 effector protein. In some embodiments, the second nuclear localization signal is linked to the Cas9 effector protein via a linker. In some embodiments, the linker is a peptide linker having 2 to 30 residues.

[0010] In some embodiments, the protein comprises two copies of the first nuclear localization signal. In some embodiments, the protein comprises three copies of the first nuclear localization signal. In some embodiments, the protein comprises two copies of the second nuclear localization signal. In some embodiments, the protein comprises three copies of the second nuclear localization signal.

[0011] In some embodiments, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. In some embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97. In some embodiments, the Cas9 effector protein comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5. In some embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 90% identical to SEQ ID NO: 98. In some embodiments, the Cas9 effector protein comprises a modified polypeptide of SEQ ID NO: 98, wherein one or more modifications are N1164R, N1265R, N1300R, N1412R, N347R, N651A, D1266R, D309R, D345R, D487R, D607R, Q1129R, Q1381A, Q1381A, Q1381R, Q661A, Q713R, Q734R, E1032G, E1032R, E1409A, E436R, E611R, E691R, E 697R, G1335R, L125R, L1264S, L1299S, K1031R, K490R, K615R, K656R, F636R, S1334A, S1334A, S1334R, S1380R, S1410R, S1413R, S634R, S638R, S711R, S1006R, S1017R, T1267A, T1267R, T551R, Y1338A, Y1338R, V1273S, V1274S, V486R, V644R, V736R, and V736Y.

[0012] The present disclosure also relates to a CRISPR-Cas system comprising: a) a Cas9 effector protein comprising: i) a first nuclear localization signal attached to an N-terminus of the Cas9 effector protein; and ii) a second nuclear localization signal attached to a C-terminus of the Cas9 effector protein; and b) a guide polynucleotide comprising a guide sequence and complexed with the Cas9 effector protein, wherein the guide sequence is capable of hybridizing to a target sequence in a eukaryotic cell.

[0013] The disclosure further relates to a CRISPR-Cas system comprising: a) a nucleic acid sequence encoding a Cas9 effector protein, the Cas9 effector protein comprising: i) a first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; and ii) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein; and b) a nucleic acid sequence encoding a guide polynucleotide comprising a guide sequence and complexed with the Cas9 effector protein, the guide sequence being capable of hybridizing to a target sequence in a eukaryotic cell.

[0014] In some embodiments of the system, the nucleotide sequences of (a) and (b) are under the control of a eukaryotic promoter. In some embodiments, the nucleic acid sequences of (a) and (b) are in a single vector.

[0015] The disclosure further relates to a CRISPR-Cas system comprising one or more vectors comprising: a) a regulatory element operably linked to one or more nucleotide sequences encoding a Cas9 effector protein, the Cas9 effector protein comprising: i) a first nuclear localization signal linked to the N-terminus of the Cas9 effector protein; and ii) a second nuclear localization signal linked to the C-terminus of the Cas9 effector protein; and b) a guide polynucleotide comprising a guide sequence and complexed with the Cas9 effector protein, the guide sequence being capable of hybridizing to a target sequence in a eukaryotic cell.

[0016] In some embodiments of the system, the regulatory elements are eukaryotic regulatory elements.

[0017] In some embodiments of the system, the first nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal. In some embodiments, the second nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the second nuclear localization signal is a bisegmental nuclear localization signal. In some embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the first nuclear localization signal and the second nuclear localization signal are each a bisegmental nuclear localization signal.

[0018] In some embodiments, the monopartite nuclear localization signal is a nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. In some embodiments, the bipartite nuclear localization signal is a classical bipartite nuclear localization signal. In some embodiments, the first nuclear localization signal is a classical bipartite nuclear localization signal and the second nuclear localization signal is a nuclear localization signal of SV40 large T antigen.

[0019] In some embodiments of the system, the first nuclear localization signal is directly linked to the Cas9 effector protein. In some embodiments, the first nuclear localization signal is linked to the Cas9 effector protein via a linker. In some embodiments, the second nuclear localization signal is directly linked to the Cas9 effector protein. In some embodiments, the second nuclear localization signal is linked to the Cas9 effector protein via a linker. In some embodiments, the linker is a peptide linker having 2 to 30 residues.

[0020] In some embodiments of the system, the protein comprises two copies of the first nuclear localization signal. In some embodiments, the protein comprises three copies of the first nuclear localization signal. In some embodiments, the protein comprises two copies of the second nuclear localization signal. In some embodiments, the protein comprises three copies of the second nuclear localization signal.

[0021] In some embodiments of the system, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. In some embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97. In some embodiments, the Cas9 effector protein comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5. In some embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 90% identical to SEQ ID NO: 98.

[0022] In some embodiments of the system, the guide polynucleotide is RNA. In some embodiments, the guide sequence is 19-30 bases in length. In some embodiments, the guide sequence is 19-25 bases in length. In some embodiments, the guide sequence is 21-26 bases in length. In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence.

[0023] In some embodiments of the system, the Cas9 effector protein generates a sticky end. In some embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 1-10 nucleotides. In some embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 2-6 nucleotides. In some embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 3-5 nucleotides.

[0024] The present disclosure provides a eukaryotic cell comprising the protein described above.The present disclosure further provides a eukaryotic cell comprising the system described above.

[0025] The present disclosure provides a delivery particle comprising the protein described above. The present disclosure further provides a delivery particle comprising the system. In some embodiments of the delivery particle, the Cas9 effector protein and the guide polynucleotide are in a complex. In some embodiments, the complex further comprises a polynucleotide comprising a tracrRNA sequence. In some embodiments, the delivery particle further comprises a lipid, a sugar, a metal, or a protein.

[0026] The present disclosure provides a vesicle comprising the protein described above. The present disclosure further provides a vesicle comprising the system described above. In some embodiments of the vesicle, the Cas9 effector protein and the guide polynucleotide are present in a complex. In some embodiments, the vesicle further comprises a polynucleotide comprising a tracrRNA sequence. In some embodiments, the vesicle is an exosome or a liposome.

[0027] The present disclosure provides a viral vector comprising the proteins described above. The present disclosure further provides a viral vector comprising the system described above. In some embodiments, the viral vector further comprises a nucleic acid sequence encoding a tracrRNA sequence. In some embodiments, the viral vector is an adenovirus particle, an adeno-associated virus particle, or a herpes simplex virus particle.

[0028] The present disclosure also provides a method for providing site-specific modification of a target sequence in a eukaryotic cell, the method comprising: a) introducing into a cell; i) nucleotides encoding a Cas9 effector protein comprising: A) a first nuclear localization signal linked to the N-terminus of the Cas9 effector protein; and B) a second nuclear localization signal linked to the C-terminus of the Cas9 effector protein; and ii) nucleotides encoding a guide polynucleotide that forms a complex with the Cas9 effector protein and comprises a guide sequence, wherein the guide sequence is capable of hybridizing to a target sequence in a host polynucleotide; b) generating sticky ends in a host polynucleotide with a Cas9 effector protein and a guide polynucleotide; c) i) together with the sticky ends of (b), or ii) ligating the 3' end of the polynucleotide sequence of interest to one sticky end and the 5' end of the polynucleotide sequence to one sticky end; thereby modifying the target sequence.

[0029] In some embodiments of the method, the first nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal. In some embodiments, the second nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the second nuclear localization signal is a bisegmental nuclear localization signal. In some embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In some embodiments, the monosegmental nuclear localization signal is a nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. In some embodiments, the bisegmental nuclear localization signal is a classical bisegmental nuclear localization signal. In some embodiments, the first nuclear localization signal and the second nuclear localization signal are each a bipartite nuclear localization signal, hi some embodiments, the first nuclear localization signal is a classical bipartite nuclear localization signal and the second nuclear localization signal is an SV40 large T antigen nuclear localization signal.

[0030] In some embodiments of the method, the first nuclear localization signal is directly linked to the Cas9 effector protein. In some embodiments, the first nuclear localization signal is linked to the Cas9 effector protein via a linker. In some embodiments, the second nuclear localization signal is directly linked to the Cas9 effector protein. In some embodiments, the second nuclear localization signal is linked to the Cas9 effector protein via a linker. In some embodiments, the linker is a peptide linker having 2-30 residues.

[0031] In some embodiments of the method, the protein comprises two copies of the first nuclear localization signal. In some embodiments, the protein comprises three copies of the first nuclear localization signal. In some embodiments, the protein comprises two copies of the second nuclear localization signal. In some embodiments, the protein comprises three copies of the second nuclear localization signal.

[0032] In some embodiments of the method, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. In some embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97. In some embodiments, the Cas9 effector protein comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5.

[0033] In some embodiments of the method, the guide polynucleotide is RNA. In some embodiments, the guide polynucleotide is 19-30 bases in length. In some embodiments, the guide polynucleotide is 19-25 bases in length. In some embodiments, the guide polynucleotide is 21-26 bases in length. In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence.

[0034] In some embodiments of the method, the Cas9 effector protein generates a sticky end. In some embodiments, the sticky end comprises a single stranded polynucleotide overhang of 1-10 nucleotides. In some embodiments, the sticky end comprises a single stranded polynucleotide overhang of 2-6 nucleotides. In some embodiments, the sticky end comprises a single stranded polynucleotide overhang of 3-5 nucleotides. In embodiments, the sticky end is blunt. In embodiments, the sticky end has a 5' single stranded polynucleotide overhang. In embodiments, the sticky end has a 3' single stranded polynucleotide overhang.

[0035] In some embodiments of the methods, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the eukaryotic cell is a plant cell.

[0036] In some embodiments of the method, the modification is a deletion of at least a portion of the target sequence. In some embodiments, the modification is a mutation of the target sequence. In some embodiments, the modification is an insertion of a sequence of interest into the target sequence.

[0037] The disclosure also provides a method of reducing degradation of a Cas9 effector protein in a cell, the method comprising: a) attaching a first nuclear localization signal to the N-terminus of the Cas9 effector protein; and b) attaching a second nuclear localization signal to the C-terminus of the Cas9 effector protein.

[0038] In an embodiment of the method, the first nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the monosegmental nuclear localization signal is a nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. In an embodiment, the bisegmental nuclear localization signal is a classical bisegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a classical bipartite nuclear localization signal and the second nuclear localization signal is the SV40 large T antigen nuclear localization signal.

[0039] The following drawings form part of the present specification and are included to further demonstrate illustrative embodiments of certain aspects of the present invention. [Brief description of the drawings]

[0040] [Figure 1] Provided are amino acid sequences of Cas9 proteins that may be used in the Cas9 effector proteins described herein. [Diagram 2] Further amino acid sequences of Cas9 proteins that may be used in the Cas9 effector proteins described herein are provided. [Figure 3A] 1 is a western blot showing expression of MHCas9 in the absence or presence of the following inhibitors: 1) the proteasome inhibitor MG132 at a concentration of 5 μM as described in Example 1; 2) the lysosomal vATPase inhibitor bafilomycin A1 at a concentration of 20 nM; or 3) the nuclear export inhibitor leptomycin B at a concentration of 10 nM. [Figure 3B] 13 is a bar graph of quantification of Western blots using MAPK for normalization for each inhibitor and control. [Figure 4] 1 is a Western blot showing expression of the following Cas9 constructs: 1) 3xSV40-MHCas9; 2) MHCas9-NLSSV40; 3) 3XNLSSV40-MHCas9-NLSSV40; and 4) bpNLS-MHCas9-SLSSV40 (SpOT-ON) as described in Example 2. GFP expressed by the cloning vector is detected as a transfection control and tubulin is detected as a gel loading control. [Diagram 5] 5A shows a Western blot of the expression of SpOT-ON (5A) or bpNLS-SpCas9-NLSSV40 (5B) as described in Example 3. Cas9 constructs were tested in the absence or presence of the following inhibitors: 1) the proteasome inhibitor MG132 at a concentration of 5 μM; 2) the lysosomal vATPase inhibitor Bafilomycin A1 at a concentration of 20 nM; or 3) the nuclear export inhibitor Leptomycin B at a concentration of 10 nM. [Figure 6] 1 shows plots of titration of DNA cleavage activity at different sites using either SpOT-ON (MHCas9) or SpCas9 as described in Example 4. [Figure 7]1 shows a bar graph of DNA cleavage rate constant k when different protospacer lengths are used, as described in Example 5. [Figure 8] 8A and 8B show bar graphs plotting the average percentage of mutant reads among mapped reads for different protospacer lengths at the EMX1 site (8A) and CD34 site (8B) as described in Example 6. Bar graphs represent the average editing efficiency ± SD of HEK293T cells from n=3 different PBMC donors targeting CD34 or EMX1 as assessed by Amplicon-Seq and RIMA analysis. Allele frequencies <0.1% were excluded from the analysis. [Figure 9] 1 shows a plot of the percentage of modified reads at off-target sites for SpOT-ON and SPCas9 as described in Example 7. [Figure 10] 1 shows a bar graph plotting the cleavage rate constants of DNA substrates with mismatches at positions 1, 2, and 3 from the PAM described in Example 8. [Figure 11] Figure 1 shows a bar graph of the average percentage of mutant reads in the mapped reads at different positions for mismatch editing of EMX1 tested using the 23-nucleotide guide RNA described in Example 9. The bar graph represents the average editing efficiency ± SD of HEK293T cells of n=3 different PBMC donors targeting EMX1, assessed by Amplicon-Seq and RIMA analysis. Allele frequencies <0.1% were excluded from the analysis. [Figure 12A-B] 12A and 12B show qualitative analysis of DNA editing at the EMX locus (FIG. 12A) and the CD34 locus (FIG. 12B) as described in Example 10. [Figure 12C] 1 shows a comparative qualitative analysis of DNA repair following SpCas9 DNA cleavage at the CD34 locus. [Figure 13]1 is a bar graph showing the percentage of non-homologous end joining (NHEJ) knock-in at the CD34 locus for substrates with different overhangs as indicated. Experiments were performed as in Example 11. Plots are shown for both potential orientations of insertion, with dark grey representing forward (expected) insertions and light grey representing reverse insertions. [Figure 14] 1 is a bar graph showing the percentage of NHEJ knock-in at the STAT1 locus for substrates with different overhangs as indicated. Experiments were performed as in Example 11. Plots are shown for both potential orientations of insertion, with dark grey representing forward (expected) insertions and light grey representing reverse insertions. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0041] The present disclosure provides a Cas9 effector protein with improved stability that includes a nuclear localization signal at both the N-terminus and C-terminus of the Cas9 effector protein. The present disclosure also provides a system that includes a Cas9 effector protein with improved stability and a nucleic acid guide sequence that complexes with the Cas9 effector protein. The present disclosure also provides a method for site-specific modification of a target sequence in a eukaryotic cell using a Cas9 effector protein with improved stability. The present disclosure further provides a method for improving the stability of a Cas9 effector protein by attaching a nuclear localization signal to both the N-terminus and C-terminus of the protein.

[0042] Without wishing to be bound by theory, it is believed that the presence of an additional nuclear localization signal in a Cas9 effector protein enhances nuclear transport of the protein. This enhanced nuclear transport may reduce the time the Cas9 effector protein spends in the cytoplasm, where it may become a substrate for lysosomal degradation, a common degradation pathway for cytoplasmic proteins. In embodiments, the Cas9 effector proteins described herein have enhanced stability but retain significant Cas9 effector activity compared to Cas9 proteins without enhanced stability.

[0043] As used herein, a protein with "improved stability" refers to a protein that has a longer life span within an in vivo environment, such as a cell, or an in vitro environment. In some embodiments, a protein with "improved stability" may be more resistant to degradation in the environment, for example, by being more resistant to cleavage of bonds within the protein, by being less exposed to proteolytic agents, such as proteases, and / or by having less substrate for proteolytic agents. In embodiments, the "improved stability" is improved compared to the protein in its unmodified state. In embodiments, the "improved stability" of the Cas9 effector proteins described herein is improved compared to a Cas9 effector protein that does not have a nuclear localization signal. In embodiments, the "improved stability" of the Cas9 effector proteins described herein is improved compared to a Cas9 effector protein that has only one nuclear localization signal. In embodiments, the "improved stability" of the Cas9 effector proteins described herein is improved compared to a Cas9 effector protein having only one nuclear localization signal attached to the N-terminus of the Cas9 effector protein. In some embodiments, the stability of the Cas9 effector protein is increased by more than 10%, more than 20%, more than 30%, more than 40%, more than 50%, more than 60%, more than 70%, more than 80%, more than 90%, more than 100%, more than 120%, more than 140%, more than 160%, more than 180%, more than 200%, more than 300%, or more than 400% after 30 minutes of expression, after 60 minutes of expression, after 90 minutes of expression, after 120 minutes of expression, after 150 minutes of expression, after 120 minutes of expression, as measured by means known to one of skill in the art for determining the amount of a protein (e.g., Western blot) or by means known to one of skill in the art for determining the amount of a protein by measuring the activity of the protein (e.g., an activity assay described herein).

[0044] As used herein, "a" or "an" may mean one or more. As used herein, when used in conjunction with the word "comprising," the words "a" or "an" may mean one or more. As used herein, "another" or "further" may mean at least a second or more.

[0045] Throughout this application, the term "about" is used to indicate that a value includes the inherent variation of error for the method / device used to determine the value or the variation that exists between test subjects. Typically, the term "about" means to encompass approximately or less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20% or more variability, depending on the context. In embodiments, a person skilled in the art will understand the level of variation indicated by the term "about" depending on the context in which it is used herein. It should also be understood that the use of the term "about" includes the specifically recited value.

[0046] Although the use of the term "or" in the claims is used to mean "and / or" unless expressly stated to refer to alternatives only or the alternatives are not mutually exclusive, the present disclosure supports a definition that refers to alternatives only and "and / or."

[0047] As used herein, the terms "comprising" (and any variants or forms of comprising, e.g., "comprise" and "comprises"), "having" (and any variants or forms of having, e.g., "have" and "has"), "including" (and any variants or forms of including, e.g., "includes" and "include") or "containing" (and any variants or forms of containing, e.g., "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.

[0048] Use of the term "for example" and its corresponding abbreviation "eg" (whether italicized or not) means that the particular term cited is representative of examples and embodiments of the present disclosure that are not intended to be limited to the specific example referenced or cited, unless expressly stated otherwise.

[0049] As used herein, "between" refers to a range that includes both ends of the range. For example, a number between x and y explicitly includes numbers x and y, and all numbers between x and y.

[0050] Cas9 effector proteins In embodiments, the present disclosure provides a Cas9 effector protein with improved stability. In embodiments, the present disclosure provides a Cas9 effector protein comprising two or more nuclear localization signals. In embodiments, the present disclosure provides a Cas9 effector protein comprising: a) a first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; and b) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein.

[0051] As described herein, Cas proteins are components of CRISPR-Cas systems that can be used for, among other things, genome editing, gene regulation, gene circuit construction, and functional genomics. Cas1 and Cas2 proteins are believed to be universal to all currently identified CRISPR systems, while Cas3, Cas9, and Cas10 proteins are believed to be specific to type I, type II, and type III CRISPR systems, respectively.

[0052] Following the initial publication of the CRISPR-Cas9 system (a type II system), Cas9 variants have been identified in various bacterial species and a number have been functionally characterized (see, e.g., Chylinski et al., "Classification and evolution of type II CRISPR-Cas systems", Nucleic Acids Research 42(10):6091-6105 (2014), Ran et al., "In vivo genome editing using Staphylococcus aureus Cas9", Nature 520(7546):186-91 (2015), and Esvelt et al., "Orthogonal Cas9 proteins for RNA-guided gene regulation and editing", Nature Methods 10(11):1116-1121 (2013), the entire contents of which are incorporated herein by reference).

[0053] The present disclosure encompasses novel effector proteins of the CRISPR-Cas9 system with improved Cas9 stability. The terms "Cas9," "Cas9 protein," and "Cas9 effector protein" are used interchangeably herein to describe effector proteins that can provide sticky ends, blunt ends, or nicked dsDNA when used in the CRISPR-Cas9 system.

[0054] In an embodiment, the nuclear localization signal is a mono-gangly nuclear localization signal, a bi-gangly nuclear localization signal, or a combination thereof. A nuclear localization signal, also called a nuclear localization sequence or NLS, is an amino acid sequence that causes a protein having that sequence to be imported into a cell nucleus. In an embodiment, a mono-gangly nuclear localization signal is a signal having a single continuous sequence recognized for nuclear transport. In an embodiment, a bi-gangly nuclear localization signal is a signal having two sequences recognized for nuclear transport separated by a spacer sequence. Examples of both mono-gangly and bi-gangly nuclear localization signals are provided herein.

[0055] In an embodiment of the Cas9 effector protein, the first nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a bisegmental nuclear localization signal.

[0056] In an embodiment, the first and second nuclear localization signals can both be monosegmental, both be bisegmental, or a mixture of monosegmental and bisegmental. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal.

[0057] In embodiments, the monosegmental nuclear localization signal is a monosegmental nuclear localization signal known in the art. In embodiments, the monosegmental nuclear localization signal is one of the monosegmental nuclear localization signals listed in Table 1, or a combination thereof.

[0058] [Table 1]

[0059] In an embodiment, the bipartite nuclear localization signal is a bipartite nuclear localization signal known in the art. In an embodiment, the bipartite nuclear localization signal is a classical bipartite nuclear localization signal. In an embodiment, the bipartite nuclear localization signal is one of the bipartite nuclear localization signals listed in Table 2, or a combination thereof.

[0060] [Table 2]

[0061] In a protein embodiment, the first nuclear localization signal is a classical bipartite nuclear localization signal (SEQ ID NO: 7) and the second nuclear localization signal is the SV40 large T antigen nuclear localization signal (SEQ ID NO: 1).

[0062] In embodiments, the nuclear localization signal is attached to the Cas9 effector protein using methods standard in the art. In embodiments, the nucleic acid sequence encoding the nuclear localization signal is placed upstream and downstream of the nucleic acid sequence encoding the Cas9 effector protein using standard molecular biology methods such as restriction enzyme digestion and ligation, resulting in the formation of a nucleic acid encoding the Cas9 effector protein that includes nuclear localization signals at its N-terminus and C-terminus. This nucleic acid can then be expressed in a cell, for example a eukaryotic cell. In other embodiments, the Cas9 effector protein that includes nuclear localization signals at its N-terminus and C-terminus is fully or partially synthesized using solid phase protein synthesis.

[0063] In the embodiment of the protein, the first nuclear localization signal is directly linked to the Cas9 effector protein. In the embodiment, the first nuclear localization signal is linked to the Cas9 effector protein via a linker. In the embodiment, the second nuclear localization signal is directly linked to the Cas9 effector protein. In the embodiment, the second nuclear localization signal is linked to the Cas9 effector protein via a linker.

[0064] In embodiments where a linker is used, the linker is a peptide linker having 2-30 residues. In embodiments, the linker is a peptide linker having 2-20 residues. In embodiments, the linker is a peptide linker having 2-15 residues. In embodiments, the linker is a peptide linker having 2-10 residues. In embodiments, the linker is a peptide linker having 2-5 residues. In embodiments, the linker is a substituted or unsubstituted C2-C 20 It may be an alkyl, alkene, or alkynyl chain.

[0065] In embodiments, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its N-terminus. In embodiments, the Cas9 effector protein comprises two or more types of nuclear localization signal on its N-terminus. In embodiments, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its C-terminus. In embodiments, the Cas9 effector protein comprises two or more types of nuclear localization signal on its C-terminus.

[0066] In an embodiment, the protein comprises two copies of the first nuclear localization signal. In an embodiment, the protein comprises three copies of the first nuclear localization signal. In an embodiment, the protein comprises two copies of the second nuclear localization signal. In an embodiment, the protein comprises three copies of the second nuclear localization signal.

[0067] The Cas9 portion of the Cas9 protein, including the first and second nuclear localization signals, may be derived from any Cas9 effector domain known in the art. In an embodiment, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. Examples of suitable type II-B Cas9 proteins are described in WO 2019 / 099943, which is incorporated herein by reference. In an embodiment, a suitable type II-B Cas9 is capable of generating sticky ends. As described herein, a type II-B CRISPR system is identified, inter alia, by the presence of the cas4 gene on the cas operon, and the type II-B Cas9 protein is of the TIGR03031 TIGRFAM protein family. Thus, in an embodiment, the Cas9 portion is of the TIGR03031 TIGRFAM protein family. In an embodiment, the Cas9 portion comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5. In an embodiment, the site-specific nuclease comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-10. Type II-B CRISPR systems are used in, for example, Legionella pneumophila, Francisella novicida, gamma proteobacterium HTCC5015, Parasutterella excrementihominis, Sutterella wadsworthensis, Sulfurospirillum sp. SCADC, Ruminobacter sp.RM87, Burkholderiales bacterium 1_1_47, Bacteroidetes oral taxon strain 274 F0058, Wolinella succinogenes, Burkholderiales bacterium YL45, Ruminobacter amylophilus, Campylobacter sp. P0111, Campylobacter sp. RM9261, Campylobacter lanienae strain RM8001, Campylobacter lanienae strain P0121, Turicimonas muris, Legionella londiniensis, Salinivibrio sharmensis, Leptospira sp. isolate FW.030, Moritella sp. isolate NORP46, Endozoicomonas sp. S-B4-1U, Tamilnaduibacter salinus, Vibrio natriegens, Arcobacter skirrowii, Francisella philomiragia, Francisella hispaniensis It is found in bacterial species such as Bacillus subtilis, Bacillus hispaniensis, and Parendozoicomonas haliclonae.

[0068] In some embodiments, Cas9 can generate double-stranded polynucleotide cleavage, e.g., double-stranded DNA cleavage. In some embodiments, Cas9 can include one or more nuclease domains, such as RuvC and HNH, and can cleave double-stranded DNA. In some embodiments, Cas9 can include a RuvC domain and an HNH domain, each of which can cleave one strand of double-stranded DNA. In some embodiments, Cas9 generates blunt ends. In some embodiments, the RuvC and HNH of the Cas nuclease cleave each DNA strand at the same position, thereby generating blunt ends. In some embodiments, Cas9 generates sticky ends. In some embodiments, the RuvC and HNH of Cas9 cleave each DNA strand at a different position (i.e., cut with an "offset"), thereby generating sticky ends. As used herein, the terms "cohesive ends," "staggered ends," or "sticky ends" refer to nucleic acid fragments having strands of unequal length. In contrast to a "blunt end," a sticky end is generated by a staggered cut on a double-stranded nucleic acid (e.g., DNA). A sticky or cohesive end has a protruding single strand with unpaired nucleotides, or "overhang," e.g., a 3' or 5' overhang.

[0069] In embodiments, the term Cas9 refers to engineered Cas9 variants, such as deadCas9-FokI, Cas9n D10A -FokI and Cas9n H840A -FokI, etc. In an embodiment of the disclosure, the Cas9 effector protein comprises: a) a first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; and b) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein.

[0070] In some embodiments, the Cas9 (e.g., the Cas9 domain of the fusion protein) comprises a nuclease-inactivated Cas9 (e.g., a Cas9 lacking DNA cleavage activity; "dCas9") that retains RNA (gRNA) binding activity and thus can bind to a target site complementary to the gRNA. In embodiments, the fusion protein comprises a linker between the dCas9 domain and the transcription regulator domain. In embodiments, the dCas9 domain is fused to a transcription activator or repressor domain to form a dCas9 transcription regulator that can be directed to a specific target site via a complementary gRNA sequence. Examples of linkers are described herein. In embodiments, the fusion protein of the dCas9 domain and the transcription regulator domain has a nuclear localization signal attached to the N-terminus of the dCas9 domain and a nuclear localization signal attached to the C-terminus of the transcription regulator domain, as described herein.

[0071] In embodiments, the dCas9 domain is a dCas9 domain that functions as a roadblock to block transcription. In embodiments, the dCas9 domain can sterically block the transcription elongation of RNA polymerase.

[0072] In an embodiment, the dCas9 domain is fused to a VP64 transcription activation domain. In an embodiment, the dCas9 domain is modified using the SunTag gene activation system, which utilizes tandem repeats of the small peptide GCN4 to recruit multiple copies of a single-chain variable fragment fused to the transcription activator VP64. In an embodiment, the dCas9 domain is modified using the synergistic activation mediator (SAM) system, where dCas9 is fused to VP64 and the sgRNA is modified to contain two MS2 RNA aptamers that recruit the transcription activator p65 and the MS2 bacteriophage coat protein (MCP) fused to heat shock factor 1 (HSF1). In an embodiment, the dCas9 domain is modified with VP64-p65-Rta (VPR) for gene activation, where dCas9 is fused to a combination VPR transcription activation domain to amplify the activation effect. In embodiments, the dCas9 domain is engineered with scRNA to simultaneously activate and repress genes, where a hybrid RNA scaffold that binds sgRNA and RNA aptamers (e.g., MS2, com, PP7) can recruit RNA binding proteins (e.g., MCP, COM, PCP) that are bound to either transcriptional activators or repressors.

[0073] In embodiments, the dCas9 domain is modified with a chemically or light-regulated dimerization system, in which a chemically or light-induced dimerization agent (e.g., PYL1::ABI, GID::GAI, and PhyB::PIF) is fused to dCas9 and a transcriptional effector, respectively. In these embodiments, gene regulation can be triggered by the addition of the corresponding chemical (e.g., abscisic acid [ABA] or gibberellin [GA]) or light. In embodiments, the dCas9 domain is modified using a split dCas system or a receptor-coupled system:I / O molecular device.

[0074] In embodiments, dCas9 is a second or third generation transcription factor as described in Xu et al., “A CRISPR-dCas Toolbox for Genetic Engineering and Synthetic Biology,” J. Mol. Biol., 2019, 431:34-47 (incorporated herein by reference).

[0075] In an embodiment, the dCas9 is a dCas9 fusion protein for epigenome editing. In an embodiment, the dCas9 for epigenome editing is a dCas9 fusion protein as described in Xu et al., J. Mol. Biol., 2019, 431:34-47 (hereby incorporated by reference). In an embodiment, the dCas9 is fused to a methyltransferase, such as DNMT3A, DNMT3B, or DNMT3L. In an embodiment, the dCas9 is fused to a KRAB domain. In an embodiment, the dCas9 is fused to a DNA demethylase, such as TET1. In an embodiment, the dCas9 is fused to a histone methyltransferase, such as PRDM9 or DOT1L. In an embodiment, the dCas9 is fused to a histone demethylase, such as LSD1. In an embodiment, the dCas9 is fused to a histone acetyltransferase, such as p300. In embodiments, dCas9 is fused to a histone deacetylase, such as HDAC1, HDAC2, HDAC3, HDAC4, HDAC5, HDAC6, HDAC7, HDAC8, HDAC9, HDAC10, or HDAC11, or SIRT1, SIRT2, SIRT3, SIRT4, SIRT5, SIRT6, or SIRT7.

[0076] In an embodiment, the dCas9 is a dCas9 fusion protein for genomic imaging. In an embodiment, the dCas9 for genomic imaging is a dCas9 fusion protein as described in Xu et al., J. Mol. Biol., 2019, 431:34-47 (hereby incorporated by reference). In an embodiment, the dCas9 is fused to a fluorescent protein, such as a green fluorescent protein, a yellow fluorescent protein, a blue fluorescent protein, a cyan fluorescent protein, an orange fluorescent protein, or a red fluorescent protein.

[0077] In embodiments, dCas9 is a dCas9 fusion protein for base editing. In embodiments, dCas9 is fused to a cytosine base editor. In embodiments, dCas9 is fused to an adenine base editor. In embodiments, dCas9 is fused to a uracil base editor. In embodiments, dCas9 is fused to a cytidine deaminase. In embodiments, dCas9 is fused to an adenine deaminase. In embodiments, dCas9 is fused to a uracil DNA glycosylase.

[0078] In embodiments, the Cas9 domain is a Cas9 nickase fusion protein for base editing. As used herein, "Cas9 nickase" refers to a Cas9 protein that cleaves only one strand of a target DNA. In embodiments, the Cas9 nickase is fused to a cytosine base editor. In embodiments, the Cas9 nickase is fused to an adenine base editor. In embodiments, the Cas9 nickase is fused to a uracil base editor. In embodiments, the Cas9 nickase is fused to a cytidine deaminase. In embodiments, the Cas9 nickase is fused to an adenine deaminase. In embodiments, the Cas9 nickase is fused to a uracil DNA glycosylase. In embodiments, the Cas9 domain is a Cas9 nickase fusion for base editing as described in U.S. Patent Application Publication Nos. 2018 / 0312828, 2018 / 0237787, and 2020 / 0010835, each of which is incorporated herein by reference.

[0079] In an embodiment, the Cas9 domain is a prime-editing Cas9 nickase fusion protein. In an embodiment, the Cas9 nickase is fused to a reverse transcriptase. In an embodiment, the Cas9 domain is a prime-editing Cas9 nickase fusion as described in WO2020 / 191248 (hereby incorporated by reference).

[0080] In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide selected from any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2.

[0081] In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 71.

[0082] In protein embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 90% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 95% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 98% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 99% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 98.

[0083] [Table 3]

[0084] In some embodiments, the Cas9 effector protein is selected from the group consisting of R1336, R1389, R668, N1164, N1265, N1300, N1412, N347, N348, N562, N565, N618, N651, D1266, D309, D345, D487, D607, D30, Q1129, Q1381, Q624, Q661, Q713, Q734, E1032, E1409, E436, E610, E611, E691, E697, G1245, G1335, H777, I1242, L125, L1162, L1264, 98, which comprises an amino acid modification at one or more of positions L1299, K1031, K443, K490, K615, K656, F1035, F620, F636, F670, S1243, S1334, S1380, S1410, S1413, S634, S638, S711, S1006, S1017, T1267, T1333, T551, T639, T639, T640, T666, T897, Y1338, Y343, Y566, V1273, V1274, V486, V644, V660, V667, V736, or a combination thereof.

[0085] In some embodiments, the amino acid modifications are the following mutations R668A, N1164R, N1265R, N1300R, N1412R, N347R, N348R, N562R, N565R, N618R, N651A, N651R, D1266R, D309R, D345R, D487R, D607R, Q1129R, Q1381A, Q1381A, Q 1381R, Q624R, Q661A, Q661R, Q713R, Q734R, E1032G, E1032R, E1409A, E1409R, E436R, E610R, E 611R, E691R, E697R, G1245R, G1335R, H777A, I1242S, L125R, L125Y, L1162S, L1264S, L1299S, K1031R, K443R, K490R, K615R, K656R, F1035R, F620R, F636R, F670Y, S1243R, S1334A, S1334A, S1334R, S1380R, S1410R, S1413R, S634R, S638A, S638R, S711R, S1006R, S1017R, T1267A, T126 7R, T1333A, T1333R, T551R, T639A, T639R, T640R, T666R, T897R, Y1338A, Y1338R, Y343R, Y566R, V1273S, V1274S, V486R, V644R, V660R, V660Y, V667R, V667S, V736R, or V736Y.In some embodiments, the amino acid modifications are the following mutations N1164R, N1265R, N1300R, N1412R, N347R, N651A, D1266R, D309R, D345R, D487R, D607R, Q1129R, Q1381A, Q1381A, Q1381R, Q661A, Q713R, Q734R, E1032G, E1032R, E1409A, E436R, E611R, E691R, E697R, G1335R, L125R, L1 264S, L1299S, K1031R, K490R, K615R, K656R, F636R, S1334A, S1334A, S1334R, S1380R, S1410R, S1413R, S634R, S638R, S711R, S1006R, S1017R, T1267A, T1267R, T551R, Y1338A, Y1338R, V1273S, V1274S, V486R, V644R, V736R, or V736Y. In some embodiments, the amino acid modifications include one or more of the following mutations: N1265R, N1300R, N1412R, D1266R, E436R, G1335R, S1334R, S1380R, S1017R, T1267R, V736R, or V736Y.

[0086] In some embodiments, the amino acid modification increases the binding affinity between the Cas9 effector protein and DNA.

[0087] CRISPR-Cas system In embodiments, the present disclosure provides a CRISPR-Cas system comprising a Cas9 effector protein with improved stability.

[0088] Generally, CRISPR or CRISPR-Cas or CRISPR systems are characterized by elements (also referred to as protospacers for endogenous CRISPR systems) that promote the formation of a CRISPR complex at the site of the target sequence. In the context of a CRISPR complex, a "target sequence" refers to a sequence that a guide polynucleotide is designed to target, e.g., have complementarity, and hybridization between the target sequence and the guide polynucleotide promotes the formation of a CRISPR complex. A section of a guide polynucleotide whose complementarity to a target sequence may be important for cleavage activity is referred to as a guide sequence herein. A target sequence may comprise any polynucleotide, e.g., a DNA or RNA polynucleotide, and may be located within a target locus of interest. In an embodiment, the target sequence is located in the nucleus or cytoplasm of a cell. In an embodiment, the target sequence is located on a chromosome (TSC). In an embodiment, the target sequence is located on a vector (TSV).

[0089] In embodiments, the present disclosure provides a CRISPR-Cas system comprising: a) a Cas9 effector protein having: i) a first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; and ii) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein; A Cas9 effector protein comprising b) a guide polynucleotide comprising a guide sequence and complexed with a Cas9 effector protein, the guide sequence being capable of hybridizing to a target sequence in a eukaryotic cell.

[0090] In embodiments, the present disclosure provides a CRISPR-Cas system comprising: a) a Cas9 effector protein having: i) a first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; and ii) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein; A nucleic acid sequence encoding a Cas9 effector protein comprising: b) a nucleic acid sequence encoding a guide polynucleotide that comprises a guide sequence and that forms a complex with a Cas9 effector protein, wherein the guide sequence is capable of hybridizing to a target sequence in a eukaryotic cell.

[0091] In this system embodiment, the nucleotide sequences of (a) and (b) are under the control of the same promoter. In an embodiment, the nucleotide sequences of (a) and (b) are under the control of different promoters.

[0092] As used herein, "promoter", "promoter sequence" or "promoter region" refers to a DNA regulatory region / sequence that can bind RNA polymerase and participate in initiating transcription of downstream coding or non-coding sequences. In some examples of the present disclosure, the promoter sequence includes the transcription initiation site and extends upstream to include a minimum number of bases or elements used to initiate transcription at a level detectable above background. In embodiments, the promoter sequence includes the transcription initiation site and a protein binding domain responsible for binding RNA polymerase. Eukaryotic promoters often, but not always, contain "TATA" and "CAT" boxes. Various promoters, including inducible promoters, can be used to drive the various vectors of the present disclosure.

[0093] In an embodiment, the nucleotide sequences of (a) and (b) are under the control of a eukaryotic promoter. In an embodiment, the nucleotide sequences of (a) and (b) are under the control of two different eukaryotic promoters. In an embodiment, at least one of the eukaryotic promoters is a promoter active in human induced pluripotent stem cells. In an embodiment, at least one of the eukaryotic promoters is EF1alpha (EF1a). In an embodiment, at least one of the eukaryotic promoters is a human cytomegalovirus (CMV) promoter. In an embodiment, at least one of the eukaryotic promoters is the doxycycline regulated promoter TRE3G. In an embodiment, the nucleotide sequences of (a) and (b) are under the control of a bacterial promoter. In an embodiment, the nucleotide sequences of (a) and (b) are under the control of two different bacterial promoters. In an embodiment, the nucleotide sequences of (a) and (b) are under the control of a viral promoter. In an embodiment, the nucleotide sequences of (a) and (b) are under the control of two different viral promoters.

[0094] In an embodiment, the nucleic acid sequences of (a) and (b) are in a single vector. In an embodiment, the nucleic acid sequences of (a) and (b) are in separate vectors.

[0095] In embodiments, the present disclosure provides a CRISPR-Cas system comprising one or more vectors: a) a Cas9 effector protein having: i) a first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; and ii) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein; a regulatory element operably linked to one or more nucleotide sequences encoding a Cas9 effector protein comprising: b) a CRISPR-Cas system comprising one or more vectors comprising a guide polynucleotide comprising a guide sequence and complexed with a Cas9 effector protein, wherein the guide sequence is capable of hybridizing to a target sequence in a eukaryotic cell.

[0096] In this system embodiment, the regulatory element is a eukaryotic regulatory element. In this system embodiment, the regulatory element is a prokaryotic regulatory element.

[0097] In embodiments, the nucleotides encoding the Cas9 effector protein and the guide polynucleotide are present on a single vector. In embodiments, the nucleotides encoding the Cas9 effector protein, the guide polynucleotide (or nucleotides that can be transcribed into a guide polynucleotide), and the tracrRNA are present on a single vector. In embodiments, the nucleotides encoding the Cas9 effector protein, the guide polynucleotide (or nucleotides that can be transcribed into a guide polynucleotide), the tracrRNA, and the direct repeat sequence are present on a single vector. In embodiments, the vector is an expression vector. In embodiments, the vector is a mammalian expression vector. In embodiments, the vector is a human expression vector. In embodiments, the vector is a plant expression vector.

[0098] In embodiments, the nucleotides encoding the Cas9 effector protein and the guide polynucleotide are a single nucleic acid molecule. In embodiments, the nucleotides encoding the Cas9 effector protein, the guide polynucleotide, and the tracrRNA are a single nucleic acid molecule. In embodiments, the nucleotides encoding the Cas9 effector protein, the guide polynucleotide, the tracrRNA, and the direct repeat sequence are a single nucleic acid molecule. In embodiments, the single nucleic acid molecule is an expression vector. In embodiments, the single nucleic acid molecule is a mammalian expression vector. In embodiments, the single nucleic acid molecule is a human expression vector. In embodiments, the single nucleic acid molecule is a plant expression vector.

[0099] "Operably linked" means that the nucleotide of interest, i.e., the nucleotide encoding the Cas9 effector protein, is linked to a regulatory element in a manner that allows for expression of the nucleotide sequence. Thus, in an embodiment, the vector is an expression vector.

[0100] In embodiments, the regulatory element is a promoter. In embodiments, the regulatory element is a bacterial promoter. In embodiments, the regulatory element is a viral promoter. In embodiments, the regulatory element is a eukaryotic regulatory element, i.e., a eukaryotic promoter. In embodiments, the eukaryotic regulatory element is a mammalian promoter.

[0101] In any of the above system embodiments, the first nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a bisegmental nuclear localization signal.

[0102] In an embodiment, the first and second nuclear localization signals can both be monosegmental, both be bisegmental, or a mixture of monosegmental and bisegmental. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal.

[0103] In any of the above systems, the monosegmental nuclear localization signal is a monosegmental nuclear localization signal known in the art. In an embodiment, the monosegmental nuclear localization signal is one of the monosegmental nuclear localization signals listed in Table 1 above (SEQ ID NOs: 1-6), or a combination thereof.

[0104] In any of the above systems, the bi-section nuclear localization signal is a bi-section nuclear localization signal known in the art. In an embodiment, the bi-section nuclear localization signal is a classical bi-section nuclear localization signal. In an embodiment, the bi-section nuclear localization signal is one of the bi-section nuclear localization signals listed in Table 2 above (SEQ ID NOs: 7-9), or a combination thereof.

[0105] In an embodiment of any of the above systems, the first nuclear localization signal is a classical bipartite nuclear localization signal (SEQ ID NO: 7) and the second nuclear localization signal is the SV40 large T antigen nuclear localization signal (SEQ ID NO: 1).

[0106] In any embodiment of the above system, the first nuclear localization signal is directly linked to the Cas9 effector protein. In an embodiment, the first nuclear localization signal is linked to the Cas9 effector protein via a linker. In an embodiment, the second nuclear localization signal is directly linked to the Cas9 effector protein. In an embodiment, the second nuclear localization signal is linked to the Cas9 effector protein via a linker.

[0107] In embodiments where a linker is used, the linker is a peptide linker having 2-30 residues. In embodiments, the linker is a peptide linker having 2-20 residues. In embodiments, the linker is a peptide linker having 2-15 residues. In embodiments, the linker is a peptide linker having 2-10 residues. In embodiments, the linker is a peptide linker having 2-5 residues. In embodiments, the linker is a substituted or unsubstituted C2-C 20 It may be an alkyl, alkene, or alkynyl chain.

[0108] In an embodiment of any of the above systems, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its N-terminus. In an embodiment, the Cas9 effector protein comprises two or more types of nuclear localization signal on its N-terminus. In an embodiment, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its C-terminus. In an embodiment, the Cas9 effector protein comprises two or more types of nuclear localization signal on its C-terminus.

[0109] In an embodiment of any of the above systems, the protein comprises two copies of the first nuclear localization signal. In an embodiment, the protein comprises three copies of the first nuclear localization signal. In an embodiment, the protein comprises two copies of the second nuclear localization signal. In an embodiment, the protein comprises three copies of the second nuclear localization signal.

[0110] In any of the above system embodiments, the Cas9 portion of the Cas9 protein comprising the first and second nuclear localization signals may be derived from any Cas9 effector domain known in the art. In an embodiment, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. Examples of suitable type II-B Cas9 proteins are described above. In an embodiment, the Cas9 portion comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5. In an embodiment, the site-specific nuclease comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-10.

[0111] In an embodiment of any of the above systems, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97 shown in Figures 1 and 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to any one of SEQ ID NOs: 10-97 shown in Figures 1 and 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to any one of SEQ ID NOs: 10-97 shown in Figures 1 and 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to any one of SEQ ID NOs: 10-97 shown in Figures 1 and 2. In an embodiment, the Cas9 effector protein comprises a polypeptide selected from any one of SEQ ID NOs: 10-97 shown in Figures 1 and 2.

[0112] In embodiments of any of the above systems, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 71.

[0113] In embodiments of any of the above systems, the Cas9 effector protein comprises a polypeptide sequence at least 90% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 95% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 98% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 99% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 98.

[0114] In other embodiments of any of the above systems, the Cas9 portion of the Cas9 effector protein comprises dCas9, i.e., an inactivated or "dead" Cas9 that lacks DNA double-strand cleavage activity. In embodiments, dCas9 can be fused to other activity domains, such as transcription factors, epigenetic regulatory proteins, or fluorescent proteins, as described elsewhere herein. In embodiments in which dCas9 is fused to another activity domain, the nuclear localization signals described herein are present at the N- and C-termini of the entire Cas9 effector protein construct.

[0115] In other embodiments of any of the above systems, the Cas9 portion of the Cas9 effector protein comprises a Cas9 nickase, a Cas9 protein that cleaves only one strand of a DNA duplex. In embodiments, the Cas9 nickase can be fused to another activity domain, such as a transcription factor, an epigenetic regulatory protein, or a fluorescent protein, as described elsewhere herein. In embodiments in which the Cas9 nickase is fused to another activity domain, the nuclear localization signals described herein are present at the N- and C-termini of the entire Cas9 effector protein construct.

[0116] The systems and methods described herein may include a guide polynucleotide. In embodiments, the guide polynucleotide is an RNA. The RNA that binds to CRISPR-Cas9 components and targets them to a specific location within the target DNA is referred to herein as a "guide RNA," "gRNA," or "small molecule guide RNA," and may also be referred to herein as a "DNA-targeting RNA." A guide polynucleotide, e.g., a guide RNA, comprises at least two nucleotide segments: at least one "DNA-binding segment" and at least one "polypeptide-binding segment." By "segment" is meant a portion, section, or region of a molecule, e.g., a contiguous stretch of nucleotides of a guide polynucleotide molecule. The definition of "segment" is not limited to a particular number of total base pairs, unless specifically defined otherwise.

[0117] In embodiments, the DNA binding segment of guide polynucleotide hybridizes with the target sequence in eukaryotic cells, but does not hybridize with the sequence in bacterial cells.As used herein, the sequence in bacterial cells refers to the polynucleotide sequence derived from bacterial organisms, i.e., the naturally occurring bacterial polynucleotide sequence, or the sequence derived from bacteria.For example, the sequence can be bacterial chromosome or bacterial plasmid, or any other polynucleotide sequence naturally found in bacterial cells.

[0118] In embodiments, the polypeptide binding segment of the guide polynucleotide binds to a Cas9 effector protein with improved stability as described herein.

[0119] In an embodiment, the guide polynucleotide is 10 to 150 nucleotides. In an embodiment, the guide polynucleotide is 20 to 120 nucleotides. In an embodiment, the guide polynucleotide is 30 to 100 nucleotides. In an embodiment, the guide polynucleotide is 40 to 80 nucleotides. In an embodiment, the guide polynucleotide is 50 to 60 nucleotides. In an embodiment, the guide polynucleotide is 10 to 35 nucleotides. In an embodiment, the guide polynucleotide is 15 to 30 nucleotides. In an embodiment, the guide polynucleotide is 20 to 25 nucleotides.

[0120] The guide polynucleotide, e.g., guide RNA, can be introduced into the target cell as an isolated molecule, e.g., an RNA molecule, or is introduced into the cell using an expression vector containing DNA encoding the guide polynucleotide, e.g., the guide RNA.

[0121] The "DNA-binding segment" (or "DNA-targeting sequence") of a guide polynucleotide, e.g., a guide RNA, comprises a nucleotide sequence that is complementary to a specific sequence within the target DNA.

[0122] A guide polynucleotide, e.g., a guide RNA, of the present disclosure may comprise a polypeptide binding sequence / segment. The polypeptide binding segment (or "protein binding sequence") of a guide polynucleotide, e.g., a guide RNA, interacts with a polynucleotide binding domain of a Cas protein of the present disclosure. Such polypeptide binding segments or sequences are known to those of skill in the art, for example, they are disclosed in U.S. Patent Application Publication Nos. 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899, and 2015 / 0071906, the disclosures of which are incorporated herein in their entireties.

[0123] In some embodiments, the polypeptide binding segment is modified to improve binding to the polypeptide of the invention. Methods for modifying polypeptide binding segments to improve binding are described in Riesenberg et al. (Nature Communications, 2021) and references therein. Optimized polypeptide binding segments of guide RNA suitable for SEQ ID NO: 98 are shown in Table 3 as SEQ ID NO: 100-107. SEQ ID NO: 99 is a polypeptide binding segment sequence suitable for SEQ ID NO: 98 before optimization. In some embodiments, the guide RNA comprises a sequence selected from SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103, SEQ ID NO: 104, SEQ ID NO: 105, SEQ ID NO: 106, or SEQ ID NO: 107.

[0124] [Table 4]

[0125] [Table 5]

[0126] In an embodiment of the present disclosure, the Cas9 effector protein and the guide polynucleotide may form a complex. A "complex" is a group of two or more associated nucleic acids and / or polypeptides. In an embodiment, the complex is formed when all components of the complex are present together, i.e., a self-assembled complex. In an embodiment, the complex is formed through chemical interactions between different components of the complex, such as hydrogen bonds. In an embodiment, the guide polynucleotide forms a complex with the Cas9 effector protein through secondary structure recognition of the guide polynucleotide by the Cas9 effector protein. In an embodiment, the Cas9 effector protein is inactive, i.e., does not exhibit nuclease activity, until it forms a complex with the guide polynucleotide. Binding of the guide RNA induces a conformational change in the Cas9 effector protein, converting it from an inactive form to an active form, i.e., a catalytically active form.

[0127] In an embodiment of any of the above systems, the guide sequence is 19-30 bases in length. In an embodiment, the guide sequence is 19-25 bases in length. In an embodiment, the guide sequence is 21-26 bases in length.

[0128] In an embodiment of any of the above systems, the guide polynucleotide further comprises a tracrRNA sequence. The "tracrRNA" or trans-activated CRISPR-RNA forms an RNA duplex with the pre-crRNA or pre-CRISPR-RNA, which is then cleaved by the RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In an embodiment, the guide RNA comprises a crRNA / tracrRNA hybrid. In an embodiment, the tracrRNA component of the guide RNA activates the Cas9 effector protein.

[0129] In an embodiment of the system disclosed herein, the Cas9 effector protein, the guide polynucleotide, and the tracrRNA can form a complex.

[0130] In embodiments of any of the above systems, the Cas9 effector protein generates a sticky end. In embodiments, the sticky end generated by the Cas9 effector protein comprises a 5' overhang. In embodiments, the sticky end generated by the Cas9 effector protein comprises a 3' overhang. In embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 1-10 nucleotides. In embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 2-6 nucleotides. In embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 3-5 nucleotides.

[0131] In embodiments, the Cas9 effector protein prefers sticky ends with multiple nucleotides at the 5' end. In embodiments, the Cas9 effector protein prefers sticky ends with three nucleotides at the 5' end. In embodiments, the Cas9 effector protein prefers sticky ends with 2, 3, 4, 5, or 6 nucleotides at the 5' end. In embodiments, this preference is in contrast to the previously used S. pyogenes Cas9 (SpCas9), which prefers a single nucleotide 5' sticky end.

[0132] In embodiments, the presence of a single nucleotide 5' sticky end can be used to directly insert a nucleic acid of interest in a particular orientation. In embodiments, the presence of three nucleotides at the 5' sticky end can be used to directly insert a nucleic acid of interest in a particular orientation. In embodiments, the presence of two, three, four, five, or six nucleotides at the 5' sticky end can be used to directly insert a nucleic acid of interest in a particular orientation.

[0133] cell In embodiments, the present disclosure provides eukaryotic cells comprising the Cas9 effector proteins described herein. In embodiments, the present disclosure also provides eukaryotic cells comprising systems comprising the Cas9 effector proteins described herein.

[0134] In an embodiment, the eukaryotic cell is an animal or human cell. In an embodiment, the eukaryotic cell is a human, rodent, or bovine cell line or cell line. Examples of such cells, cell lines or cell lines include, but are not limited to, mouse myeloma (NS0) cell lines, Chinese hamster ovary (CHO) cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, NIH3T3, PC12, BHK (baby hamster kidney cells), VERO, SP2 / 0, YB2 / 0, Y0, C127, L cells, COS, e.g., COS1 and COS7, QC1-3, HEK-293, VERO, PER.C6, HeLA, EB1, EB2, EB3, oncolytic or hybridoma cell lines. In an embodiment, the eukaryotic cell is a CHO cell line. In an embodiment, the eukaryotic cell is a CHO cell. In embodiments, the cell is a CHO-K1 cell, a CHO-K1 SV cell, a DG44 CHO cell, a DUXB11 CHO cell, a CHOS, a CHO GS knockout cell, a CHO FUT8 GS knockout cell, a CHOZN or a CHO derived cell. A CHO GS knockout cell (e.g., a GSKO cell) is, for example, a CHO-K1 SV GS knockout cell. A CHO FUT8 knockout cell is, for example, Potelligent® CHOK1 SV (Lonza Biologics, Inc.). The eukaryotic cell may be an avian cell, cell line or cell strain, such as, for example, an EBx® cell, EB14, EB24, EB26, EB66, or EBvl3.

[0135] In an embodiment, the eukaryotic cell is a human cell. In an embodiment, the human cell is a stem cell. The stem cell can be, for example, a pluripotent stem cell, such as an embryonic stem cell (ESC), an adult stem cell, an induced pluripotent stem cell (iPSC), a tissue-specific stem cell (e.g., a hematopoietic stem cell), and a mesenchymal stem cell (MSC). In an embodiment, the human cell is any differentiated form of a cell described herein. In an embodiment, the eukaryotic cell is a cell derived from any primary cell in culture. In an embodiment, the cell is a stem cell or a stem cell line.

[0136] In embodiments, the eukaryotic cell is a hepatocyte, such as a human hepatocyte, an animal hepatocyte, or a non-parenchymal cell.For example, the eukaryotic cell can be an attached metabolic test human hepatocyte, an attached induction test human hepatocyte, an attached Qualyst Transporter Certified™ human hepatocyte, a suspension test human hepatocyte (for example, 10-donor and 20-donor pooled hepatocytes), a human hepatic Kupffer cell, a human hepatic stellate cell, a dog hepatocyte (for example, single and pooled beagle hepatocytes), a mouse hepatocyte (for example, CD-1 and C57BI / 6 hepatocytes), a rat hepatocyte (for example, Sprague-Dawley, Wistar Han, and Wistar hepatocytes), a monkey hepatocyte (for example, Cynomolgus or Rhesus hepatocytes), a cat hepatocyte (for example, domestic shorthair hepatocytes), and a rabbit hepatocyte (for example, New Zealand White hepatocytes).

[0137] In an embodiment, the eukaryotic cell is a plant cell. For example, the plant cell may be from a crop plant, such as cassava, maize, sorghum, wheat or rice. The plant cell may be from an algae, a tree or a vegetable. The plant cell may be from a monocotyledonous or dicotyledonous plant or from a crop or cereal plant, a productive plant, a fruit or a vegetable. For example, the plant cell may be from a tree, such as a citrus tree, e.g., an orange, grapefruit or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree, e.g., an almond or walnut or pistachio tree; a plant of the Solanum genus, i.e., potato; a plant of the Brassica genus, a Lactuca plant, a Spinacia plant; a plant of the Capsicum genus; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, and the like.

[0138] delivery particles In embodiments, the present disclosure provides a delivery particle comprising a Cas9 effector protein as described herein. In embodiments, the present disclosure also provides a delivery particle comprising a system comprising a Cas9 effector protein as described herein.

[0139] In embodiments where the delivery particle comprises a system described herein, the Cas9 effector protein and the guide polynucleotide are present in a complex. In embodiments, the complex further comprises a polynucleotide comprising a tracrRNA sequence.

[0140] In embodiments, the delivery particle is a lipid-based system, a liposome, a micelle, a microvesicle, an exosome, or a gene gun. In embodiments, the delivery particle comprises a Cas9 effector protein and a guide polynucleotide. In embodiments, the delivery particle comprises a Cas9 effector protein and a guide polynucleotide, and the Cas9 effector protein and the guide polynucleotide are in a complex. In embodiments, the delivery particle comprises a polynucleotide encoding a Cas9 effector protein, a polynucleotide encoding a guide polynucleotide, and a polynucleotide comprising a tracrRNA. In embodiments, the delivery particle comprises a Cas9 effector protein, a guide polynucleotide, and a tracrRNA. In embodiments, the delivery particle comprises a polynucleotide encoding one or more Cas9 effector proteins, a polynucleotide encoding one or more guide polynucleotides, and a polynucleotide encoding a tracrRNA.

[0141] In embodiments, the delivery particle further comprises lipid, sugar, metal, or protein. In embodiments, the delivery particle is a lipid envelope. In embodiments, the delivery particle is a sugar-based particle, such as GalNAc. In embodiments, the delivery particle is a nanoparticle. Examples of nanoparticles are described herein. Preparation of delivery particles is further described in U.S. Patent Application Publication Nos. 2011 / 0293703, 2012 / 0251560, and 2013 / 0302401; and U.S. Patent Nos. 5,543,158, 5,855,913, 5,895,309, 6,007,845, and 8,709,843, each of which is incorporated herein by reference in its entirety.

[0142] Vesicle In embodiments, the present disclosure provides a vesicle comprising a Cas9 effector protein as described herein. In embodiments, the present disclosure also provides a vesicle comprising a system comprising a Cas9 effector protein as described herein.

[0143] In embodiments where the vesicle comprises a system described herein, the Cas9 effector protein and the guide polynucleotide are present in a complex. In embodiments, the complex further comprises a polynucleotide comprising a tracrRNA sequence.

[0144] A "vesicle" is a small structure in a chamber with fluid enclosed by a lipid bilayer. Examples of vesicles are provided herein. In embodiments, the vesicle comprises a Cas9 effector protein and a guide polynucleotide. In embodiments, the vesicle comprises a Cas9 effector protein and a guide polynucleotide, where the Cas9 effector protein and the guide polynucleotide are in a complex. In embodiments, the vesicle comprises a polynucleotide encoding a Cas9 effector protein, a polynucleotide encoding a guide polynucleotide, and a polynucleotide comprising a tracrRNA. In embodiments, the vesicle comprises a Cas9 effector protein, a guide polynucleotide, and a tracrRNA. In embodiments, the vesicle comprises a polynucleotide encoding one or more Cas9 effector proteins, a polynucleotide encoding one or more guide polynucleotides, and a polynucleotide encoding a tracrRNA.

[0145] In embodiments, the vesicle is an exosome or liposome. In embodiments, the Cas9 effector protein is delivered to cells via exosomes. Exosomes are endogenous nanovesicles (i.e., having a diameter of about 30 to about 100 nm) that can transport RNA and proteins and deliver RNA to the brain and other target organs. Engineered exosomes for delivery of exogenous biological material into target organs have been described, for example, by Alvarez-Erviti et al., Nature Biotechnology 29:341 (2011), El-Andaloussi et al., Nature Protocols 7:2112-2116 (2012), and Wahlgren et al., Nucleic Acids Research 40(17):e130 (2012), each of which is incorporated herein by reference in its entirety.

[0146] In embodiments, Cas9 effector proteins are delivered to cells via liposomes. Liposomes are spherical vesicular structures with at least one lipid bilayer and can be used as vehicles for the administration of nutrients and pharmaceuticals. Liposomes are often composed of phospholipids, particularly phosphatidylcholine, but also other lipids, such as egg phosphatidylethanolamine. Types of liposomes include, but are not limited to, multilamellar vesicles, small unilamellar vesicles, large unilamellar vesicles, and spiral-wound vesicles. See, e.g., Spuch and Navarro, "Liposomes for Targeted Delivery of Active Agents against Neurodegenerative Diseases (Alzheimer's Disease and Parkinson's Disease), Journal of Drug Delivery 2011, Article ID 469679 (2011). Liposomes for delivery of biological materials, e.g., CRISPR-Cas components, have been described, e.g., by Morrissey et al., Nature Biotechnology 23(8):1002-1007 (2005), Zimmerman et al., Nature Letters 441:111-114 (2006), and Li et al., Gene Therapy 19:775-780 (2012), each of which is incorporated herein by reference in its entirety.

[0147] Viral Vectors In embodiments, the present disclosure provides viral vectors comprising the Cas9 effector proteins described herein. In embodiments, the present disclosure also provides viral vectors comprising systems comprising the Cas9 effector proteins described herein.

[0148] In embodiments where the viral vector comprises a system described herein, the Cas9 effector protein and the guide polynucleotide are present in a complex. In embodiments, the complex further comprises a polynucleotide comprising a tracrRNA sequence.

[0149] In embodiments, the viral vector is an adenovirus particle, an adeno-associated virus particle, or a herpes simplex virus particle. In embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus. Examples of viral vectors are provided herein. Viral transfection with adeno-associated virus (AAV) and lentivirus vectors (administration can be local, targeted, or systemic) has been used as a delivery method for in vivo gene therapy. In embodiments of the present disclosure, Cas effector proteins are expressed intracellularly by transduced cells.

[0150] In embodiments, the viral vector comprises a Cas9 effector protein and a guide polynucleotide. In embodiments, the viral vector comprises a Cas9 effector protein and a guide polynucleotide, and the Cas9 effector protein and the guide polynucleotide are in a complex. In embodiments, the viral vector comprises a polynucleotide encoding a Cas9 effector protein, a polynucleotide encoding a guide polynucleotide, and a polynucleotide comprising a tracrRNA. In embodiments, the viral vector comprises a Cas9 effector protein, a guide polynucleotide, and a tracrRNA. In embodiments, the viral vector comprises a polynucleotide encoding one or more Cas9 effector proteins, a polynucleotide encoding one or more guide polynucleotides, and a polynucleotide encoding a tracrRNA.

[0151] Methods for Providing Site-Specific Modification of a Target Sequence In embodiments, the present disclosure provides a method for providing site-specific modification of a target sequence in a eukaryotic cell; the method comprising: a) in cells; i) a Cas9 effector protein comprising: A) A first nuclear localization signal attached to the N-terminus of the Cas9 effector protein; B) a second nuclear localization signal attached to the C-terminus of the Cas9 effector protein; A nucleotide sequence encoding a Cas9 effector protein comprising: ii) introducing a nucleotide that forms a complex with a Cas9 effector protein and encodes a guide polynucleotide comprising a guide sequence, wherein the guide sequence is capable of hybridizing to a host polynucleotide; b) generating sticky ends in a host polynucleotide with a Cas9 effector protein and a guide polynucleotide; c) i) together with the sticky end of (b), or ii) ligating the 3' end of the polynucleotide sequence of interest to one sticky end and the 5' end of the polynucleotide sequence to one sticky end; thereby modifying the target sequence.

[0152] "Modifications" of the target sequence encompass single nucleotide substitutions, multiple nucleotide substitutions, insertions (ie, knock-ins) and deletions (ie, knock-outs) of the nucleic acid, frameshift mutations, and other nucleic acid modifications.

[0153] In embodiments, the modification is a deletion of at least a portion of the target sequence. The target sequence can be cleaved at two different sites to generate complementary sticky ends, and the complementary sticky ends can be religated, thereby removing the portion of the sequence between the two sites.

[0154] In an embodiment, the modification is a mutation of the target sequence. Site-specific mutagenesis in eukaryotic cells is achieved by the use of site-specific nucleases that promote homologous recombination of an exogenous polynucleotide template (also called a "donor polynucleotide" or "donor vector") containing the mutation of interest. In an embodiment, the sequence of interest (SoI) contains the mutation of interest.

[0155] In an embodiment, the modification is the insertion of a sequence of interest (SoI) into the target sequence. The SoI can be introduced as an exogenous polynucleotide template. In an embodiment, the exogenous polynucleotide template comprises a sticky end. In an embodiment, the exogenous polynucleotide template comprises a sticky end that is complementary to the sticky end in the target sequence.

[0156] The exogenous polynucleotide template can be of any suitable length, for example, about or at least about 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 500, or 1000 or more nucleotides in length. In embodiments, the exogenous polynucleotide template is complementary to a portion of the polynucleotide that comprises the target sequence. The exogenous polynucleotide template, when optimally aligned, overlaps with one or more nucleotides (e.g., about or at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides) of the target sequence. In embodiments, when the exogenous polynucleotide template and a polynucleotide comprising a target sequence are optimally aligned, the nearest nucleotide of the exogenous polynucleotide template is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 100, 1500, 2000, 2500, 5000, 10000 or more nucleotides of the target sequence.

[0157] In embodiments, the exogenous polynucleotide is DNA, such as a DNA plasmid, a bacterial artificial chromosome (BAC), a yeast artificial chromosome (YAC), a viral vector, a linear single-stranded or double-stranded piece of DNA, an oligonucleotide, a PCR fragment, a naked nucleic acid or a nucleic acid complexed with a delivery vehicle, such as a liposome.

[0158] In an embodiment, the exogenous polynucleotide is inserted into the target sequence using the cell's endogenous DNA repair pathway. Endogenous DNA repair pathways include the non-homologous end joining (NHEJ) pathway, the microhomology-mediated end joining (MMEJ) pathway, and the homology-directed repair (HDR) pathway. NHEJ, MMEJ, and HDR pathways repair double-stranded DNA breaks. In NHEJ, a homologous template is not required to repair the break in DNA. NHEJ repair can be error-prone, but errors are reduced when the DNA break contains a compatible overhang. NHEJ and MMEJ are DNA repair pathways that are mechanistically distinct due to the different subsets of DNA repair enzymes involved in each of them. Unlike NHEJ, which can be as accurate as error-prone, MMEJ is always error-prone and results in both deletions and insertions at the site under repair. MMEI-associated deletions result from microhomologies (2-10 base pairs) on either side of the double-strand break. In contrast, HDR requires a homologous template to direct repair, but HDR repair is typically high fidelity and low error prone. In embodiments, the error prone nature of NHEJ and MMEJ repair is exploited to introduce non-specific nucleotide substitutions in the target sequence. In embodiments, Cas9 effector protein cleaves the target sequence in a manner that facilitates HDR repair.

[0159] During the repair process, an exogenous polynucleotide template comprising an SoI can be introduced into the target sequence. In an embodiment, an exogenous polynucleotide template comprising an SoI flanked by upstream and downstream sequences is introduced into the cell, the upstream and downstream sequences sharing sequence similarity with either side of the site of integration in the target sequence. In an embodiment, the exogenous polynucleotide comprising an SoI comprises, for example, a mutant gene. In an embodiment, the exogenous polynucleotide comprises a sequence that is endogenous or exogenous to the cell. In an embodiment, the SoI comprises a polynucleotide that codes for a protein, or a non-coding sequence, such as a microRNA. In an embodiment, the SoI is operably linked to a regulatory element. In an embodiment, the SoI is a regulatory element. In an embodiment, the SoI comprises a resistance cassette, for example, a gene that confers resistance to an antibiotic. In an embodiment, the SoI comprises a mutation of the wild-type target sequence. In an embodiment, the SoI disrupts or corrects the target sequence by creating a frameshift mutation or a nucleotide substitution. In an embodiment, the SoI comprises a marker. Introduction of a marker into the target sequence can facilitate screening for targeted integration. In embodiments, the marker is a restriction site, a fluorescent protein, or a selection marker. In embodiments, the SoI is introduced as a vector containing the SoI.

[0160] The upstream and downstream sequences in the exogenous polynucleotide template are selected to promote homologous recombination between the target sequence and the exogenous polynucleotide. The upstream sequence is a nucleic acid sequence that shares sequence similarity with the sequence upstream of the target site for integration (i.e., the target sequence). Similarly, the downstream sequence is a nucleic acid sequence that shares sequence similarity with the sequence downstream of the target site for integration. Thus, in an embodiment, the exogenous polynucleotide template containing SoI inserts into the target sequence by homologous recombination at the upstream and downstream sequences. In an embodiment, the upstream and downstream sequences in the exogenous polynucleotide template have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the upstream and downstream sequences of the targeted genome sequence, respectively. In an embodiment, the upstream or downstream sequence has about 20 to 2000 base pairs, or about 50 to 1750 base pairs, or about 100 to 1500 base pairs, or about 200 to 1250 base pairs, or about 300 to 1000 base pairs, or about 400 to about 750 base pairs, or about 500 to 600 base pairs. In an embodiment, the upstream or downstream sequence has about 50, about 100, about 250, about 500, about 100, about 1250, about 1500, about 1750, about 2000, about 2250, or about 2500 base pairs.

[0161] In embodiments, the modification in the target sequence is the inactivation of the expression of the target sequence in cells.For example, when CRISPR complex binds to the target sequence, the target sequence is inactivated, so that the sequence is not transcribed and the encoded protein is not produced, or the sequence does not function as the wild-type sequence functions.For example, protein or microRNA coding sequence can be inactivated, so that the protein is not produced.

[0162] In embodiments, a regulatory sequence can be inactivated so that it no longer functions as a regulatory sequence. Examples of regulatory sequences include promoters, transcription terminators, enhancers and other regulatory elements as described herein. Inactivated target sequences can include deletion mutations (i.e., deletion of one or more nucleotides), insertion mutations (i.e., insertion of one or more nucleotides) or nonsense mutations (i.e., replacement of a single nucleotide with another nucleotide such that a stop codon is introduced). In embodiments, inactivation of a target sequence results in a "knockout" of the target sequence.

[0163] In an embodiment of the method, the first nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a bisegmental nuclear localization signal.

[0164] In embodiments of the method, the first and second nuclear localization signals can both be monosegmental, both be bisegmental, or a mixture of monosegmental and bisegmental. In embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In embodiments, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal. In embodiments, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal.

[0165] In an embodiment of the method, the monosegmental nuclear localization signal is a monosegmental nuclear localization signal known in the art. In an embodiment, the monosegmental nuclear localization signal is one of the monosegmental nuclear localization signals listed in Table 1 above (SEQ ID NOs: 1-6), or a combination thereof.

[0166] In an embodiment of the method, the bi-section nuclear localization signal is a bi-section nuclear localization signal known in the art. In an embodiment, the bi-section nuclear localization signal is a classical bi-section nuclear localization signal. In an embodiment, the bi-section nuclear localization signal is one of the bi-section nuclear localization signals listed in Table 2 above (SEQ ID NOs: 7-9), or a combination thereof.

[0167] In an embodiment of the method, the first nuclear localization signal is a classical bipartite nuclear localization signal (SEQ ID NO: 7) and the second nuclear localization signal is the SV40 large T antigen nuclear localization signal (SEQ ID NO: 1).

[0168] In an embodiment of the method, the first nuclear localization signal is directly linked to the Cas9 effector protein.In an embodiment, the first nuclear localization signal is linked to the Cas9 effector protein via a linker.In an embodiment, the second nuclear localization signal is directly linked to the Cas9 effector protein.In an embodiment, the second nuclear localization signal is linked to the Cas9 effector protein via a linker.

[0169] In embodiments where a linker is used, the linker is a peptide linker having 2-30 residues. In embodiments, the linker is a peptide linker having 2-20 residues. In embodiments, the linker is a peptide linker having 2-15 residues. In embodiments, the linker is a peptide linker having 2-10 residues. In embodiments, the linker is a peptide linker having 2-5 residues. In embodiments, the linker is a substituted or unsubstituted C2-C 20 It may be an alkyl, alkene, or alkynyl chain.

[0170] In an embodiment of the method, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its N-terminus. In an embodiment, the Cas9 effector protein comprises two or more types of nuclear localization signal on its N-terminus. In an embodiment, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its C-terminus. In an embodiment, the Cas9 effector protein comprises two or more types of nuclear localization signal on its C-terminus.

[0171] In an embodiment of the method, the protein comprises two copies of the first nuclear localization signal. In an embodiment, the protein comprises three copies of the first nuclear localization signal. In an embodiment, the protein comprises two copies of the second nuclear localization signal. In an embodiment, the protein comprises three copies of the second nuclear localization signal.

[0172] In an embodiment of the method, the Cas9 portion of the Cas9 protein comprising the first and second nuclear localization signals can be derived from any Cas9 effector domain known in the art. In an embodiment, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. Examples of suitable type II-B Cas9 proteins are described above. In an embodiment, the Cas9 portion comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5. In an embodiment, the site-specific nuclease comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-10.

[0173] In an embodiment of the method, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide selected from any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2.

[0174] In embodiments of the methods, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 71.

[0175] In embodiments of the methods, the Cas9 effector protein comprises a polypeptide sequence at least 90% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 95% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 98% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 99% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 98.

[0176] In other embodiments of the methods, the Cas9 portion of the Cas9 effector protein comprises dCas9, i.e., an inactivated or "dead" Cas9 that lacks DNA double-strand cleavage activity. In embodiments, dCas9 can be fused to other activity domains, such as transcription factors, epigenetic regulatory proteins, or fluorescent proteins, as described elsewhere herein. In embodiments in which dCas9 is fused to another activity domain, the nuclear localization signals described herein are present at the N- and C-termini of the entire Cas9 effector protein construct.

[0177] In other embodiments of the method, the Cas9 portion of the Cas9 effector protein comprises a Cas9 nickase, a Cas9 protein that cleaves only one strand of a DNA duplex. In embodiments, the Cas9 nickase can be fused to another activity domain, such as a transcription factor, an epigenetic regulatory protein, or a fluorescent protein, as described elsewhere herein. In embodiments in which the Cas9 nickase is fused to another activity domain, the nuclear localization signals described herein are present at the N- and C-termini of the entire Cas9 effector protein construct.

[0178] In embodiments, the methods include the use of a guide polynucleotide as described herein. In embodiments of the methods, the guide polynucleotide is RNA.

[0179] In an embodiment of any of the above systems, the guide sequence is 19-30 bases in length. In an embodiment, the guide sequence is 19-25 bases in length. In an embodiment, the guide sequence is 21-26 bases in length.

[0180] In embodiments of any of the above systems, the guide polynucleotide further comprises a tracrRNA sequence as described herein. In embodiments of the systems disclosed herein, the Cas9 effector protein, the guide polynucleotide, and the tracrRNA can form a complex.

[0181] In embodiments of the methods, the Cas9 effector protein generates a sticky end. In embodiments, the sticky end generated by the Cas9 effector protein comprises a 5' overhang. In embodiments, the sticky end generated by the Cas9 effector protein comprises a 3' overhang. In embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 1-10 nucleotides. In embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 2-6 nucleotides. In embodiments, the sticky end comprises a single-stranded polynucleotide overhang of 3-5 nucleotides.

[0182] In embodiments of the methods, the eukaryotic cell is an animal or human cell. In embodiments, the eukaryotic cell is an animal cell as described herein. In embodiments, the eukaryotic cell is a human cell. In embodiments, the eukaryotic cell is a human cell as described herein. In embodiments, the eukaryotic cell is a plant cell. In embodiments, the eukaryotic cell is a plant cell as described herein.

[0183] In an embodiment of the method, the modification is a deletion of at least a portion of the target sequence. In an embodiment, the modification is a mutation of the target sequence. In an embodiment, the modification is an insertion of a sequence of interest into the target sequence. In an embodiment, the modification is a modification described herein.

[0184] In embodiments of the method, the modification results in reduced off-target effects. In embodiments of the method, the modification results in reduced off-target effects compared to the off-target effects caused by S. pyogenes Cas9 (SpCas9).

[0185] The present disclosure also provides a method for providing site-specific modification of a target sequence in a eukaryotic cell with reduced off-target effects; the method comprising: a) introducing into a cell; i) nucleotides encoding a Cas9 effector protein comprising: A) a first nuclear localization signal linked to the N-terminus of the Cas9 effector protein; and B) a second nuclear localization signal linked to the C-terminus of the Cas9 effector protein; and ii) nucleotides encoding a guide polynucleotide that forms a complex with the Cas9 effector protein and comprises a guide sequence, wherein the guide sequence is capable of hybridizing to a target sequence in a host polynucleotide; b) generating sticky ends in a host polynucleotide with a Cas9 effector protein and a guide polynucleotide; c) i) together with the sticky ends of (b), or ii) ligating the 3' end of the polynucleotide sequence of interest to one sticky end and the 5' end of the polynucleotide sequence to one sticky end; thereby modifying the target sequence while reducing off-target effects.

[0186] In an embodiment of the method, the modification results in a reduced off-target effect compared to the off-target effect caused by S. pyogenes Cas9 (SpCas9).In an embodiment of the method, the modification results in a reduced off-target effect compared to the off-target effect caused by wild-type S. pyogenes Cas9 (SpCas9).

[0187] Methods for reducing degradation of Cas9 effector proteins In embodiments, the present disclosure provides a method of reducing degradation of a Cas9 effector protein in a cell, comprising: a) attaching a first nuclear localization signal to the N-terminus of a Cas9 effector protein; b) attaching a second nuclear localization signal to the C-terminus of the Cas9 effector protein; The present invention provides a method comprising:

[0188] In embodiments, the linking can be performed as described herein. In embodiments, a nucleic acid sequence encoding a nuclear localization signal is placed upstream and downstream of a nucleic acid sequence encoding a Cas9 effector protein using standard molecular biology methods such as restriction enzyme digestion and ligation, resulting in the formation of a nucleic acid encoding a Cas9 effector protein that includes nuclear localization signals at its N-terminus and C-terminus. This nucleic acid can then be expressed in a cell, for example, a eukaryotic cell. In other embodiments, a Cas9 effector protein that includes nuclear localization signals at its N-terminus and C-terminus is fully or partially synthesized using solid phase protein synthesis.

[0189] In an embodiment of the method, the first nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the first nuclear localization signal is a bisegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a monosegmental nuclear localization signal. In an embodiment, the second nuclear localization signal is a bisegmental nuclear localization signal.

[0190] In embodiments of the method, the first and second nuclear localization signals can both be monosegmental, both be bisegmental, or a mixture of monosegmental and bisegmental. In embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In embodiments, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal. In embodiments, the first nuclear localization signal is a monosegmental nuclear localization signal and the second nuclear localization signal is a monosegmental nuclear localization signal. In embodiments, the first nuclear localization signal is a bisegmental nuclear localization signal and the second nuclear localization signal is a bisegmental nuclear localization signal.

[0191] In an embodiment of the method, the monosegmental nuclear localization signal is a monosegmental nuclear localization signal known in the art. In an embodiment, the monosegmental nuclear localization signal is one of the monosegmental nuclear localization signals listed in Table 1 above (SEQ ID NOs: 1-6), or a combination thereof.

[0192] In an embodiment of the method, the bi-section nuclear localization signal is a bi-section nuclear localization signal known in the art. In an embodiment, the bi-section nuclear localization signal is a classical bi-section nuclear localization signal. In an embodiment, the bi-section nuclear localization signal is one of the bi-section nuclear localization signals listed in Table 2 above (SEQ ID NOs: 7-9), or a combination thereof.

[0193] In an embodiment of the method, the first nuclear localization signal is a classical bipartite nuclear localization signal (SEQ ID NO: 7) and the second nuclear localization signal is the SV40 large T antigen nuclear localization signal (SEQ ID NO: 1).

[0194] In an embodiment of the method, the first nuclear localization signal is directly linked to the Cas9 effector protein.In an embodiment, the first nuclear localization signal is linked to the Cas9 effector protein via a linker.In an embodiment, the second nuclear localization signal is directly linked to the Cas9 effector protein.In an embodiment, the second nuclear localization signal is linked to the Cas9 effector protein via a linker.

[0195] In embodiments where a linker is used, the linker is a peptide linker having 2-30 residues. In embodiments, the linker is a peptide linker having 2-20 residues. In embodiments, the linker is a peptide linker having 2-15 residues. In embodiments, the linker is a peptide linker having 2-10 residues. In embodiments, the linker is a peptide linker having 2-5 residues. In embodiments, the linker is a substituted or unsubstituted C2-C 20 It may be an alkyl, alkene, or alkynyl chain.

[0196] In an embodiment of the method, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its N-terminus. In an embodiment, the Cas9 effector protein comprises two or more types of nuclear localization signal on its N-terminus. In an embodiment, the Cas9 effector protein comprises two or more copies of a nuclear localization signal on its C-terminus. In an embodiment, the Cas9 effector protein comprises two or more types of nuclear localization signal on its C-terminus.

[0197] In an embodiment of the method, the protein comprises two copies of the first nuclear localization signal. In an embodiment, the protein comprises three copies of the first nuclear localization signal. In an embodiment, the protein comprises two copies of the second nuclear localization signal. In an embodiment, the protein comprises three copies of the second nuclear localization signal.

[0198] In an embodiment of the method, the Cas9 portion of the Cas9 protein comprising the first and second nuclear localization signals can be derived from any Cas9 effector domain known in the art. In an embodiment, the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system. Examples of suitable type II-B Cas9 proteins are described above. In an embodiment, the Cas9 portion comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-5. In an embodiment, the site-specific nuclease comprises a domain matching the TIGR03031 protein family with an E-value cutoff of 1E-10.

[0199] In an embodiment of the method, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2. In an embodiment, the Cas9 effector protein comprises a polypeptide selected from any one of SEQ ID NOs: 10-97 shown in FIG. 1 and FIG. 2.

[0200] In embodiments of the methods, the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 90% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 98% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises a polypeptide sequence having at least 99% identity to SEQ ID NO: 71. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 71.

[0201] In embodiments of the methods, the Cas9 effector protein comprises a polypeptide sequence at least 90% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 95% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 98% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises a polypeptide sequence at least 99% identical to SEQ ID NO: 98. In embodiments, the Cas9 effector protein comprises SEQ ID NO: 98.

[0202] All references cited herein, e.g., patents, patent applications, articles, textbooks, etc., and the references cited therein, are incorporated by reference in their entirety, unless they have already been cited. EXAMPLES

[0203] Example 1 - Cas9 is a substrate for lysosomal degradation The Cas9 protein from the gut metagenomic sequence MH0245 (MHCas9) - as described in WO2019099943 (incorporated herein by reference) - was cloned into a plasmid encoding three copies of the SV40 monopartite nuclear localization signal (NLS; SEQ ID NO:1) to form 3xSV40-MHCas9, a Cas9 protein with three SV40 NLSs attached at its N-terminus.

[0204] The plasmids were transfected into HEK293T cells. After the cells were cultured in DMEM+10% FBS medium for 24 hours, either of the following was added to the cell culture: 1) the proteasome inhibitor MG132 at a concentration of 5 μM; 2) the lysosomal vATPase inhibitor Bafilomycin A1 at a concentration of 20 nM; or 3) the nuclear export inhibitor Leptomycin B at a concentration of 10 nM. Untreated cells were used as controls. The cells were harvested and total protein was subsequently extracted.

[0205] The day before transfection, HEK293T cells were seeded at a density of 25,000 cells per well on a 96-well plate. 20 hours after seeding, the cells were transfected with the above plasmids. 48 hours after transfection, 100 μL of medium was added to the cells. 60 hours after transfection, the cells were harvested.

[0206] After recovery, total levels of 3xSV40-MHCas9 were analyzed using Western blots compared to blots of mitogen-activated protein kinases (MAPKs) for normalization of band intensity. The blots are shown in Figure 3A. Western blots were quantified and normalized protein expression was plotted as shown in Figure 3B.

[0207] As can be seen in Figure 3B, blocking lysosomal function with bafilomycin A1 led to increased levels of MHCas9, suggesting that MHCas9 is degraded within the lysosome.

[0208] Example 2 - Adding an NLS to Cas9 to prevent degradation MHCas9, as described in Example 1, was cloned into a plasmid encoding a nuclear localization signal to form four different Cas9 effector protein constructs: 1) 3xSV40-MHCas9, as described in Example 1; 2) MHCas9 with a single SV40 NLS at the C-terminus (MHCas9-NLSSV40); 3) MHCas9 with three SV40 NLSs at the N-terminus and a single SV40 NLS at the C-terminus (3XNLSSV40-MHCas9-NLSSV40); and 4) MHCas9 with a single bipartite NLS (SEQ ID NO: 7) at the N-terminus and a single SV40 NLS at the C-terminus (bpNLS-MHCas9-SLSSV40). The plasmid expressed green fluorescent protein (GFP), which was detected for normalization of transfection.

[0209] HEK293T cells were transfected and grown as described in Example 1, but were not treated with any inhibitors. Cells were harvested and Western blots were performed to detect tubulin as a gel loading control and to normalize the transfection amount using GFP. The blots are shown in Figure 4. As can be seen, adding an NLS to the C-terminus of Cas9 increases the stability of the protein in vivo. The bpNLS-MHCas9-SLSSV40 protein was selected for further studies and named SpOT-ON.

[0210] Example 3 - Cas9 constructs with NLS at both termini prevent lysosomal degradation To test the effect of NLS on other Cas9 proteins, S. pyogenes Cas9 (SpCas9) was cloned into the same vector as Construct 4 in Example 2, forming the bpNLS-SpCas9-NLSSV40 construct.

[0211] Cells expressing SpOT-ON and bpNLS-SpCas9-NLSSV40 were left untreated or grown in the presence of the inhibitors MG132, bafilomycin A1, or leptomycin at the same concentrations as used in Example 1. Cells were harvested and Western blots were performed using MAPK for band intensity normalization as described in Example 1. The blot for SpOT-ON is shown in Figure 5A and the blot for bpNLS-SpCas9-NLSSV40 is shown in Figure 5B.

[0212] As can be seen in the blot, the same levels of protein are detected whether the samples are treated with inhibitors or inhibitors that did not significantly slow the degradation in Example 1 (MG132 and leptomycin B). The enhanced nuclear targeting provided by the additional NLS signal may prevent the protein from being degraded in the cytoplasm. Furthermore, the addition of NLS signals at the N- and C-termini results in improved stability for both MHCas9 and SPCas9, suggesting that this technique can be generally applied to increase the stability of all types of Cas9 proteins.

[0213] Example 4 - SpOT-ON has the same DNA cleavage activity as unmodified Cas9 The DNA cleavage activity of SpOT-ON was compared to Cas9 lacking the NLS and found to be similar.

[0214] The cleavage activity of SpOT-On and Streptococcus pyogenes Cas9 protein (SpyCas9) was measured in vitro. Cas9 ribonucleoprotein (RNP) targeting the 20 nt protospacer was mixed with fluorescently labeled target DNA and a loading control lacking the protospacer adjacent motif (PAM). The reaction was incubated at 37 C and aliquots were taken at different time points, quenched and resolved using capillary electrophoresis. The fraction of digested DNA was quantified, normalized to the loading control and zero time point, and then plotted against time. Both enzymes digested the target DNA to the same extent. Analysis of the data (not shown) determined the following rate constants (k): SpyCas9, k = 0.224 and SpOT-ON, k = 0.004.

[0215] The results show that SpOT-On Cas9 can digest target DNA to the same extent as SpyCas9 in vitro, however, this cleavage occurs slower with SpOT-On Cas9 than with SpyCas9, as seen by the different rate constants.

[0216] Example 5 - SpOT-ON has the same editing activity as unmodified Cas9 The DNA gene editing activity of SpOT-ON was compared to Cas9 lacking the NLS and found to be similar.

[0217] Gene editing activity was compared for SpOT-ON and SpCas9. HEK293T cells were transfected with expression vectors expressing Cas9 variants and guide RNAs for HEK3, HEK4, EMX1, and FANCF. CD34 was used as the insertion site and STAT1 was used as the deletion site. Cells were cultured for 72 hours and then lysed to obtain DNA. Deep amplicon sequencing was performed to evaluate the editing that occurred. As can be seen in Figure 6, the editing efficiency was similar for SpOT-ON and SpCas9.

[0218] Example 6 - Determination of optimal protospacer length for SpOT-ON It was hypothesized that the slow DNA cleavage by SpOT-On Cas9 seen in Example 4 was caused by suboptimal sgRNA design, specifically the protospacer length. The in vitro cleavage experiment of Example 4 was repeated for SpOT-On Cas9 RNPs formed with a series of sgRNAs with various protospacer target sequence lengths targeting the same sequence.

[0219] To optimize the reaction efficiency, the cleavage activity of guide RNAs with various spacer lengths was determined and plotted using the method described in Example 4. The bars represent the calculated rate constants, and the error bars represent the standard error of fitting. The results are shown in Figure 7.

[0220] As can be seen in Figure 7, SpOT-On Cas9 targeting shorter protospacers (18-20 nt) results in less efficient DNA cleavage. However, RNPs with longer guides digest DNA 10-50 times faster, suggesting that a target sequence of at least 21 nucleotides is required for optimal activity of SpOT-On Cas9.

[0221] Further studies were performed to determine the optimal protospacer length in vivo. Cleavage activity was tested in vivo at two different target sites: EMX1 and CD34.

[0222] The day before transfection, HEK293T cells were seeded at a density of 25,000 cells per well on a 96-well plate. 20 hours after seeding, cells were transfected with the above plasmids. 48 hours after transfection, 100 μL of medium was added to the cells. 60 hours after transfection, cells were harvested using QuickExtract DNA extraction solution (Lucigen). Deep targeted amplicon sequencing was performed. The bar graph shown in Figure 8 shows the average percentage of mutant reads in the mapped reads. Replicate number n=3 (cells were separated into three stocks, then transfected and analyzed separately).

[0223] As can be seen in Figure 8, the optimal protospacer length for SpOT-ON is 19-23 nucleotides, with 21 nucleotides showing peak activity.

[0224] Example 7 - SpOT-ONs show reduced off-target DNA editing Gene editing activity was compared for SpOT-ON and SpCas9 using a method similar to that described in Example 5. HEK293T cells were transfected with expression vectors expressing Cas9 variants and guide RNAs for HEK3, HEK4, EMX1, and FANCF. Cells were cultured for 72 hours and then lysed to obtain DNA. Deep amplicon sequencing was performed to evaluate the editing that occurred. Analysis of off-target editing was performed using Crispresso2 pool analysis. The 14 off-target sites analyzed were those determined by Tsai et al. (Nat Biotechnol. 2015 Feb;33(2):187-197). The plot of off-target analysis is shown in Figure 9.

[0225] As can be seen in Figure 9, SpOT-ONs showed a significantly reduced percentage of editing at off-target sites compared to SpCas9, indicating that SpOT-ONs are better at discriminating between on-target and off-target sequences.

[0226] Example 8 – Analysis of off-target DNA editing Further studies were performed to investigate how mismatches in the substrate DNA affect the kinetics of DNA cleavage. To study the specificity of SpyCas9 and SpOT-ON, DNA substrates with single base pair substitutions in the target sequence were generated. The activity of the Cas9 enzyme was measured for perfectly matched and mismatched DNA substrates at positions 1, 2, and 3 from the PAM. The experiments were performed as described in the examples above, using optimal guides. The cleavage rate constants for each DNA substrate were calculated and plotted in Figure 10.

[0227] As shown in Figure 10, the Cas9 enzyme digests mismatched DNA substrates more slowly than perfectly matched DNA substrates. A mismatch immediately adjacent to the PAM significantly reduces the cleavage rate for both Spy Cas9 and SpOT-On Cas9. A more distal mismatch only slightly reduced Spy Cas9 activity, whereas SpOT-On Cas9 was inhibited by at least 10-fold. These data suggest that SpOT-On Cas9 is a more specific enzyme in vitro than Spy Cas9, which may explain the lower off-target genome editing activity of SpOT-On Cas9 in vivo.

[0228] Example 9 - Mismatch tolerance in vivo Mismatch tolerance of SpOT-ON in vivo, mismatch editing of EMX1 was tested using a 23-nucleotide guide RNA in HEK293T cells.

[0229] The day before transfection, HEK293T cells were seeded at a density of 25,000 cells per well on a 96-well plate. 20 hours after seeding, cells were transfected with the above plasmids. 48 hours after transfection, 100 μL of medium was added to the cells. 60 hours after transfection, cells were harvested using QuickExtract DNA extraction solution (Lucigen). Deep targeted amplicon sequencing was performed. The bar graph shown in Figure 11 shows the average percentage of mutant reads in the mapped reads. Replicate number n=3 (cells were separated into three stocks, then transfected and analyzed separately).

[0230] As can be seen in Figure 11, mismatches at positions 1-10 (1 being closest to the PAM) were not tolerated and resulted in no editing or very low editing (>0.7%). Mismatches between positions 11-21 showed moderate editing efficiencies of up to 20%. A mismatch at position 22 resulted in editing efficiencies similar to those of sgRNAs without mismatches (~55%).

[0231] Example 10 - DNA editing and analysis of cleavage sites For qualitative evaluation of DNA repair results, DNA editing at EMX and CD34 loci was further analyzed. Cells were seeded and grown as described in Example 9. NGS results of amplicon sequencing analyzed using RIMA are shown in Figure 12A for EMX1, Figure 12B for CD34, and Figure 12C for CD34 control with SpCas9.

[0232] These results demonstrated that SpOT-ON cleaved DNA to generate three-nucleotide overhangs in HEK293T cells.

[0233] Example 11 - Knock-in experiments Experiments were performed to evaluate the efficiency of directional non-homologous end joining (NHEJ) mediated knock-in of oligos with blunt ends or different overhangs at two target sites: CD34 and STAT1. DNA PK (M9831 / VX-984) inhibitor was added to half of the samples at a final concentration of 1 μM as an NHEJ inhibitor to demonstrate that NHEJ was occurring. SpOT-ON Cas9 was compared to SpCas9. Cells were seeded and grown as described in Example 9. DNA was analyzed using deep targeted amplicon sequencing.

[0234] Results for knock-in at the CD34 locus are shown in FIG. 13. Results for knock-in at the STAT1 locus are shown in FIG. 14. As can be seen in FIG. 13 and FIG. 14, SpOT-ON shows the best activity with a selection substrate with a 3 nucleotide 5' overhang (gray box), while SpCas9 shows the best activity with a selection substrate with a 1 nucleotide 5' overhang (white box). Plots are shown for both potential orientations of insertion, with dark grey representing forward (expected) insertion and light grey representing reverse insertion. As seen in the DNAPK column, DNA-PK inhibitor treatment completely inhibits oligo-donor insertion, proving that the knock-in is NHEJ-mediated.

[0235] As shown in the data, blunt-ended dsDNA oligos are integrated in both forward and reverse orientations after introducing a double-stranded break with SpCas9 or SpOT-ON Cas9. NHEJ-mediated insertion is also seen when short homology arms of 3 nucleotides are introduced at both ends of the oligo.

[0236] SpOT-ON Cas9 efficiently enables targeted integration of dsDNA oligos with 1 bp 5' overhangs, 3-nucleotide and 4-nucleotide overhangs, with the most efficient directed 3-nucleotide overhangs. SpCas9 integrates dsDNA with 1-nucleotide overhangs more efficiently, whereas dsDNA with 3-nucleotide or 4-nucleotide overhangs is less efficiently integrated. These results suggest that SpOT-ON Cas9 DSBs result primarily in 3-nucleotide 5' overhangs, whereas SpCas9 exhibits staggered cleavage with 1-nucleotide overhangs.

[0237] The ligation efficiency of dsDNA oligos with 3' overhangs (3 nucleotides) is less than 5% with SpCas9 and SpOT-ON Cas9, further supporting the theory that 5' overhangs are generated.

[0238] Example 12 Comparison of Various SpOT-ON Enzymes Experiments were performed to improve the genome editing efficiency of SpOT-ON by introducing single amino acid substitutions in residues near the DNA binding site. SpOT-ON residues near the DNA binding site were defined by modeling the SpOT-ON structure using Alpha Fold2 (Jumper et al, Nature 2021) (see Table Figure 16A). A subset of residues was selected for follow-up and mutagenesis. For example, negative charges (aspartic acid (ASP or D) or glutamic acid (GLU or E)) were removed and / or positive charges (arginine (ARG or R) or lysine (LYS or K)) were introduced to increase binding of the SpOT-ON complex to its target DNA. SpOT-ON variants were generated by mutagenesis and transfected into cells as described in Example 5. Table 3 shows the mutations that were tested for their effect on SpOT-ON activity. Overall, mutations near the PAM-interacting region (CTD domain) of SpOT-ON as well as in the REC3 domain show improved enzymatic activity.

[0239]

Table 6

[0240]

Table 7

[0241]

Table 8

[0242]

Table 9

[0243]

Table 10

Claims

1. A Cas9 effector protein comprising: a) a first nuclear localization signal (NLS) attached to the N-terminus of said Cas9 effector protein; and b) a second NLS attached to the C-terminus of said Cas9 effector protein. The Cas9 effector protein according to claim 1, which comprises the above.

2. The protein according to claim 1, wherein the first NLS is a monopartite NLS.

3. The protein according to claim 1, wherein the first NLS is a bipartite NLS.

4. The protein according to claim 1, wherein the second NLS is a monopartite NLS.

5. The protein according to claim 1, wherein the second NLS is a bipartite NLS.

6. The protein according to claim 1, wherein the first NLS is a bipartite NLS and the second NLS is a monopartite NLS.

7. The protein according to claim 1, wherein the first and / or the second NLS is a monopartite NLS, and the monopartite NLS is the NLS of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof.

8. The protein according to claim 1, wherein the first and / or the second NLS is a bipartite NLS, and the bipartite NLS is a classical bipartite NLS.

9. The protein according to claim 6, wherein the first NLS is a classical bipartite NLS and the second NLS is the NLS of SV40 large T antigen.

10. The protein according to claim 1, wherein the first NLS is directly bound to the Cas9 effector protein.

11. The protein according to claim 1, wherein the first NLS is bound to the Cas9 effector protein via a linker.

12. The protein according to claim 1, wherein the second NLS is directly bound to the Cas9 effector protein.

13. The protein according to claim 1, wherein the second nuclear localization signal is bound to the Cas9 effector protein via a linker.

14. The protein according to claim 1, wherein the first and / or the second nuclear localization signal is bound to the Cas9 effector protein via a linker, and the linker is a peptide linker having 2 to 30 residues.

15. The protein according to claim 1, wherein the protein comprises two copies of the first nuclear localization signal.

16. The protein according to claim 1, wherein the protein comprises three copies of the first nuclear localization signal.

17. The protein according to claim 1, wherein the protein comprises two copies of the second nuclear localization signal.

18. The protein according to claim 1, wherein the protein comprises three copies of the second nuclear localization signal.

19. The protein according to any one of claims 1 to 18, wherein the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10 to 97.

20. The protein according to any one of claims 1 to 18, wherein the Cas9 effector protein comprises a domain that matches the TIGR03031 protein family having an E-value cutoff of 1E-5.

21. The protein according to any one of claims 1 to 18, wherein the Cas9 effector protein comprises a polypeptide sequence that is at least 90% identical to SEQ ID NO:

98.

22. The protein according to claim 21, wherein the polypeptide contains one or more modifications selected from N1164R, N1265R, N1300R, N1412R, N347R, N651A, D1266R, D309R, D345R, D487R, D607R, Q1129R, Q1381A, Q1381A, Q1381R, Q661A, Q713R, Q734R, E1032G, E1032R, E1409A, E436R, E611R, E691R, E697R, G1335R, L125R, L1264S, L1299S, K1031R, K490R, K615R, K656R, F636R, S1334A, S1334A, S1334R, S1380R, S1410R, S1413R, S634R, S638R, S711R, S1006R, S1017R, T1267A, T1267R, T551R, Y1338A, Y1338R, V1273S, V1274S, V486R, V644R, V736R, and V736Y.

23. A CRISPR-Cas system comprising: a) A Cas9 effector protein comprising: i) A first nuclear localization signal bound to the N-terminus of the Cas9 effector protein; ii) A second nuclear localization signal bound to the C-terminus of the Cas9 effector protein; A Cas9 effector protein, and b) A guide polynucleotide comprising a guide sequence and forming a complex with the Cas9 effector protein, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell. A CRISPR-Cas system comprising the above.

24. A CRISPR-Cas system comprising: a) A nucleic acid sequence encoding a Cas9 effector protein comprising: i) A first nuclear localization signal bound to the N-terminus of the Cas9 effector protein; ii) A second nuclear localization signal bound to the C-terminus of the Cas9 effector protein; And b) A nucleic acid sequence encoding a guide polynucleotide comprising a guide sequence and forming a complex with the Cas9 effector protein, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell. A CRISPR-Cas system comprising the above.

25. The system according to claim 24, wherein the nucleotide sequences of (a) and (b) are under the control of a eukaryotic promoter. **Claim 26** The system according to claim 24, wherein the nucleic acid sequences of (a) and (b) are in a single vector. **Claim 27** A CRISPR-Cas system comprising one or more vectors, wherein: a) a Cas9 effector protein, wherein: i) a first nuclear localization signal linked to the N-terminus of the Cas9 effector protein; ii) a second nuclear localization signal linked to the C-terminus of the Cas9 effector protein; a regulatory element operably linked to one or more nucleotide sequences encoding the Cas9 effector protein, the regulatory element comprising the first and second nuclear localization signals; b) a guide polynucleotide comprising a guide sequence and forming a complex with the Cas9 effector protein, wherein the guide sequence is capable of hybridizing to a target sequence in a eukaryotic cell. A CRISPR-Cas system comprising one or more vectors. **Claim 28** The system according to claim 27, wherein the regulatory element is a eukaryotic regulatory element. **Claim 29** The system according to any one of claims 23 to 28, wherein the first nuclear localization signal is a monopartite nuclear localization signal. **Claim 30** The system according to any one of claims 23 to 28, wherein the first nuclear localization signal is a bipartite nuclear localization signal. **Claim 31** The system according to any one of claims 23 to 28, wherein the second nuclear localization signal is a monopartite nuclear localization signal. **Claim 32** The system according to any one of claims 23 to 28, wherein the second nuclear localization signal is a bipartite nuclear localization signal. **Claim 33** The system according to any one of claims 23 to 28, wherein the first nuclear localization signal is a bipartite nuclear localization signal and the second nuclear localization signal is a monopartite nuclear localization signal. **Claim 34** The system according to any one of claims 23 to 28, wherein the first nuclear localization signal and the second nuclear localization signal are each a bipartite nuclear localization signal. **Claim 35**: The system according to claim 27, wherein the first and / or the second nuclear localization signal is a monopartite nuclear localization signal, and the monopartite nuclear localization signal is a nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. **Claim 36**: The system according to claim 27, wherein the first and / or the second nuclear localization signal is a bipartite nuclear localization signal, and the bipartite nuclear localization signal is a classical bipartite nuclear localization signal. **Claim 37** The system according to claim 33, wherein the first nuclear localization signal is a classical bipartite nuclear localization signal, and the second nuclear localization signal is a nuclear localization signal of SV40 large T antigen. **Claim 38** The system according to any one of claims 23 to 28, wherein the first nuclear localization signal is directly bound to the Cas9 effector protein. **Claim 39** The system according to any one of claims 23 to 28, wherein the first nuclear localization signal is bound to the Cas9 effector protein via a linker. **Claim 40** The system according to any one of claims 23 to 28, wherein the second nuclear localization signal is directly bound to the Cas9 effector protein. **Claim 41** The system according to any one of claims 23 to 28, wherein the second nuclear localization signal is bound to the Cas9 effector protein via a linker. **Claim 42**: The system according to claim 27, wherein the first and / or the second nuclear localization signal is bound to the Cas9 effector protein via a linker, and the linker is a peptide linker having 2 to 30 residues. **Claim 43** The system according to any one of claims 23 to 28, wherein the protein comprises two copies of the first nuclear localization signal. **Claim 44** The system according to any one of claims 23 to 28, wherein the protein comprises three copies of the first nuclear localization signal. **Claim 45** The system according to any one of claims 23 to 28, wherein the protein comprises two copies of the second nuclear localization signal. **Claim 46** The system according to any one of claims 23 to 28, wherein the protein comprises three copies of the second nuclear localization signal. **Claim 47** The system according to any one of claims 23 to 28, wherein the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system.

48. The system according to any one of claims 23 to 28, wherein the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10 to 97.

49. The system according to any one of claims 23 to 28, wherein the Cas9 effector protein comprises a domain that matches the TIGR03031 protein family having an E-value cutoff of 1E-5.

50. The system according to any one of claims 23 to 28, wherein the Cas9 effector protein comprises a polypeptide sequence that is at least 90% identical to SEQ ID NO:

98.

51. The system according to any one of claims 23 to 28, wherein the guide polynucleotide is RNA.

52. The system according to claim 51, wherein the guide sequence is 19 to 30 bases in length.

53. The system according to claim 51, wherein the guide sequence is 19 to 25 bases in length.

54. The system according to claim 51, wherein the guide sequence is 21 to 26 bases in length.

55. The system according to any one of claims 23 to 28, wherein the guide polynucleotide further comprises a tracrRNA sequence.

56. The system according to any one of claims 23 to 28, wherein the Cas9 effector protein generates sticky ends.

57. The system according to claim 56, wherein the sticky ends comprise a single-stranded polynucleotide overhang of 1 to 10 nucleotides.

58. The system according to claim 56, wherein the sticky ends comprise a single-stranded polynucleotide overhang of 2 to 6 nucleotides.

59. The system according to claim 57, wherein the sticky ends comprise a single-stranded polynucleotide overhang of 3 to 5 nucleotides.

60. A eukaryotic cell comprising the protein according to any one of claims 1 to 18.

61. A eukaryotic cell comprising the system according to any one of claims 23 to 28.

62. A delivery particle comprising the protein according to any one of claims 1 to 18.

63. A delivery particle comprising the system according to any one of claims 23 to 28.

64. The delivery particle according to claim 63, wherein the Cas9 effector protein and the guide polynucleotide are present in a complex. **Claim 65** The delivery particle according to claim 64, wherein the complex further comprises a polynucleotide comprising a tracrRNA sequence. **Claim 66** The delivery particle according to claim 62, further comprising a lipid, a sugar, a metal, or a protein. **Claim 67** A vesicle comprising the protein according to any one of claims 1 to 18. **Claim 68** A vesicle comprising the system according to any one of claims 23 to 28. **Claim 69** The vesicle according to claim 68, wherein the Cas9 effector protein and the guide polynucleotide are present in a complex. **Claim 70** The vesicle according to claim 68, further comprising a polynucleotide comprising a tracrRNA sequence. **Claim 71** The vesicle according to claim 67, wherein the vesicle is an exosome or a liposome. **Claim 72** A viral vector comprising the protein according to any one of claims 1 to 18. **Claim 73** A viral vector comprising the system according to any one of claims 23 to 28. **Claim 74** The viral vector according to claim 73, further comprising a nucleic acid sequence encoding a tracrRNA sequence. **Claim 75** The viral vector according to claim 72, wherein the viral vector is an adenovirus particle, an adeno-associated virus particle, or a herpes simplex virus particle. **Claim 76** A method for providing site-specific modification of a target sequence in a eukaryotic cell, the method comprising: a) into the cell: i) a nucleotide encoding a Cas9 effector protein, wherein the Cas9 effector protein comprises: A) a first nuclear localization signal bound to the N-terminus of the Cas9 effector protein; and B) a second nuclear localization signal bound to the C-terminus of the Cas9 effector protein; and a nucleotide encoding a guide polynucleotide that forms a complex with the Cas9 effector protein and comprises a guide sequence that is capable of hybridizing to a host polynucleotide; and ii) introducing into the cell a nucleotide encoding a guide polynucleotide that forms a complex with the Cas9 effector protein and comprises a guide sequence that is capable of hybridizing to a host polynucleotide; b) generating sticky ends in the host polynucleotide by the Cas9 effector protein and the guide polynucleotide; c) i) together with the sticky ends of (b), or ii) ligating the 3' end of the polynucleotide sequence of interest to one sticky end and the 5' end of the said polynucleotide sequence of interest to one sticky end; thereby modifying the said target sequence, A method comprising. **Claim 77** The method according to claim 76, wherein the first nuclear localization signal is a monopartite nuclear localization signal. **Claim 78** The method according to claim 76, wherein the first nuclear localization signal is a bipartite nuclear localization signal. **Claim 79** The method according to claim 76, wherein the second nuclear localization signal is a monopartite nuclear localization signal. **Claim 80** The method according to claim 76, wherein the second nuclear localization signal is a bipartite nuclear localization signal. **Claim 81** The method according to claim 76, wherein the first nuclear localization signal is a bipartite nuclear localization signal and the second nuclear localization signal is a monopartite nuclear localization signal. **Claim 82** The method according to claim 76, wherein the first and / or the second nuclear localization signal is a monopartite nuclear localization signal, and the monopartite nuclear localization signal is the nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. **Claim 83** The method according to claim 76, wherein the first and / or the second nuclear localization signal is a bipartite nuclear localization signal, and the bipartite nuclear localization signal is a classical bipartite nuclear localization signal. **Claim 84** The method according to claim 76, wherein the first nuclear localization signal and the second nuclear localization signal are each a bipartite nuclear localization signal. **Claim 85** The method according to claim 81, wherein the first nuclear localization signal is a classical bipartite nuclear localization signal and the second nuclear localization signal is the nuclear localization signal of SV40 large T antigen. **Claim 86** The method according to claim 76, wherein the first nuclear localization signal is directly bound to the Cas9 effector protein. **Claim 87** The method according to claim 76, wherein the first nuclear localization signal is bound to the Cas9 effector protein via a linker. **Claim 88** The method according to claim 76, wherein the second nuclear localization signal is directly bound to the Cas9 effector protein. **Claim 89** The method according to claim 76, wherein the second nuclear localization signal is bound to the Cas9 effector protein via a linker.

90. The method according to claim 76, wherein the first and / or the second nuclear localization signal is bound to the Cas9 effector protein via a linker, and the linker is a peptide linker having 2 to 30 residues.

91. The method according to claim 76, wherein the protein comprises two copies of the first nuclear localization signal.

92. The method according to claim 76, wherein the protein comprises three copies of the first nuclear localization signal.

93. The method according to claim 76, wherein the protein comprises two copies of the second nuclear localization signal.

94. The method according to claim 76, wherein the protein comprises three copies of the second nuclear localization signal.

95. The method according to claim 76, wherein the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system.

96. The method according to any one of claims 76 to 95, wherein the Cas9 effector protein comprises a polypeptide sequence having at least 95% identity to any one of SEQ ID NOs: 10 to 97.

97. The method according to any one of claims 76 to 95, wherein the Cas9 effector protein comprises a domain that matches the TIGR03031 protein family having an E-value cutoff of 1E-5.

98. The method according to any one of claims 76 to 95, wherein the guide polynucleotide is RNA.

99. The method according to claim 98, wherein the guide polynucleotide is 19 to 30 bases in length.

100. The method according to claim 98, wherein the guide polynucleotide is 19 to 25 bases in length.

101. The method according to claim 98, wherein the guide polynucleotide is 21 to 26 bases in length.

102. The method according to any one of claims 76 to 95, wherein the guide polynucleotide further comprises a tracrRNA sequence.

103. The method according to any one of claims 76 to 95, wherein the Cas9 effector protein generates sticky ends.

104. The method according to any one of claims 76 to 95, wherein the sticky end comprises a single-stranded polynucleotide overhang of 1 to 10 nucleotides.

105. The method according to any one of claims 76 to 95, wherein the sticky end comprises a single-stranded polynucleotide overhang of 2 to 6 nucleotides.

106. The method according to any one of claims 76 to 95, wherein the sticky end comprises a single-stranded polynucleotide overhang of 3 to 5 nucleotides.

107. The method according to any one of claims 76 to 95, wherein the sticky end comprises a blunt end.

108. The method according to any one of claims 76 to 95, wherein the sticky end has a 5' single-stranded polynucleotide overhang.

109. The method according to any one of claims 76 to 95, wherein the sticky end has a 3' single-stranded polynucleotide overhang.

110. The method according to any one of claims 76 to 95, wherein the eukaryotic cell is an animal or human cell.

111. The method according to any one of claims 76 to 95, wherein the eukaryotic cell is a human cell.

112. The method according to any one of claims 76 to 95, wherein the eukaryotic cell is a plant cell.

113. The method according to any one of claims 76 to 95, wherein the modification is a deletion of at least a part of the target sequence.

114. The method according to any one of claims 76 to 95, wherein the modification is a mutation of the target sequence.

115. The method according to any one of claims 76 to 95, wherein the modification is an insertion of a target sequence into the target sequence.

116. A method for reducing the degradation of the Cas9 effector protein in a cell, comprising: a) binding a first nuclear localization signal to the N-terminus of the Cas9 effector protein; b) binding a second nuclear localization signal to the C-terminus of the Cas9 effector protein. The method comprising.

117. The method according to claim 116, wherein the first nuclear localization signal is a monopartite nuclear localization signal.

118. The method according to claim 116, wherein the first nuclear localization signal is a bipartite nuclear localization signal.

119. The method according to claim 116, wherein the second nuclear localization signal is a monopartite nuclear localization signal.

120. The method according to claim 116, wherein the second nuclear localization signal is a bipartite nuclear localization signal. **Claim 121** The method according to claim 116, wherein the first nuclear localization signal is a bipartite nuclear localization signal and the second nuclear localization signal is a monopartite nuclear localization signal. **Claim 122** The method according to any one of claims 117, 119, or 121, wherein the monopartite nuclear localization signal is the nuclear localization signal of SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, TUS protein, or a combination thereof. **Claim 123** The method according to any one of claims 118, 120, or 121, wherein the bipartite nuclear localization signal is a classical bipartite nuclear localization signal. **Claim 124** The method according to claim 121, wherein the first nuclear localization signal is a classical bipartite nuclear localization signal and the second nuclear localization signal is the nuclear localization signal of SV40 large T antigen.