Improved methods and compositions for crispr interference and activation

EP4684009A1Pending Publication Date: 2026-01-28ALGEN BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024775675
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-18
Filing Date
2024-03-20
Publication Date
2026-01-28

AI Technical Summary

Technical Problem

Current CRISPR/Cas systems face challenges in efficiently targeting specific nucleic acid sequences in mammalian cells due to limitations in nuclease activity and specificity, particularly in directing sequence-specific binding and modulating gene expression without inducing DNA cleavage.

Method used

Development of a nuclease-deficient class 2, type V Cas polypeptide fusion protein with reduced nuclease activity, combined with an effector domain and engineered guide nucleic acid, to facilitate targeted modulation of gene expression by forming complexes with specific nucleic acid sequences and regulating transcription without DNA cleavage.

Benefits of technology

The solution enables robust and specific regulation of gene expression, achieving significant reduction in mRNA levels for targeted genes and enhanced transcriptional activation, outperforming existing CRISPR systems in certain contexts by maintaining nuclease activity and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024020801_26092024_PF_FP
    Figure US2024020801_26092024_PF_FP
Patent Text Reader

Abstract

Described herein are methods and compositions utilizing CRISPR / Cas systems. In some cases, the utilization includes using such systems for gene activation or interference.
Need to check novelty before this filing date? Find Prior Art

Description

IMPROVED METHODS AND COMPOSITIONS FOR CRISPR INTERFERENCE AND ACTIVATIONCROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 453,559 filed March 21, 2023 and U.S. Provisional Application No. 63 / 583,507, filed September 18, 2023, which applications are incorporated by reference herein in their entireties.BACKGROUND

[0002] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as an adaptive immune system in microbes. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally comprise two parts: (i) an array of short repetitive sequences (30-40bp) separated by equally short spacer sequences, which encode the RNA-based targeting element; and (ii) ORFs encoding the Cas encoding the nuclease polypeptide directed by the RNA-based targeting element alongside accessory proteins / enzymes. Efficient nuclease targeting of a particular target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (the target seed) and the RNA guide; and (ii) the presence of a protospacer-adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM usually being a sequence not commonly represented within the host genome). In some cases, CRISPR / Cas systems can be modified to direct sequence-specific binding using complementary hybridization of the RNA guide.SUMMARY

[0003] In some aspects, the present disclosure provides for a fusion protein comprising: (a) a nuclease-deficient class 2, type V Cas polypeptide: (i) comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 35, 36, or a variant thereof; or (ii) comprising a TSL or C-terminal domain deletion, or any combination thereof relative to SEQ ID NO: 1; and (b) an effector domain comprising or consisting of a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 40-45, or a variant thereof. In some embodiments, the fusion protein further comprises a nuclear localization signal (NLS). In some embodiments, said NLS is between (a) and (b). In some embodiments, said NLS is near toan N- or C-terminus of said fusion protein. In some embodiments, (a) is near to said N-terminus of said fusion protein and (b) is near to said C-terminus of said fusion protein. In some embodiments, (a) is near to said C-terminus of said fusion protein and (b) is near to said N- terminus of said fusion protein. In some embodiments, (a) and (b) are in order from N- to C- terminus. In some cases, a Cas endonuclease described herein with reduced nuclease activity is catalytically dead. In some cases, a Cas endonuclease described herein with reduced nuclease activity has a reduction of nuclease activity of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.5% relative to a wild-type version of the Cas endonuclease. In some embodiments, (b) and (a) are in order from N- to C-terminus. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of residue 925 or 929 relative to SEQ ID NO: 1 to a non-wild-type residue. In some embodiments, said nonwild-type residue comprises alanine, glycine, lysine, or arginine. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises an A925K or an I929K mutation, or any combination thereof relative to SEQ ID NO: 1. In some embodiments, said nuclease- deficient class 2, type V Cas polypeptide comprises said A925K mutation and said I929K mutation relative to SEQ ID NO: 1. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of at least one residue corresponding to residues 659, 756, 922 of SEQ ID NO: 1, or any combination thereof, to a non-wild-type residue. In some embodiments, said non-wild-type residue is a glycine or an alanine residue. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises said sequence having at least 80% sequence identity to any one of SEQ ID NOs: 35, 36, or a variant thereof. In some embodiments, said nuclease- deficient class 2, type V Cas polypeptide comprises said TSL or C-terminal domain deletion, or any combination thereof. In some embodiments, said effector domain comprises a sequence having at least 80% sequence identity to SEQ ID NO: 45 or a variant thereof and comprises a mutation of at least one of the residues mutated in SEQ ID NO: 46 relative to SEQ ID NO: 45.

[0004] In some aspects, the present disclosure provides for a polynucleotide or nucleic acid encoding any of the fusion proteins described herein. In some embodiments, the nucleic acid is codon optimized for expression in a mammalian or human cell.

[0005] In some aspects, the present disclosure provides for a system comprising: (a) any of the fusion proteins described herein or a polynucleotide encoding any of the fusion proteins described herein; (b) an engineered guide nucleic acid or a template polynucleotide encodingsaid engineered guide nucleic acid, wherein said engineered guide nucleic acid configured to form a complex with said nuclease-deficient class 2, type V Cas polypeptide, comprising: (i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence; (ii) a scaffold nucleic acid sequence configured to bind to said nuclease-deficient class 2, type V Cas polypeptide. In some embodiments, said engineered guide nucleic acid comprises (ii) and (i) in order from 5’ to 3’. In some embodiments, said engineered guide nucleic acid comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 6-12 or 47-52, or a variant thereof. In some embodiments, said class 2, type V Cas polypeptide sequence is a class 2, type V-E polypeptide sequence. In some embodiments, said anti -target nucleic acid sequence comprises at least 16-22 nucleotides.

[0006] In some aspects, the present disclosure provides for a polypeptide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 1, 35, 36, or a variant thereof or comprising a TSL or C-terminal domain deletion relative to SEQ ID NO: 1, or any combination thereof. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of residue 925 or 929 relative to SEQ ID NO: 1 to a non-wild-type residue. In some embodiments, said non-wild-type residue comprises alanine, glycine, lysine, or arginine. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises an A925K or an I929K mutation, or any combination thereof relative to SEQ ID NO: 1. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises said A925K mutation and said I929K mutation relative to SEQ ID NO: 1.

[0007] In some aspects, the present disclosure provides for a polynucleotide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, atleast about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 47-52 a variant thereof.

[0008] In some aspects, the present disclosure provides for a polynucleotide comprising or encoding a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 47-52, or a variant thereof.

[0009] In some aspects, the present disclosure provides for a system, comprising: (a) a polypeptide comprising a class 2, type V Cas polypeptide sequence or a polynucleotide encoding said class 2, type V Cas polypeptide sequence; and (b) a polynucleotide any of the engineered guide nucleic acids described herein or a template polynucleotide encoding any of the engineered guide nucleic acids described herein. In some embodiments, said class 2, type V Cas polypeptide comprises any of the polypeptides described herein. In some embodiments, said class 2, type V Cas polypeptide sequence is a class 2, type V-E polypeptide sequence. In some embodiments, said anti -target nucleic acid sequence comprises at least 16-22 nucleotides.

[0010] In some aspects, the present disclosure provides for a method of binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, comprising contacting to said cell any of the systems described herein. In some embodiments, said DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to 5'-TCN-3' to said target nucleic acid sequence. In some embodiments, said contacting comprises transfecting said cell with said system. In some embodiments, said contacting comprises virally transducing said cell with said system. In some embodiments, said contacting comprises virally transducing said cell with one of (a) or (b) and transfecting said cell with the other of (a) or (b). In some embodiments, (a) further comprises an effector domain comprising a transcriptional repression domain, further comprising repressing transcription of said target nucleic acid sequence. In some embodiments, (a) further comprises an effector domain comprising a transcriptional activation domain, further comprising activating transcription of said target nucleic acid sequence.

[0011] In some aspects, the present disclosure provides for a cell comprising the any of the systems described herein, any of the polynucleotides described herein, or any of the polypeptides described herein. In some embodiments, said cell is a eukaryotic cell.

[0012] In some aspects, the present disclosure provides for a vector comprising or encoding any of the polynucleotides or nucleic acids described herein. In some embodiments, said vector is a plasmid or a viral vector. In some embodiments, said vector is a viral vector, wherein said viral vector is an adeno-associated virus (AAV) vector.

[0013] In some aspects, the present disclosure provides for a fusion protein comprising: (a) a nuclease-deficient class 2, type V Cas polypeptide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to SEQ ID NO: 1, or a variant thereof; (b) an effector domain comprising a transcriptional activation domain or a transcriptional repression domain. In some embodiments, said effector domain comprises a transcriptional repression domain. In some embodiments, said effector domain comprises a transcriptional repression domain, wherein said transcriptional repression domain comprises a Kriippel-associated box (KRAB) domain, a chromo domain, a chromo-shadow domain, a CBX domain, a Methyl-CpG Binding Protein 2 (MECP2) domain, a SIN3 Transcription Regulator Family Member A (SIN3 A) domain, a Histone deacetylase HDT1 (HDT1) domain, a Methyl-CpG Binding Domain Protein 2 (MBD2) domain, a Protein Phosphatase 1 Regulatory Subunit 8 (PPP1R8) domain, or a Chromobox 5 (CBX5) domain, a DNA Cytosine Methyltransferase (DCM) domain, a DNA Cytosine Methyltransferase-like (DCML) domain or any combination thereof. In some embodiments, said effector domain comprises a transcriptional activation domain. In some embodiments, said effector domain comprises a transcriptional activation domain, wherein said transcriptional activation domain comprises a Kriippel-associated box (KRAB) domain, a transcript! on / trans-activating domain (TAD), a RelA / p65 gene activating domain (p65), a FOXO3 domain, or a viral gene activation domain, any combination thereof. In some embodiments, said transcriptional activation domain comprises a viral gene activation domain, wherein said viral gene activation protein / domain comprises Tegument protein VP16 (VP16), or Replication and Transcription Activator (RTA) domain. In some embodiments, said fusion protein comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-5, 13-19, or a variant thereof. In some embodiments, the fusion protein further comprises a nuclear localization signal (NLS). In some embodiments, the NLS is between (a) and (b). In some embodiments, the NLS is near to the N- or C-terminus of the fusion protein. In some embodiments, (a) is near to the N-terminus of the fusion protein and (b) is near to the C-terminus of the fusion protein. In some embodiments, (a) is near to the C-terminus of the fusion protein and (b) is near to the N-terminus of the fusion protein. In some embodiments, (a) and (b) are in order from N- to C-terminus. In some embodiments, (b) and (a) are in order from N- to C-terminus. In some embodiments, said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of at least one residue corresponding to residues 659, 756, 922 of SEQ ID NO: 1, or any combination thereof, to a non-wild-type residue. In some embodiments, said non-wild-type residue is a glycine or an alanine residue. In some embodiments, aid nuclease- deficient class 2, type V Cas polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof.

[0014] In some aspects, the present disclosure provides for a polynucleotide encoding any of the fusion proteins described herein.

[0015] In some aspects, the present disclosure provides for a system comprising: (a) a fusion protein according to any of the aspects or embodiments described herein or a polynucleotide encoding a fusion protein according to any of the aspects or embodiments described herein; (b) an engineered guide nucleic acid or a template polynucleotide encoding said engineered guide nucleic acid, wherein said engineered guide nucleic acid configured to form a complex with said nuclease-deficient class 2, type V Cas polypeptide, comprising: (i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence; (ii) a scaffold nucleic acid sequence configured to bind to said nuclease-deficient class 2, type V Cas polypeptide. In some embodiments, said engineered guide nucleic acid comprises (ii) and (i) in order from 5’ to 3’. In some embodiments, said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 6-12 or a variant thereof. In some embodiments, said engineered guide nucleic acid comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to SEQ ID NO: 6, or a variant thereof, further comprising at least one mutation according to Table 2 relative to SEQ ID NO: 6. In some embodiments, said anti -target sequence comprises at least 16-22 nucleotides. In some embodiments, said fusion protein comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at leastabout 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to SEQ ID NO: 14, and further comprises a polypeptide having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to SEQ ID NO: 15.

[0016] In some aspects, the present disclosure provides for a polynucleotide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% to SEQ ID NO: 6, or a variant thereof, further comprising at least one mutation according to Table 2 relative to SEQ ID NO: 6. In some embodiments, said polynucleotide comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 7-12, or a variant thereof. In some embodiments, the polynucleotide further comprises an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence, which anti -target sequence is oriented 3’ to said sequence having at least 80% sequence identity to SEQ ID NO: 6 or said sequence at least 90% sequence identity to any one of SEQ ID NOs: 7-12, or a variant thereof. In some embodiments, said anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

[0017] In some aspects, the present disclosure provides for a polynucleotide encoding any of the polynucleotides described herein.

[0018] In some aspects, the present disclosure provides for a system, comprising: (a) a polypeptide comprising a class 2, type V Cas polypeptide sequence or a polynucleotide encoding said class 2, type V Cas polypeptide sequence; and (b) a polynucleotide according to any of the aspects or embodiments described herein, or a template polynucleotide encoding said polynucleotide according to any of the aspects or embodiments described herein. In some embodiments, said class 2, type V Cas polypeptide sequence is a class 2, type V-E polypeptide sequence. In some embodiments, said polypeptide further comprises an effector domain comprising a transcriptional activation domain or a transcriptional repression domain. In some embodiments, said polypeptide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-5, 13-19, or a variant thereof. In some embodiments, said fusionprotein comprises a sequence having at least 80% sequence identity to SEQ ID NO: 14, further comprising a polypeptide having at least 80% sequence identity to SEQ ID NO: 15. In some embodiments, said polypeptide is nuclease-deficient. In some embodiments, said polypeptide comprises a substitution of at least one residue corresponding to residues 659, 756, 922 of SEQ ID NO: 1, or any combination thereof, to a non-wild-type residue. In some embodiments, said non-wild-type residue is a glycine or an alanine residue. In some embodiments, said polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof.

[0019] In some aspects, the present disclosure provides for a polypeptide comprising a sequence having at least 80% identity to SEQ ID NO: 1 comprising a substitution of at least two residues corresponding to residues 659, 756, or 922 of SEQ ID NO: 1 to a non-wild-type residue. In some embodiments, said non-wild-type residue is a glycine or an alanine residue. In some embodiments, said polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof. In some embodiments, said polypeptide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2-5, 13-19, or a variant thereof. In some embodiments, said polypeptide further comprises an effector domain comprising a transcriptional activation domain or a transcriptional repression domain. In some embodiments, said polypeptide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-5, 12-19, or a variant thereof.

[0020] In some aspects, the present disclosure provides for a polynucleotide encoding any of the polypeptides described herein.

[0021] In some aspects, the present disclosure provides for a system comprising: (a) a polypeptide according to any of the aspects or embodiments described herein or a polynucleotide encoding a polypeptide according to any of the aspects or embodiments described herein; (b) an engineered guide nucleic acid or a template polynucleotide encoding said engineered guide nucleic acid, wherein said engineered guide nucleic acid is configured to form a complex with said polypeptide, comprising: (i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence; and (ii) a scaffold nucleic acid sequence configured to bind to said polypeptide. In some embodiments, said engineered guide nucleic acid comprises (ii) and (i) in order from 5’ to 3’. In some embodiments, said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 6-12 or a variant thereof. In some embodiments, said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to SEQ ID NO: 6, or a variantthereof, further comprising at least one mutation according to Table 2 relative to SEQ ID NO: 6. In some embodiments, said anti -target sequence comprises at least 16-22 nucleotides. In some embodiments, said polypeptide comprises a sequence having at least 80% sequence identity to SEQ ID NO: 14, further comprising a polypeptide having at least 80% sequence identity to SEQ ID NO: 15.

[0022] In some aspects, the present disclosure provides for a method of binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, comprising contacting to said cell any of the systems described herein. In some embodiments, said DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to 5’-TCN-3’ to said target nucleic acid sequence. In some embodiments, said contacting comprises transfecting said cell with said system. In some embodiments, said contacting comprises virally transducing said cell with said system. In some embodiments, said contacting comprises virally transducing said cell with one of (a) or (b) and transfecting said cell with the other of (a) or (b). In some embodiments, (a) further comprises an effector domain comprising a transcriptional repression domain, further comprising repressing transcription of said target nucleic acid sequence. In some embodiments, (a) further comprises an effector domain comprising a transcriptional activation domain, further comprising activating transcription of said target nucleic acid sequence.

[0023] In some aspects, the present disclosure provides for a cell comprising any of the systems described herein, any of the fusion proteins described herein, or any of the polynucleotides described herein. In some embodiments, said cell is a eukaryotic cell.

[0024] In some aspects, the present disclosure provides for a vector comprising any of the polynucleotides described herein. In some embodiments, said vector is a plasmid or a viral vector. In some embodiments, said vector is a viral vector, wherein said viral vector is an adeno-associated virus (AAV) vector.

[0025] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0026] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:

[0028] Figure 1 (FIG. 1) depicts a schematic of genetic units for targeted modulation of mammalian gene expression using dPlmCasl2e protein fusions. To modulate mammalian gene expression patterns a nuclease inactivated (dead) Casl2e enzyme is fused to a mammalian nuclear localization signal (NLS) and effector domain via linker regions of variable length. Effector domains include documented gene regulatory domains (e.g. from vertebrate and viral genomes) that repress or activate gene expression. Expression of this fusion protein is driven by a robust and constitutive RNA polymerase II promoter. Simultaneously, a Casl2e single-guide RNA molecule with a variable and programmable ‘targeting or spacer sequence’ is expressed from a constitutive RNA polymerase III promoter. This ribonucleoprotein complex is directed by the targeting / spacer sequence to sites in the human genome containing both the target sequence and an upstream (5’) nucleotide motif consisting of or comprising TCN.

[0029] Figure 2 (FIG. 2) illustrates nuclease inactivating mutations in PlmCasl2e. Shown is a structure based on Tsuchida et al. (Molecular Cell 82, no. 6 (March 2022): 1199-1209. e6, which is incorporated by reference in its entirety herein; PDB: 7WB1) showing Casl2e bound to gRNA and double-stranded DNA; this structure illustrates that three amino acids (D659, E756, and D922; shown as spheres) cluster near the non-target DNA strand to contribute to DNA cleavage through DNA binding and hydrolysis. These residues are mutated to alanine to create PlmCasl2e3Anuclease inactive variant.

[0030] Figure 3 (FIG. 3) depicts putative models of PlmCasl2e gRNA variants (SEQ ID NOs: 6-11) . The putative folding pattern for Casl2e gRNA variants is based on electron microscopy of homologous ribonucleoprotein complexes (Liu et al. Nature 566, no. 7743 (February 2019): 218-23). The proposed Casl2e gRNA variants include substitution mutations to improve transcription in mammalian cells, replacing the sequence of RNA loops such that they adoptmore energetically favorable geometries, and extending the stem of hairpins. Potential sites of mutation and possible substitute residues (e.g. mutations A-K in Table 2) are show in bold and denoted with *.

[0031] Figure 4 (FIG. 4) depicts hypothetical illustrations of Casl2e gRNA variants within the ribonucleoprotein complex, revealing motivation for mutations. Substitutions to improve transcription in mammalian cells (top right) are based on findings that homopolymeric thymidine repeats cause premature transcriptional termination (Bogenhagen and Brown. Cell 24, no. 1 (April 1981): 261-70, which is incorporated by reference in its entirety herein). Sites chosen are unlikely to compromise complex formation based on structural data. Replacing a three-nucleotide RNA loop with a four-nucleotide motif commonly found in mammalian RNA sequences (bottom left) is likely to result in more stable gRNA and / or complex formation. Cas9 gRNAs have shown increased activity after extending stem loops (bottom right) which may increase the binding affinity of the gRNA-Cas complex.

[0032] Figure 5 (FIG. 5) depicts an experimental schematic for testing dPlmCasl2e-based gene modulation via transient plasmid expression. Shown is a schematic of experimental methods used to test gene repression activity of PlmCasl2e3A-HA-2xNLS-BFP-KRAB in human cells. Lipid particles were used to deliver many copies of a DNA plasmid encoding a Casl2e gRNA targeting the non-native GFP gene, four different native genes, or targeting nowhere in the human genome. These plasmids were delivered to engineered cells containing non-native genomic copies of PlmCasl23A-HA-2xNLS-BFP-KRAB and GFP. All the RNA molecules were extracted from cells transfected with plasmids and relative transcript levels for the genes were determined via qPCR.

[0033] Figure 6 (FIG. 6) illustrates that targeting PlmCasl2e3A-KRAB to five different genomic loci in human cells via transient expression of gRNAs causes gene silencing. Human tumor cells expressing dSpCas9-HA-2xNLS-BFP-KRAB or PlmCasl2e3A-HA-2xNLS-BFP- KRAB were transiently transfected with plasmids encoding the matched gRNAs targeting the transcription start sites of an ectopic GFP reporter gene or endogenous human genes (RPA1, STK33, GUCY1A2, and CD164). Cells were harvested after 48 hours and quantities of target RNAs were determined via quantitative PCR and normalized to cells transfected with a plasmid encoding a non-targeting control gRNA that does not bind the human genome. Shown is a bar graph of GFP, RPA1, STK33, GUCY1A2, or CD164 RNA expression relative to a non-targeting control as affected by either the PlmCasl2e3A-KRAB construct or an analogous dCas9-based construct.

[0034] Figure 7 (FIG. 7) depicts an experimental schematic for testing dPlmCasl2e-based gene modulation via viral transduction into host genome. Shown is a schematic of experimental methods that were used to test gene repression activity of PlmCasl2e3A-NLS-KRABZNF10in human cells. Non-replicating viral particles were used to insert multiple genes into the genome of human cells: (a) a gene conferring resistance to the toxin puromycin; and (b) a Casl2e gRNA targeting the human RPA1 gene or targeting nowhere in the human genome. Viral particles form in and are secreted by HEK293T cells and then media containing non-replicating viral particles is harvested. Engineered cells containing non-native genomic copies of PlmCasl23A-NLS- KRABZNF1° and GFP were infected with viral particles and were exposed to the toxin puromycin to isolate a pure population of gRNA expressing cells. RNA was extracted from these pure populations and relative RPA1 transcript levels were determined via qPCR.

[0035] Figure 8 (FIG. 8) shows that targeting PlmCasl2e3A-KRAB to RPA1 locus in human cells via viral transduction of gRNAs causes gene silencing. Human cells expressing dSpCas9- HA-2xNLS-BFP-KRAB or PlmCasl2e3A-HA-2xNLS-BFP-KRAB were infected with lentiviral particles encoding the matched gRNAs targeting the transcription start site of the endogenous RPA1 gene. Cells were harvested after 48 hours and quantity of RPA1 RNA was determined via quantitative PCR and normalized to cells infected with lentiviral particles encoding a nontargeting control gRNA that does not bind the human genome. Shown is a bar graph of RPA1 RNA expression relative to a non-targeting control.

[0036] Figure 9 (FIG. 9) depicts an experimental schematic for testing dPlmCasl2e-based gene modulation via transient plasmid expression. Shown is a schematic of experimental methods used to test gene activation activity of PlmCasl2e3A-2xNLS-VP64-p65-Rta (dPlmCasl2e-VPR) in human cells. Lipid particles were used to deliver a DNA plasmid encoding a Casl2e gRNA targeting six different native genes or targeting nowhere in the human genome. These plasmids were delivered to engineered cells containing non-native genomic copies of PlmCasl23A- 2xNLS-VP64-p65-Rta. All the RNA molecules were extracted from cells transfected with plasmids and relative transcript levels for the genes were determined via qPCR.

[0037] Figure 10 (FIG. 10) illustrates that targeting dPlmCasl2e-VPR to six different genomic loci in human cells via transient expression of gRNAs causes gene activation. Human tumor cells expressing dSpCas9-2xNLS-VP64-p65-Rta or dPlmCasl2e-2xNLS-VP64-p65-Rta were transiently transfected with plasmids encoding the matched gRNAs targeting the transcription start sites of endogenous human genes (GCK, CD164, GUCY1A2, KDSR, PPAT, and STK17A). Cells were harvested after 48 hours and quantities of target RNAs were determined via quantitative PCR and normalized to cells transfected with a plasmid encoding a non-targeting control gRNA that does not bind the human genome. Shown is a bar graph of GCK, CD164, GUCY1A2, KDSR, PPAT, or STK17A RNA expression relative to a non-targeting control as affected by either the dPlmCasl2e-VPR construct or an analogous dCas9-based construct.

[0038] Figure 11 (FIG. 11) illustrates a sequence homology-based rationale for identifying PlmCasl2e amino acids required for nuclease activity. (Top) Domain structure of PlmCasl2e within two-dimensional sequence space. This includes a non-target DNA strand loading (NTSL) domain, an Oligobinding domain (OBD), a RuVC nuclease domain interrupted in sequence order by a target DNA strand loading domain. (Bottom) Sequence alignment of eight diverse type V Cas nuclease (RuVC) domains. Positions with well-conserved amino acids are highlighted and catalytic residues D659, E765, and D922 that were mutated in Casl2e3Aare specifically called out. Additionally, two well conserved residues A925 and 1929 expected also to contribute to nuclease activity are called out.

[0039] Figure 12 (FIG. 12) illustrates a ribonucleoprotein structure-based rationale for identifying PlmCasl2e amino acids required for nuclease activity. (Top) Evolutionarily constrained residues A925 and 1929 were specifically identified because models based on cryoelectron microscopy data suggest that substituting these positions with lysine may provide steric and electrostatic interference with DNA cleavage. (Center) Analysis of RNP structure also indicates that the C-terminal target DNA strand loading (TSL) domain and part of the nuclease domain are distal from the PlmCasl2e:gRNA interface and the gRNA:DNA hybrid suggesting their removal may inactivate DNA cleavage without compromising DNA binding. (Bottom) Sequence diagram of potential nuclease dead PlmCasl2e variants.

[0040] Figure 13 (FIG. 13) shows that cells containing ectopic, genomic copies of green fluorescent protein (GFP), neomycin resistance (NeoR), and a PlmCasl2e variant were transduced with lentiviral particles encoding PlmCasl2e gRNAs targeting the coding region of GFP or NeoR. Successful viral transduction is ensured by a puromycin resistance gene that was delivered with the gRNA and allows cells to survive puromycin treatment. Nuclease activity is then determined either through direct DNA sequencing of the target genomic DNA, or through functional loss of the NeoR gene due to editing.

[0041] Figure 14 (FIG. 14) depicts experimental results for nuclease activity of PlmCasl2e variants. (Left) NeoR knockout assay showed that wild-type PlmCasl2e caused 40% loss in drug resistance relative to the nuclease silenced PlmCasl2e3A. The panel of PlmCasl2e variants exhibited the entire range of nuclease activity. A dot represents the average of four experiments with a single gRNA, and the bar is the average of the two NeoR gRNAs. (Right) PCRamplification of gRNA target sites in the genome followed by deconvolution capillary sequencing confirmed results of the NeoR knockout assay. The values represent the average aberrant reads observed in two experiments.

[0042] Figure 15 (FIG. 15) depicts an experimental schematic to measure gene modulating activity of PlmCasl2e variants and effector fusions. Human tumor-derived cells are transduced with dPlmCasl2e variants fused to different effector domains and positively selected for a cotransduced blasticidin resistance gene. A sub-population of cells is transiently transfected with a plasmid encoding a non-targeting PlmCasl2e gRNA, while a separate sub -population is transfected with a plasmid encoding a gRNA targeting an endogenous human gene. RNA is extracted, converted to cDNA, and then relative amounts of mRNA molecules are measured via qPCR, assaying both gene inhibition and activation.

[0043] Figure 16 (FIG. 16) illustrates that PlmCasl2e nuclease inactive variants provide more robust transcriptional repression than dSpCas9 when fused to KRABZNF1°. Human tumor- derived cells are transduced with dPlmCasl2e variants or dSpCas9 fused to KRABZNF1° were transduced with 12 gRNAs targeting six endogenous human genes and mRNA content was assayed 72 hours later via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range. Numbers are paired t-tests comparing dSpCas9 values to the dPlmCasl2e variants. All values were determined from three experimental replicates.

[0044] Figure 17 (FIG. 17) illustrates that PlmCasl2e nuclease inactive variants provide more robust transcriptional activation than dSpCas9 when fused to VPR. Human tumor-derived cells are transduced with dPlmCasl2e variants or dSpCas9 fused to VPR were transduced with 12 gRNAs targeting six endogenous human genes and mRNA content was assayed 72 hours later via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range and any points outside that range are individually plotted.Numbers are paired t-tests comparing dSpCas9 values to the dPlmCasl2e variants. All values were determined from two experimental replicates.

[0045] Figure 18 (FIG. 18) illustrates that putative secondary structure and rationale for mutations made to PlmCasl2e gRNA scaffold. The single-guide RNA derived from bacterial sequences (vl.O) was modified in several ways to improve transcription and stability of the molecule. A variant (v2.1) was generated to contain modifications that increase base-pairing at key locations, remove polypyrimidine stretches that cause RNA polymerase III termination, and provide increased phosphodiester backbone for contact with the Casl2e protein. Further iterations were made (v2.25) to include features with more thermodynamically favorable geometries, ensure more consistent transcription initiation, and remove more polypyrimidine stretches. Finally, a set of third generation variants were designed (v3.0) to enable high- throughput next generation sequencing of the Casl2e gRNA.

[0046] Figure 19 (FIG. 19) illustrates that modifications of Casl2e gRNA (v2.1) provide more robust transcriptional silencing than bacterially derived sequences (vl.O) when fused to dCasl2e-KRABZNF10. Human tumor-derived cells are transduced with PlmCasl2e3Afused to KRABZNF1° were transfected with 12 vl.O or v2.1 gRNAs encoding the same spacer sequences that target six endogenous human genes. The mRNA content of a population was assayed 72 hours later via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range. The displayed number is a paired t-test comparing the two gRNA variants. Most vl.O values are duplicated from figure 7 and all values were determined from three experimental replicates.

[0047] Figure 20 (FIG. 20) illustrates that modifications of Casl2e gRNA (v2.1) provide more robust transcriptional activation than bacterially derived sequences (vl.O) when fused to dCasl2e-NLS-VPR. Human tumor-derived cells are transduced with PlmCasl2e3Afused to VPR were transfected with 12 vl .0 or v2.1 gRNAs encoding the same spacer sequences that target six endogenous human genes. The mRNA content of a population was assayed 72 hours later via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box andwhiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range and any points outside that range are individually plotted. The displayed number is a paired t-test comparing the two gRNA variants. Some vl.O values are duplicated from figure 8 and all values were determined from two experimental replicates.

[0048] Figure 21 (FIG. 21) illustrates that different truncations of the PPP1R8 Transcription Repression Domain (TReDppplR8) silence gene expression similarly or better than KRABZNF10 when fused to dPlmCasl2e. Human tumor-derived cells are transduced with PlmCasl2e3Afused to KRABZNF1°, or two fragments of PPP1R8, TReDppplR8or mini-TReDppplR8were transfected with 12 gRNAs targeting six endogenous human genes. The mRNA content of the populations were assayed 72 hours later via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range and any points outside that range are individually plotted. The displayed number is a paired t-test comparing the TReDppplR8variants to KRABZNF1°. Most KRABZNF1° values are duplicated from figure 7 and all values were determined from three experimental replicates.

[0049] Figure 22 (FIG. 22) illustrates that the FOXO3 Transcription Activation Domain (TADFOXO3) can activate gene expression similarly or better than VPR when fused to dPlmCasl2e. Human tumor-derived cells are transduced with PlmCasl2e3Afused to VPR or TADFOXO3were transfected with 12 gRNAs targeting six endogenous human genes. The mRNA content of the populations were assayed 72 hours later via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range. The displayed number is a paired t-test comparing the TADFOXO3VPR. Some VPR values are duplicated from figure 8 and all values were determined from two experimental replicates.

[0050] Figure 23 (FIG. 23) illustrates that more robust transcriptional silencing is achieved, on average, using newly described dPlmCasl2e, gRNA scaffold, and repressor domain variants. Three gene silencing systems were compared in human tumor-derived cells. The Cas9 system comprised a dSpCas9-KRABZNF10fusion and 12 Cas9 sgRNAs targeted to six endogenoushuman genes. The Casl2e system included a PlmCasl2e3A-KRABZNF10fusion and 12 Casl2e vl.O sgRNAs targeting the same six human genes near Cas9 target sites but shifted due to unique PAM limitations. The Algen Modified system (Algen Mods) includes a PlmCasl2eAC- TReDppplR8fusion and 12 Casl2e v2.1 sgRNAs containing the exact spacer sequences as vl.O. The mRNA content of the populations were assayed 72 hours after gRNA transfection via quantitative real-time PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range. The brackets indicate which populations were compared for a paired t-test. Most Cas9 and Casl2e values are duplicated from Figure 7 and all values were determined from three experimental replicates.

[0051] Figure 24 (FIG. 24) illustrates that more robust transcriptional activation is achieved, on average, using newly described dPlmCasl2e, gRNA scaffold, and activation domain variants. Three gene silencing systems were compared in human tumor-derived cells. The Cas9 system comprised a dSpCas9-VPR fusion and 12 Cas9 sgRNAs targeted to six endogenous human genes. The Casl2e system included a PlmCasl2e3A-VPR fusion and 12 Casl2e vl.O sgRNAs targeting the same six human genes near Cas9 target sites but shifted due to unique PAM limitations. The Algen Modified system (Algen Mods) includes a PlmCasl2eAC-TADFOXO3fusion and 12 Casl2e v2.1 sgRNAs containing the exact spacer sequences as vl.O. The mRNA content of the populations were assayed 72 hours after gRNA transfection via quantitative realtime PCR. (Left) a bar represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with a gRNA targeting nowhere in the human genome. (Right) Individual gene values were collated and represented as a box and whiskers plots with the box denoting the median and interquartile range, while the whiskers mark 1.5 of the interquartile range and any points outside that range are individually plotted. The brackets indicate which populations were compared for a paired t-test. Most Cas9 and Casl2e values are duplicated from figure 8 and all values were determined from two experimental replicates.DETAILED DESCRIPTIONOverview

[0052] In some aspects herein, a CRISPR / Cas ribonucleoprotein complex identified in the bacterial phylum Planctomycetes (PlmCasl2e) was modified to alter mammalian geneexpression behavior in a targeted manner (FIG. 1). By fusing two RNA components of the Plm CRISPR / Cas system and mutating the ‘spacer’ or ‘targeting’ sequence of this ribonucleoprotein complex, it was modified to target arbitrary sites in the human genome. In addition, amino acids within the PlmCasl2e sequence were altered to enable DNA binding but prevent DNA cleavage, referred to as ‘dead’ PlmCasl2e or “dPlmCasl2e” (FIGs. 2, 11, and 12) to favor transcriptionbased effects of the complex instead of DNA editing, homologous recombination, or DNA repair. We combine this precise genomic targeting behavior with a diverse set of vertebrate (e.g. mammalian) and viral proteins / domains that are documented as regulating nuclear localization and transcriptional activities in human cells (FIG. 1). Altogether this allows precise PlmCasl2e- based modulation of mammalian gene expression.

[0053] This targeted nature of PlmCasl2e-based gene modulation relies on a synthetic RNA molecule referred to as a guide or single-guide RNA (gRNA or sgRNA). An initial molecule based on the naturally occurring bacterial sequence was used to experimentally verify this technology; however, key changes to this molecule were, or remain to be tested to respond to the unique demands of mammalian gene regulation and laboratory experimentation (FIGs. 3, 4, and 18). These modifications are proposed to optimize the Casl2e gRNA for biosynthesis, stability, improved complex formation, and high throughput nucleic acid experimentation in mammalian cells and are informed in part from cryo-electron microscopy studies.

[0054] The fundamental PlmCasl2e-based mammalian gene modulating tool was tested in human tumor derived cells by repressing the expression of a non-native GFP gene, and endogenous genes (RPA1, STK33, GUCY1A2, CD164) (FIGs. 5, 6, 7, and 8). The dPlmCasl2e protein was fused to a gene silencing ‘KRAB’ domain from the human ZNFIO / Koxl protein and inserted into the genome of human tumor cells. Casl2e gRNA molecules were then delivered to these cells using two methods, lipid-based transfection of plasmids (FIGs. 5 and 6), or viral transduction into the host cell genome (FIGs. 7 and 8). These gRNAs targeted a non-native GFP gene or multiple endogenous human genes (RPA1, STK33, GUCY1A2, CD164) and were compared to a control gRNA that does not bind to the human genome in any location. When the quantities of mRNA transcripts from targeted genes were compared to the non-targeting control we observed robust gene repression (e.g. 25-99% reduction in mRNA expression). This behavior was also compared relative to the dSpCas9 system, and it was found that dPlmCasl2e repressed genes similarly or more robustly than dSpCas9 in some instances (FIGs. 6 and 8).

[0055] Definitions

[0056] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0057] The practice of some methods disclosed herein employ, unless otherwise indicated, techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA. See for example Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4thEdition (2012); the series Current Protocols in Molecular Biology (F. M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6thEdition (R.I. Freshney, ed. (2010)) (which is entirely incorporated by reference herein).

[0058] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising”.

[0059] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within one or more than one standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0060] The term “or”, as used herein, is intended to be an inclusive “or”.

[0061] As used herein, a “cell” generally refers to a biological cell. A cell may be the basic structural, functional or biological unit of a living organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include: a prokaryotic cell, eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a single-cell eukaryotic organism, a protozoa cell, a cell from a plant, an algal cell, an animal cell, a cell from an invertebrate animal, a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), or a cell from amammal (e.g., a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, a human, etc.).

[0062] The term “nucleotide,” as used herein, generally refers to a base-sugar-phosphate combination. A nucleotide may comprise a synthetic nucleotide. A nucleotide may comprise a synthetic nucleotide analog. Nucleotides may be monomeric units of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates such as dATP, dCTP, diTP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [aS]dATP, 7-deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance on the nucleic acid molecule containing them. Nucleotides can also be labeled or marked by chemical modification. A chemically- modified single nucleotide can be biotin-dNTP. Some non-limiting examples of biotinylated dNTPs can include, biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin- 11-dCTP, biotin- 14-dCTP), and biotin-dUTP (e.g., biotin- 11-dUTP, biotin- 16-dUTP, biotin-20-dUTP). In some cases, nucleotides can encompass modifications at the base moiety or modifications at the phosphate moiety (e.g. modifications intended to influence stability of nucleotide polymers).

[0063] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to generally refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multistranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may exist in a cell-free environment. A polynucleotide may be a gene or fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may comprise one or more analogs (e.g., altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some nonlimiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queuosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA),ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro- RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides may be interrupted by non-nucleotide components.

[0064] The terms “transfection” or “transfected” generally refer to introduction of a nucleic acid into a cell by non-viral based methods. The nucleic acid molecules may be gene sequences encoding complete proteins or functional portions thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0065] The terms “transduction” or “transduced” generally refer to introduction of a nucleic acid into a cell by viral-based methods. The nucleic acid molecules may be gene sequences encoding complete proteins or functional portions thereof.

[0066] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to generally refer to a polymer of at least two amino acid residues joined by peptide bond(s). This term does not connote a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The terms include amino acid chains of any length, including full length proteins, and proteins with or without secondary or tertiary structure (e.g., domains). The terms also encompass an amino acid polymer that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation such as conjugation with a labeling component. The terms “amino acid” and “amino acids,” as used herein, generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogues. Modified amino acids may include natural amino acids and non-natural amino acids, which have been chemically modified to include a group or a chemical moiety not naturally present on the amino acid. Amino acid analogues may refer to amino acid derivatives. The term “amino acid” includes both D-amino acids and L-amino acids.

[0067] The term “expression”, as used herein, generally refers to the process by which a nucleic acid sequence or a polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides may becollectively referred to as “gene product.” If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0068] A “vector” as used herein, generally refers to a macromolecule or association of macromolecules that comprises or associates with a polynucleotide and which may be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. The vector generally comprises genetic elements, e.g., regulatory elements, operatively linked to a gene to facilitate expression of the gene in a target.

[0069] As used herein, a “guide nucleic acid” can generally refer to a nucleic acid that may hybridize to another nucleic acid. A guide nucleic acid may be RNA. A guide nucleic acid may be DNA. The guide nucleic acid may be programmed to bind to a sequence of nucleic acid site- specifically. The nucleic acid to be targeted, or the target nucleic acid, may comprise nucleotides. The guide nucleic acid may comprise nucleotides. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. The strand of a doublestranded target polynucleotide that is complementary to and hybridizes with the guide nucleic acid may be called the complementary strand. The strand of the double-stranded target polynucleotide that is complementary to the complementary strand, and therefore may not be complementary to the guide nucleic acid may be called noncomplementary strand. A guide nucleic acid may comprise a polynucleotide chain and can be called a “single guide nucleic acid.” A guide nucleic acid may comprise two polynucleotide chains and may be called a “double guide nucleic acid.” If not otherwise specified, the term “guide nucleic acid” may be inclusive, referring to both single guide nucleic acids and double guide nucleic acids. A guide nucleic acid may comprise a segment that can be referred to as an “anti-target nucleic acid sequence” or a “nucleic acid-targeting sequence.” A guide nucleic acid may comprise a subsegment that may be referred to as a “protein binding segment” or “protein binding sequence” or “scaffold nucleic acid sequence”.

[0070] The term “scaffold nucleic acid sequence”, when used in the context of an engineered guide nucleic acid, can generally refer to a nucleic acid with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity or sequence similarity to a wild type example sequence that is configured to form a complex with a Cas polypeptide (e.g., a direct-repeat sequence). “Scaffold nucleic acid sequence” may refer to a modified form of a scaffold nucleic acid sequence (e.g. a direct-repeat sequence) that can comprise a nucleotide change such as a deletion, insertion, or substitution, variant, mutation, or chimera.

[0071] The term “sequence identity” or “percent identity” in the context of two or more nucleic acids or polypeptide sequences, generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same, when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, e.g., BLASTP using parameters of a wordlength (W) of 3, an expectation I of 10, and the BLOSUM62 scoring matrix setting gap costs at existence of 11, extension of 1, and using a conditional compositional score matrix adjustment for polypeptide sequences longer than 30 residues.

[0072] Included in the current disclosure are proteins, nucleic acids, or variants thereof. In some cases, such variants have at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to the parent protein or nucleic acid, or the SEQ ID NO specifying said parent protein or nucleic acid.

[0073] Included in the current disclosure are variants of any of the polypeptides described herein with one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be accomplished by substituting amino acids with similar hydrophobicity, polarity, and R chain length for one another. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by locating amino acid residues that have been mutated between species (e.g. non-conserved residues) without altering the basic functions of the encoded proteins. Such conservatively substituted variants may include variants with at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of the polypeptide protein sequences described herein.

[0074] In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can encompass sequences with substitutions such that the activity of critical active site residues of the endonuclease are not disrupted. In some embodiments, a functional variant of any of the proteins described herein lacks substitution of at least one conserved or functional residue.

[0075] Conservative substitution tables providing functionally similar amino acids are available from a variety of references (see, for e.g., Creighton, Proteins: Structures and Molecular Properties (W H Freeman & Co.; 2nd edition (December 1993), which is incorporated by reference in its entirety herein). The following eight groups contain amino acids that are conservative substitutions for one another:

[0076] 1) Alanine (A), Glycine (G);

[0077] 2) Aspartic acid (D), Glutamic acid (E);

[0078] 3) Asparagine (N), Glutamine (Q);

[0079] 4) Arginine (R), Lysine (K);

[0080] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);

[0081] 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);

[0082] 7) Serine (S), Threonine (T); and

[0083] 8) Cysteine (C), Methionine (M)

[0084] In some cases, a Cas endonuclease described herein In may comprise a variant having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N- or C- terminus of said endonuclease. The NLS may be appended N-terminal or C-terminal to any one the Cas endonuclease sequences described herein, or to a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of the Cas endonuclease sequences. The NLS may be an SV40 large T antigen NLS. The NLS may be a c-myc NLS. The NLS may be a nucleophosmin NLS. The NLS may be a bipartite fusion of two different NLS sequences. The NLS can comprise a sequence with at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs:65-80. The NLS can comprise a sequence substantially identical to any one of SEQ ID NOs: 65-80.

[0085] In some aspects, the present disclosure includes any of the polypeptides or nucleic acids described in Table 1. In some cases, the present disclosure includes nucleic acids encoding any of the polypeptides or nucleic acids described in Table 1. In some cases, the present disclosure includes any of the polypeptide domains identified as separate entities in Table 1, or any combination thereof.Table 1 : Sequences of Genes and Components Described HereinTable 2: Description of PlmCasl2e gRNA individual or combinatorial mutations described hereinTable 3: Sequences of guide RNA targeting (e.g. anti-target) sequences described hereinTable 4: Example NLS Sequences that can be used with Cas Effectors According to the DisclosureEXAMPLES

[0086] Example 1. - Design and Generation of Nucleic Acid Reagents

[0087] Based on experiments with Casl2e isolated from Deltaproteobacteria (Nature 566, no. 7743 (February 2019): 218-23, which is incorporated by reference herein for all purposes^ three amino acids (D659, E756, and D922) were mutated within Casl2e isolated from Planctomycetes (PlmCasl2e, SEQ ID NO: 1) to alanine (SEQ ID NO: 2). These mutations render the nuclease inactive or ‘dead’ and thus is referred to as “PlmCasl2e3A”, “Casl2e3A”, “dead PlmCasl2e”, “dPlmCasl2e” or “dCasl2e”. The corresponding DNA sequence was codon optimized for expression in human cells and inserted into a third-generation lentiviral packaging plasmid such that when translated it produces a protein fusion containing: dPlmCasl2e, the hemagglutinin epitope tag (HA), two nuclear localization signals (NLS), blue fluorescent protein (BFP), and the KRAB domain from the human protein ZNFIO / Koxl. This fusion protein was referred to as “dPlmCasl2e-HA-2xNLS-BFP-KRAB” or “dPlmCasl2e-KRAB”. While the dPlmCasl2e-HA- 2xNLS-BFP-KRAB fusion was the primary unit of experimental testing, the fusion of dPlmCasl2e, two nuclear localization signals, and an effector domain such as KRAB was sufficient for gene modulation unit. To this lentiviral packaging plasmid DNA encoding a separately expressed blasticidin S-resistance gene was then inserted to allow for enrichment of successfully transduced cells. As a positive control, an equivalent lentiviral packaging plasmid containing a nuclease deficient or “dead” mutant of Cas9 from Streptococcus pyogenes (dSpCas9) was generated. This protein product is referred to as “dSpCas9-HA-2xNLS-BFP- KRAB” or “dSpCas9-KRAB”.

[0088] Lentiviral packaging plasmids were also generated to express PlmCasl2e and SpCas9 guide RNAs (gRNAs), sometimes referred to as single-guide RNAs (sgRNAs). The specific DNA binding activities of PlmCasl2e and SpCas9 naturally rely on a pair of annealed RNA molecules, the tracrRNA and crRNA. For targeting the human genome, the sequences of the SpCas9 or PlmCasl2e tracrRNA and crRNA were fused with a synthetic RNA sequence to generate an sgRNA that was inserted downstream of the human U6 promoter. The PlmCasl2e sgRNA comprised a 101-nucleotide constant scaffold followed by a variable 20-nucleotide targeting or spacer sequence at the 3’ end of the molecule. As with other Cas enzymes, the DNA targeting site of dPlmCasl2e-KRAB is specified by the sequence of a protospacer adjacent motif (PAM) and the targeting sequence of the gRNA in close proximity. To design PlmCasl2e gRNA targets we identified a 23 -nucleotide sequence starting with the PAM sequence TCN at the 5’ end and followed by 20 nucleotides that compose a unique sequence within the human genome(e.g., TCNNNNNNNNNNNNNNNNNNNNN). The final 20 nucleotides excluding the PAM (5’ TCN) comprised the targeting or spacer sequence that was inserted at the 3’ end of the Casl2e sgRNAs. For SpCas9 and PlmCasl2e sgRNA sequences a universal lentiviral packaging plasmid that produces a gRNA targeting nowhere in the human genome was generated. These vectors served as a non-targeting control but also a scaffold where different targeting sequences can be easily swapped. In addition to gRNAs, these vectors also produce puromycin N- acetyltransferase (puroR) which confers resistance to the toxin puromycin. Using this approach, we generated SpCas9 gRNA plasmids to target four human genes (RPA1, STK33, GUCY1A2, and CD164) and a GFP gene. We also generated PlmCasl2e gRNA plasmids to target these same genes.

[0089] Example 2.- Generation of Cellular Reagents

[0090] Replication deficient lentiviral particles were generated in Hek293T cells by cotransfecting either Cas-Effector fusion donor plasmids with plasmids encoding the env gene from vesicular stomatitis virus G and the gag / pol genes from human immunodeficiency virus 1. These viral particles were used to insert the DNA encoding these fusion proteins into the genome of cells derived from a human tumor. In some cases, these cells also contained a non- endogenous copy of a green fluorescent protein (GFP) and a neomycin resistance gene (NeoR). Successfully transduced cells were isolated by treating the cell populations with blasticidin S. Replication deficient lentiviral particles were also generated for gRNA vectors targeting GFP and NeoR using the same method. Human tumor-derived cells expressing the PlmCasl2e variants described above were exposed to the gRNA containing viral particles and then successfully modified cells were selected through puromycin treatment..

[0091] Example 3 - Testing dPlmCasl2e-mediated gene modulation using transient plasmid expression

[0092] Lipofectamine 3000 was used according to manufacturer’s instructions to transiently transfect lentiviral packaging plasmids encoding SpCas9 gRNAs targeting GFP, RPA1, STK33, GUCY1A2, CD164, or non-targeting controls into human tumor cells expressing dSpCas9- KRAB and a non-endogenous GFP gene. Simultaneously, Lipofectamine 3000 was used to transiently transfect lentiviral packaging plasmids encoding PlmCasl2e gRNAs targeting GFP, RPA1, STK33, GUCY1A2, CD164, or non-targeting controls into human tumor cells expressing dPlmCasl2e-KRAB and a non-endogenous GFP gene. Cells were incubated for 48 hours, then all cellular RNA was harvested, and poly-adenylated mRNA was converted into DNA. Quantitative PCR was then used to measure the relative abundance of mRNA molecules transcribed from the GFP gene in cells transfected with plasmids encoding a non-targetingcontrol and cells transfected with plasmids encoding a GFP targeting gRNA (FIG. 6). This same experiment was performed by quantifying the relative abundance of mRNA molecules transcribed from the other four targeted genes (RPA1, STK33, GUCY1A2, CD164) in cells transfected with plasmids encoding a non-targeting control gRNA and cells transfected with plasmids encoding a gene-specific targeting gRNA (FIG. 6).

[0093] Example 4.- Testing dPlmCasl2e-mediated gene modulation using viral transduction into host genome

[0094] Replication-deficient lentiviral particles encoding gRNAs were generated in HEK293T cells by co-transfecting a gRNA plasmid with plasmids encoding the env gene from vesicular stomatitis virus G and the gag / pol genes from human immunodeficiency virus 1. Viral particles were generated for SpCas9 non-targeting and RPA1 targeting gRNAs. Similarly, viral particles were generated for PlmCasl2e non-targeting and RPA1 targeting gRNAs.

[0095] These replication deficient viral particles were used to insert DNA encoding one of these gRNAs and puroR into the genome of tumor cells expressing a non-endog enous GFP gene and dSpCas9-KRAB or dPlmCasl2e-KRAB. The population was enriched for cells containing gRNAs in their genome by exposing cells to puromycin 48 hours after exposure to viral particles. After five more days all cellular RNA was harvested, and poly-adenylated mRNA was converted into DNA. Quantitative PCR was then used to measure the relative abundance of RPA1 mRNA molecules in cells containing a genomic copy of a non-targeting control gRNA and cells containing a genomic copy of an RPA1 targeting gRNA (FIG. 8).

[0096] Example 5.- Testing dPlmCasl2e-mediated gene activation using transient plasmid expression

[0097] A PlmCasl2e-based mammalian gene modulating tool was tested in human tumor derived cells by activating the expression of six endogenous genes: CD164, GCK, GUCY1A2, STK17A, KDSR, and PPAT. The dPlmCasl2e protein was fused to three gene activation domains, VP64 (a quadruple repeat of a 13 amino acid sequence derived from the herpes simplex virus protein VP 16 that is documented as recruiting transcription initiation factors of RNA polymerase II), p65 / RelA (a 201 -amino acid sequence derived from the human gene RELA that heterodimerizes with transcription factor NF-kappa-B to activate gene expression), and Rta (a 190-amino acid sequence isolated from the Rta protein encoded by the Epstein-Barr virus R also documented to recruit RNA polymerase Il-specific transcription factors)(SEQ ID NO: 20). Fusion PlmCasl2e gRNA molecules were then delivered to these cells using lipid- based transfection of plasmids (Figure 10). These gRNAs targeted endogenous human genes (CD164, GCK, GUCY1 A2, STK17A, KDSR, and PPAT, SEQ ID NOs: 21-26) and werecompared to a control gRNA that does not bind to the human genome in any location. When mRNA transcripts from the gene targets were compared to the non-targeting control robust gene activation (e.g. 2.6 to 1800-fold increase in mRNA molecules) was observed. When this degree of activation was similarly compared to an analogous dSpCas9 system targeting the same genes (gRNAs SEQ ID NOs: 27-32) also fused to VP64, p65 / RelA, and Rta, it was found that the dPlmCasl2e module was a more potent activator. At five of six genes dPlmCasl2e stimulated 6- 270 times more transcription than dSpCas9 (Figure 10). Altogether we find dPlmCasl2e- mediated gene activation is successful in all cases tested and provides superior transcriptional amplification relative to dSpCas9 for many targets.

[0098] Example 6.- Testing PlmCaslle-nuclease activity (general procedure)

[0099] Cells genetically modified to contain GFP, NeoR, and a PlmCasl2e variant, were infected lentiviral particles containing one of two PlmCasl2e gRNAs targeting the coding region of NeoR. Successful integration of these gRNAs was ensured through puromycin treatment. After puromycin selection cells were incubated for 3 days to allow NeoR editing to take place. Cells were split into two populations, one that was exposed to a neomycin analog, G418, and one untreated. After 3 days the effect of G418 on cell proliferation was determined by counting cell numbers. Nuclease active variants were configured to target the NeoR gene for editing and sometimes deletion, making them sensitive to the drug, while nuclease dead variants were configured to retain resistance and be unaffected by the drug.

[0100] The untreated populations were also collected, and their genomic DNA extracted. Using PCR amplification and deconvolution of capillary sequencing reads we were able to ascertain the fraction of the genomic DNA that contained aberrant sequence due to nuclease editing. This same process was repeated for cells expressing a PlmCasl2e gRNA targeting GFP.

[0101] Example 7.- Testing dPlmCaslle-mediated gene modulation using transient plasmid expression (general procedure)

[0102] Lipofectamine 3000 was used to transiently transfect lentiviral packaging plasmids encoding SpCas9 or PlmCasl2e gRNAs targeting 10 endogenous human genes or nontargeting controls into tumor cells expressing a Cas-Effector fusion protein. Cells were incubated for 72 hours, then all cellular RNA was harvested as a crude lysate, and polyadenylated mRNA was specifically converted into complementary DNA (cDNA). The amount of cDNA was quantified using a fluorescence-based qPCR assay in an Agilent AriaMX Real- Time PCR system. This was performed in cells transfected with plasmids encoding a nontargeting control and cells transfected with plasmids encoding a targeting gRNA to determinethe relative abundance of target mRNA molecules. This was repeated for successive Cas- effector fusions and gRNA variants.

[0103] Example 8.- Design of additional nuclease-deficient (e.g. nuclease dead, nuclease inactive) mutants of PlmCaslle

[0104] Three negative charged amino acids (D659, E756, and D922) within PlmCasl2e were identified above as essential residues required to coordinate Mg2+ions and facilitate DNA hydrolysis. Consistent with their documented function, aligning the sequences of diverse type V Cas nuclease domains (FIG. 11) shows these three residues are invariant among orthologs, and interestingly five other positions show similar conservation, indicating they may regulate type V Cas nuclease activity. We took particular interest in A925 and 1929 because their position in documented protein structures (see e.g. Tsuchida et al. 2022, Molecular Cell 82, no. 6 (March 2022): 1199-1209. e6. doi.org / 10.1016 / j.molcel.2022.02.002. which is incorporated by reference herein for all purposes^ suggests a lysine substitution at either position may disrupt the electrostatic pocket required to coordinate Mg2+ions and likely prevent DNA cleavage (FIG.12). Analyzing the structures of the PlmCasl2e ribonucleic-protein complex also suggested that truncating part of the protein’s C-terminus may compromise nuclease activity without perturbing the PlmCasl2e:gRNA interaction, not the ability of the gRNA to hybridize with its DNA target (FIG. 12). Thus, 7 PlmCasl2e variants were assayed for nuclease activity: PlmCasl2WT, PlmCasl23A(D569A, E756A, D922A), PlmCasl2ATSL(delete aa 643-816), PlmCasl2AC(delete aa 643-980), PlmCasl2A925K, PlmCasl2I929K, and PlmCasl22K(A925K, I929K). The PlmCasl2e variants were targeted to an ectopic neomycin resistance gene in human cells, such that a nuclease active enzyme will damage the drug resistance gene and cells will die upon treatment (FIG. 13). This functional assay revealed that PlmCasl2A925K, PlmCasl22K, and PlmCasl2ACretained drug resistance much like the dead enzyme PlmCasl23A, while a fraction of PlmCasl2WTand PlmCasl21929cells lost resistance due to DNA damage and died upon treatment with the neomycin analog G418 (FIG. 14). Finally, PlmCasl2ATSLresulted in an intermediate phenotype, suggesting compromised but not ablated nuclease activity. This functional assay was validated by directly sequencing the DNA targeted by the enzymes and the percent of sequencing reads that were aberrant due to erroneous DNA damage repair following nuclease activity were quantified (FIG. 14).

[0105] Example 9.- Generation of additional CRISPRi domains

[0106] PlmCasl23A, PlmCasl2AC, and PlmCasl22Kwere fused to transcriptional repressor KRABZNF1° or transcriptional activator VPR to measure their ability to regulate gene expression relative to the well characterized SpCas9 system. In this experiment, gRNA targetingnowhere in the genome, or six different endogenous human genes were transiently transfected into human tumor cells stably expressing one of the dPlmCasl2e fusions. 48-72 hours later the RNA was harvested from the cell populations and measured via quantitative PCR (FIG. 15). This revealed that all three dPlmCasl2e variants were more potent gene regulators than dSpCas9. Three days after transfection dSpCas9 fused to KRABZNF1° reduced mRNA levels by -35% on average while the PlmCasl2e variants with the same fusion reduced mRNA levels nearly 65% (FIG. 16). Similarly, when dSpCas9 was fused to VPR cells contained ~3-fold more target mRNA, while in cells containing the PlmCasl2e variants fused to VPR mRNA expression was increased by -11-fold (FIG. 17). However, PlmCasl22Kshowed relatively lower performance among the dPlmCasl2 variants. The success of PlmCasl2ACas a gene regulator is of particular interest as it has improved activity over SpCas9, and is 20% smaller than PlmCasl2WTmaking it an easier product to deliver for genetic manipulations.

[0107] Example 10.- Generation of additional guide RNA scaffold mutants

[0108] The targeted nature of PlmCasl2e-based gene modulation relies on a synthetic RNA molecule referred to as a guide or single-guide RNA (gRNA or sgRNA). An initial molecule based on the naturally occurring bacterial sequence was used to experimentally verify this technology; however, key changes were, and continue to be, made to this molecule to respond to the unique demands of mammalian gene regulation (FIG. 18). These modifications are proposed to optimize the Casl2e gRNA for biosynthesis, stability, improved RNP complex formation, and to enable high-throughput next generation sequencing in mammalian cells and are informed in part from cryo-electron microscopy studies (Bogenhagen and Brown 1981, Antao et al. 1991, Repogle et al. 2020, Tsuchida et al. 2022). The first round of rationally designed gRNA modifications improved the potency of both gene repression (FIG. 19) and gene activation (FIG. 20) within the dPlmCasl2e gene modulating system. Either version 1.0 gRNA based on bacterial sequences or the version 2.1 gRNA with unique modifications was transfected into cell expressing PlmCasl23Afused to KRABZNF1° or VPR. Studying 12 different gRNAs targeting 6 genes suggested that a fraction of version 1 gRNAs were highly active and version 2.1 modifications did not improve their activity; however, for poorly performing version 1 gRNAs the modifications in version 2.1 significantly improved both gene inhibition (FIG. 19) and gene activation (FIG. 20).

[0109] Example 11.- Generation of Additional Activation and Effector Domains for CRISPRi

[0110] We next tested if the PlmCasl2e gene modulating system was compatible with multiple effector domains by fusing PlmCasl2e3Ato the Transcription Repressor Domain(TReD) of human protein, PPP1R8 / Nippl, or the Transcription Activation Domain (TAD) of human protein, F0X03. The transcriptional repression ability of the full length TReDppplR8and a truncated mini-TReDppplR8variant were compared to the field-standard KRABZNF1° repressor at six different human genes. Both variants provided at least equal gene silencing to KRABZNF1° as observed by similar median reductions in mRNA levels, however, when looking at specific genes the PPP1R8 domains appeared to function more robustly at a subset of genomic loci (FIG. 21). This trend was also observed for gene activation. Both VPR and TADFOXO3effectors are successful gene amplifiers and function somewhat distinctly at different genes (FIG. 22). For some targets VPR is more potent while at others TADFOXO3is stronger. For that subset of genes we observed a nearly 3-fold increase in mRNA when using VPR, but and ~6-fold increase in mRNA levels with TADFOXO3(FIG. 22).

[0111] Finally, we show that the components of these gene modulating systems are modular and can be combined together to create an improved technology. PlmCasl2eACwas fused to TReDppplR8and expressed in human cancer cells that were then transduced with version 2.1 gRNA plasmids. The combination of all three novel components gave robust reductions in target mRNA levels, such that all targets were reduced by >50% (FIG. 23). The same was true for gene activation where combining PlmCasl2eACwith TADFOXO3and version 2.1 gRNAs resulted in a dramatic increase to mRNA levels. As mentioned above, dSpCas9-VPR fusions activated ~3-fold more target mRNA, while PlmCasl2e3A-VPR with version 1.0 gRNAs stimulated mRNA expression by ~11-fold, while PlmCasl2eAC-TADFOXO3and version 2.1 gRNAs resulted in a nearly 50-fold increase to target mRNA levels (FIG. 24).

[0112] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A fusion protein comprising:(a) a nuclease-deficient class 2, type V Cas polypeptide comprising a sequence having at least 80% sequence identity to SEQ ID NO: 1, or a variant thereof;(b) an effector domain comprising a transcriptional activation domain or a transcriptional repression domain.

2. The fusion protein of claim 1, wherein said effector domain comprises a transcriptional repression domain.

3. The fusion protein of claim 1, wherein said effector domain comprises a transcriptional repression domain, wherein said transcriptional repression domain comprises a Krtippel-associated box (KRAB) domain, a chromo domain, a chromo-shadow domain, a CBX domain, a Methyl-CpG Binding Protein 2 (MECP2) domain, a SIN3 Transcription Regulator Family Member A (SIN3 A) domain, a Histone deacetylase HDT1 (HDT1) domain, a Methyl- CpG Binding Domain Protein 2 (MBD2) domain, a Protein Phosphatase 1 Regulatory Subunit 8 (PPP1R8) domain, or a Chromobox 5 (CBX5) domain, a DNA Cytosine Methyltransferase (DCM) domain, a DNA Cytosine Methyltransferase-like (DCML) domain or any combination thereof.

4. The fusion protein of claim 1, wherein said effector domain comprises a transcriptional activation domain.

5. The fusion protein of claim 1, wherein said effector domain comprises a transcriptional activation domain, wherein said transcriptional activation domain comprises a Krtippel-associated box (KRAB) domain, a transcription / trans-activating domain (TAD), a VP64 domain, a Tegument protein VP 16 (VP 16) domain, a RelA / p65 gene activating domain (p65), a Rta domain, a FOXO3 domain, a viral gene activation domain, or any combination thereof.

6. The fusion protein of claim 5, wherein said transcriptional activation domain comprises a viral gene activation domain, wherein said viral gene activation protein / domain comprises a VP64 domain, a RelA / p65 gene activating domain (p65), or a Replication and Transcription Activator (RTA) domain.

7. The fusion protein of any one of claims 1-6, wherein said fusion protein comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-5, 13- 19, 20, or a variant thereof.

8. The fusion protein of any one of claims 1-7, further comprising a nuclear localization signal (NLS).

9. The fusion protein of claim 8, wherein said NLS is between (a) and (b).

10. The fusion protein of claim 8, wherein said NLS is near to an N- or C-terminus of said fusion protein.

11. The fusion protein of any one of claims 1-10, wherein (a) is near to said N- terminus of said fusion protein and (b) is near to said C-terminus of said fusion protein.

12. The fusion protein of any one of claims 1-10, wherein (a) is near to said C- terminus of said fusion protein and (b) is near to said N-terminus of said fusion protein.

13. The fusion protein of any one of claims 1-10, wherein (a) and (b) are in order from N- to C-terminus.

14. The fusion protein of any one of claims 1-10, wherein (b) and (a) are in order from N- to C-terminus.

15. The fusion protein of any one of claims 1-14, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of at least one residue corresponding to residues 659, 756, 922 of SEQ ID NO: 1, or any combination thereof, to a non-wild-type residue.

16. The fusion protein of claim 15, wherein said non-wild-type residue is a glycine or an alanine residue.

17. The fusion protein of claim 15 or 16, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof.

18. A polynucleotide encoding the fusion protein of any one of claims 1-17.

19. A system comprising:(a) a fusion protein according to any one of claims 1-17 or a polynucleotide encoding a fusion protein according to any one of claims 1-17;(b) an engineered guide nucleic acid or a template polynucleotide encoding said engineered guide nucleic acid, wherein said engineered guide nucleic acid configured to form a complex with said nuclease-deficient class 2, type V Cas polypeptide, comprising:(i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence;(ii) a scaffold nucleic acid sequence configured to bind to said nuclease- deficient class 2, type V Cas polypeptide.

20. The system of claim 19, wherein said engineered guide nucleic acid comprises (ii) and (i) in order from 5’ to 3’.

21. The system of claim 19 or 20, wherein said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 6-12 or a variant thereof.

22. The system of claim 19 or 20, wherein said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to SEQ ID NO: 6, or a variant thereof, further comprising at least one mutation according to Table 2 relative to SEQ ID NO: 6.

23. The system of any one of claims 19-22, wherein said anti -target sequence comprises at least 16-22 nucleotides.

24. The system of any one of claims 19-22, wherein said fusion protein comprises a sequence having at least 80% sequence identity to SEQ ID NO: 14, further comprising a polypeptide having at least 80% sequence identity to SEQ ID NO: 15.

25. A polynucleotide comprising a sequence having at least 80% sequence identity to SEQ ID NO: 6, or a variant thereof, further comprising at least one mutation according to Table 2 relative to SEQ ID NO: 6.

26. The polynucleotide according to claim 25, wherein said polynucleotide comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 7-12, or a variant thereof.

27. The polynucleotide of claim 25 or 26, further comprising an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence, which anti-target sequence is oriented 3’ to said sequence having at least 80% sequence identity to SEQ ID NO: 6 or said sequence at least 90% sequence identity to any one of SEQ ID NOs: 7-12, or a variant thereof.

28. The polynucleotide of claim 27, wherein said anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

29. A polynucleotide encoding the polynucleotide of any one of claims 25-28.

30. A system, comprising:(a) a polypeptide comprising a class 2, type V Cas polypeptide sequence or a polynucleotide encoding said class 2, type V Cas polypeptide sequence;(b) a polynucleotide according to any one of claims 25-28, or a template polynucleotide encoding said polynucleotide according to any one of claims 25-28.

31. The system of claim 30, wherein said class 2, type V Cas polypeptide sequence is a class 2, type V-E polypeptide sequence.

32. The system of any one of claims 30-31, wherein said polypeptide further comprises an effector domain comprising a transcriptional activation domain or a transcriptional repression domain.

33. The system of any one of claims 30-32, wherein said polypeptide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-5, 13-19,20, or a variant thereof.

34. The system of claim 33, wherein said fusion protein comprises a sequence having at least 80% sequence identity to SEQ ID NO: 14, further comprising a polypeptide having at least 80% sequence identity to SEQ ID NO: 15.

35. The system of any one of claims 30-34, wherein said polypeptide is nuclease- deficient.

36. The system of any one of claims 30-35, wherein said polypeptide comprises a substitution of at least one residue corresponding to residues 659, 756, 922 of SEQ ID NO: 1, or any combination thereof, to a non-wild-type residue.

37. The system of claim 36, wherein said non-wild-type residue is a glycine or an alanine residue.

38. The system of any one of claims 36-37, wherein said polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof.

39. A polypeptide comprising a sequence having at least 80% identity to SEQ ID NO: 1 comprising a substitution of at least two residues corresponding to residues 659, 756, or 922 of SEQ ID NO: 1 to a non-wild-type residue.

40. The polypeptide of claim 39, wherein said non-wild-type residue is a glycine or an alanine residue.

41. The polypeptide of claim 40, wherein said polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof.

42. The polypeptide of any one of claims 39-41, wherein said polypeptide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2-5, 13-19, or a variant thereof.

43. The polypeptide of any one of claims 39-42, wherein said polypeptide further comprises an effector domain comprising a transcriptional activation domain or a transcriptional repression domain.

44. The polypeptide of any one of claims 39-43, wherein said polypeptide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-5, 12-19, or a variant thereof.

45. A polynucleotide encoding the polypeptide of any one of claims 39-44.

46. A system comprising:(a) a polypeptide according to any one of claims 39-44 or a polynucleotide encoding a polypeptide according to any one of claims 39-44;(b) an engineered guide nucleic acid or a template polynucleotide encoding said engineered guide nucleic acid, wherein said engineered guide nucleic acid is configured to form a complex with said polypeptide, comprising:(i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence;(ii) a scaffold nucleic acid sequence configured to bind to said polypeptide.

47. The system of claim 46, wherein said engineered guide nucleic acid comprises (ii) and (i) in order from 5’ to 3’.

48. The system of any one of claims 46-47, wherein said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 6- 12 or a variant thereof.

49. The system of any one of claims 46-47, wherein said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to SEQ ID NO: 6, or a variant thereof, further comprising at least one mutation according to Table 2 relative to SEQ ID NO: 6.

50. The system of any one of claims 46-49, wherein said anti-target sequence comprises at least 16-22 nucleotides.

51. The system of any one of claims 46-50, wherein said polypeptide comprises a sequence having at least 80% sequence identity to SEQ ID NO: 14, further comprising a polypeptide having at least 80% sequence identity to SEQ ID NO: 15.

52. A method of binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, comprising contacting to said cell the system of any one of claims 18- 24, 30-38, or 46-51.

53. The method of claim 52, wherein said DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to 5’-TCN-3’ to said target nucleic acid sequence.

54. The method of claim 52 or 53, wherein said contacting comprises transfecting said cell with said system.

55. The method of claim 52 or 53, wherein said contacting comprises virally transducing said cell with said system.

56. The method of claim 52 or 53, wherein said contacting comprises virally transducing said cell with one of (a) or (b) and transfecting said cell with the other of (a) or (b).

57. The method of any one of claims 52-56, wherein (a) further comprises an effector domain comprising a transcriptional repression domain, further comprising repressing transcription of said target nucleic acid sequence.

58. The method of any one of claims 52-56, wherein (a) further comprises an effector domain comprising a transcriptional activation domain, further comprising activating transcription of said target nucleic acid sequence.

59. A cell comprising the system of any one of claims 19-24, 30-38, or 46-51, the fusion protein of any one of claims 1-17, or the polynucleotide of any one of claims 18, 25-29, or 45.

60. The cell of claim 59, wherein said cell is a eukaryotic cell.

61. A vector comprising the polynucleotide of any one of claims 18, 25-29, or 45.

62. The vector of claim 61, wherein said vector is a plasmid or a viral vector.

63. The vector of claim 61, wherein said vector is a viral vector, wherein said viral vector is an adeno-associated virus (AAV) vector.

64. A fusion protein comprising:(a) a nuclease-deficient class 2, type V Cas polypeptide(i) comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 35, 36, or a variant thereof; or(ii) comprising a TSL or C-terminal domain deletion, or any combination thereof relative to SEQ ID NO: 1; and(b) an effector domain comprising or consisting of a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 40-45, or a variant thereof.

65. The fusion protein of claim 64, further comprising a nuclear localization signal (NLS).

66. The fusion protein of claim 64, wherein said NLS is between (a) and (b).

67. The fusion protein of claim 64, wherein said NLS is near to an N- or C-terminus of said fusion protein.

68. The fusion protein of any one of claims 64-67, wherein (a) is near to said N- terminus of said fusion protein and (b) is near to said C-terminus of said fusion protein.

69. The fusion protein of any one of claims 64-67, wherein (a) is near to said C- terminus of said fusion protein and (b) is near to said N-terminus of said fusion protein.

70. The fusion protein of any one of claims 64-67, wherein (a) and (b) are in order from N- to C-terminus.

71. The fusion protein of any one of claims 64-67, wherein (b) and (a) are in order from N- to C-terminus.

72. The fusion protein of any one of claims 64-71, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of residue 925 or 929 relative to SEQ ID NO: 1 to a non-wild-type residue.

73. The fusion protein of claim 72, wherein said non-wild-type residue comprises alanine, glycine, lysine, or arginine.

74. The fusion protein of claim 73, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises an A925K or an I929K mutation, or any combination thereof relative to SEQ ID NO: 1.

75. The fusion protein of claim 74, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises said A925K mutation and said I929K mutation relative to SEQ ID NO: 1.

76. The fusion protein of any one of claims 64-75, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of at least one residue corresponding to residues 659, 756, 922 of SEQ ID NO: 1, or any combination thereof, to a non-wild-type residue.

77. The fusion protein of claim 76, wherein said non-wild-type residue is a glycine or an alanine residue.

78. The fusion protein of claim 77, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises at least one mutation corresponding to a mutation of a D659A, D756A, or D922A mutation of SEQ ID NO: 1, a variant thereof, or any combination thereof.

79. The fusion protein of any one of claims 64-78, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises said sequence having at least 80% sequence identity to any one of SEQ ID NOs: 35, 36, or a variant thereof.

80. The fusion protein of any one of claims 64-78, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises said TSL or C-terminal domain deletion, or any combination thereof.

81. The fusion protein of any one of claims 64-78, wherein said effector domain comprises a sequence having at least 80% sequence identity to SEQ ID NO: 45 and comprises a mutation of at least one of the residues mutated in SEQ ID NO: 46 relative to SEQ ID NO: 45.

82. A polynucleotide encoding the fusion protein of any one of claims 64-81.

83. A system comprising:(a) a fusion protein according to any one of claims 1-17 or claims 64-81, or a polynucleotide encoding a fusion protein according to any one of claims 1-17 or claims 64-81;(b) an engineered guide nucleic acid or a template polynucleotide encoding said engineered guide nucleic acid, wherein said engineered guide nucleic acid configured to form a complex with said nuclease-deficient class 2, type V Cas polypeptide, comprising:(i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence;(ii) a scaffold nucleic acid sequence configured to bind to said nuclease- deficient class 2, type V Cas polypeptide.

84. The system of claim 83, wherein said engineered guide nucleic acid comprises (ii) and (i) in order from 5’ to 3’.

85. The system of claim 83 or 84, wherein said engineered guide nucleic acid comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 6-12 or 47-52 a variant thereof.

86. The system of any one of claims 83-85, wherein said class 2, type V Cas polypeptide sequence is a class 2, type V-E polypeptide sequence.

87. The system of any one of claims 83-86, wherein said anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

88. A polypeptide comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 35, 36, or a variant thereof or comprising a TSL or C-terminal domain deletion relative to SEQ ID NO: 1, or any combination thereof.

89. The polypeptide of claim 88, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises a substitution of residue 925 or 929 relative to SEQ ID NO: 1 to a nonwild-type residue.

90. The polypeptide of claim 88 or 89, wherein said non-wild-type residue comprises alanine, glycine, lysine, or arginine.

91. The polypeptide of claim 90, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises an A925K or an I929K mutation, or any combination thereof relative to SEQ ID NO: 1.

92. The polypeptide of claim 91, wherein said nuclease-deficient class 2, type V Cas polypeptide comprises said A925K mutation and said I929K mutation relative to SEQ ID NO: 1.

93. A polynucleotide comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 47-52 a variant thereof.

94. A polynucleotide comprising or encoding a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 47-52, or a variant thereof.

95. A system, comprising:(a) a polypeptide comprising a class 2, type V Cas polypeptide sequence or a polynucleotide encoding said class 2, type V Cas polypeptide sequence;(b) a polynucleotide according to any one of claims 25-28 or claim 94, or a template polynucleotide encoding said polynucleotide according to any one of claims 25-28 or claim 94.

96. The system of claim 95, wherein said class 2, type V Cas polypeptide comprises a polypeptide according to any one of claims 88-92.

97. The system of any one of claims 95-96, wherein said class 2, type V Cas polypeptide sequence is a class 2, type V-E polypeptide sequence.

98. The system of any one of claims 95-97, wherein said anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

99. A method of binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, comprising contacting to said cell the system of any one of claims 83- 87 or claims 95-98.

100. The method of claim 99, wherein said DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to 5’-TCN-3’ to said target nucleic acid sequence.

101. The method of claim 99 or 100, wherein said contacting comprises transfecting said cell with said system.

102. The method of any one of claims 99-101, wherein said contacting comprises virally transducing said cell with said system.

103. The method of any one of claims 99-102, wherein said contacting comprises virally transducing said cell with one of (a) or (b) and transfecting said cell with the other of (a) or (b).

104. The method of any one of claims 99-103, wherein (a) further comprises an effector domain comprising a transcriptional repression domain, further comprising repressing transcription of said target nucleic acid sequence.

105. The method of any one of claims 99-103, wherein (a) further comprises an effector domain comprising a transcriptional activation domain, further comprising activating transcription of said target nucleic acid sequence.

106. A cell comprising the system of any one of claims 83-87 or claims 95-98.

107. The cell of claim 106, wherein said cell is a eukaryotic cell.

108. A vector comprising the polynucleotide of claim 94.

109. The vector of claim 108, wherein said vector is a plasmid or a viral vector.

110. The vector of claim 109 wherein said vector is a viral vector, wherein said viral vector is an adeno-associated virus (AAV) vector.