Improved methods and compositions for CRISPR interference and activation

By introducing a fusion protein of nuclease-deficient type V Cas peptides and effector domains into the CRISPR/Cas system, the limitation of the CRISPR/Cas system in targeting specific nucleic acid sequences in mammalian cells has been overcome, achieving efficient and safe gene regulation, specific binding, and transcriptional regulation.

CN121548639APending Publication Date: 2026-02-17ALGEN BIOTECHNOLOGIES INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202480034104.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-18
Filing Date
2024-03-20
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems require strict complementary hybridization and the presence of PAM sequences when targeting specific nucleic acid sequences, which limits their application in mammalian cells. Furthermore, nuclease activity may lead to nonspecific cleavage and genomic interference.

Method used

A fusion protein containing a nuclease-deficient class 2 V-type Cas polypeptide and an effector domain was designed. By expressing it in mammalian cells, the engineered nucleic acid hybridizes with the target nucleic acid sequence to achieve specific binding and gene regulation, including transcriptional repression or activation.

Benefits of technology

This technology enables efficient and specific gene regulation targeting nucleic acid sequences in mammalian cells, reducing non-specific cleavage and improving the precision and safety of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121548639A_ABST
    Figure CN121548639A_ABST
Patent Text Reader

Abstract

Described herein are methods and compositions that utilize a CRISPR / Cas system. In some cases, the utilization includes the use of such systems for gene activation or interference.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 453,559, filed March 21, 2023, and U.S. Provisional Application No. 63 / 583,507, filed September 18, 2023, the entire contents of which are incorporated herein by reference. BACKGROUND

[0002] The CRISPR / Cas system is an RNA-mediated nuclease complex that has been described as playing a role in adaptive immune systems in microorganisms. The CRISPR / Cas system exists in its natural environment in a CRISPR (clustered regularly interspaced short palindromic repeat) operon or locus, typically comprising two parts: (i) a series of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode an RNA-based targeting element; and (ii) an open reading frame (ORF) that encodes a nuclease polypeptide guided by the RNA-based targeting element as well as ancillary proteins / enzymes. Effective nuclease targeting of a particular target nucleic acid sequence typically requires: (i) complementary hybridization between the first 6-8 nucleic acids of the target (the target seed) and the RNA guide sequence; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined region near the target seed (PAMs are typically sequences that are uncommon in the host genome). In some cases, the CRISPR / Cas system can be modified to use complementary hybridization of an RNA guide sequence to direct sequence-specific binding. SUMMARY

[0003] In some aspects, the present disclosure provides a fusion protein comprising: (a) a nuclease-deficient Class 2, Type V Cas polypeptide, the polypeptide: (i) comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 35, 36, or a variant thereof; or (ii) comprising a deletion of a TSL or C-terminal domain relative to SEQ ID NO: 1, or any combination thereof; and (b) an effector domain comprising or consisting of a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 40-45, or a variant thereof. In some embodiments, the fusion protein further comprises a nuclear localization signal (NLS). In some embodiments, the NLS is located between (a) and (b). In some embodiments, the NLS is proximal to the N-terminus or C-terminus of the fusion protein. In some embodiments, (a) is proximal to the N-terminus of the fusion protein and (b) is proximal to the C-terminus of the fusion protein. In some embodiments, (a) is proximal to the C-terminus of the fusion protein and (b) is proximal to the N-terminus of the fusion protein. In some embodiments, (a) and (b) are arranged in order from N-terminus to C-terminus. In some cases, the Cas endonuclease with reduced nuclease activity described herein is catalytically dead. In some cases, the Cas endonuclease with reduced nuclease activity described herein has at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.5% reduction in nuclease activity relative to a wild-type Cas endonuclease. In some embodiments, (b) and (a) are arranged in order from N-terminus to C-terminus. In some embodiments, the nuclease-deficient Class 2, Type V Cas polypeptide comprises a substitution of residue 925 or 929 relative to SEQ ID NO: 1 with a non-wild-type residue. In some embodiments, the non-wild-type residue comprises alanine, glycine, lysine, or arginine. In some embodiments, the nuclease-deficient Class 2, Type V Cas polypeptide comprises an A925K or I929K mutation relative to SEQ ID NO: 1, or any combination thereof. In some embodiments, the nuclease-deficient Class 2, Type V Cas polypeptide comprises the A925K mutation and the I929K mutation relative to SEQ ID NO: 1.In some embodiments, the nuclease-deficient Class 2 Type V Cas polypeptide comprises a substitution of at least one residue corresponding to residue 659, 756, 922, or any combination thereof, of SEQ ID NO: 1, with a non-wild type residue. In some embodiments, the non-wild type residue is a glycine or alanine residue. In some embodiments, the nuclease-deficient Class 2 Type V Cas polypeptide comprises at least one mutation corresponding to a D659A, D756A, or D922A mutation, a variant thereof, or any combination thereof, of SEQ ID NO: 1. In some embodiments, the nuclease-deficient Class 2 Type V Cas polypeptide comprises the sequence having at least 80% sequence identity to any one of SEQ ID NO: 35, 36, or a variant thereof. In some embodiments, the nuclease-deficient Class 2 Type V Cas polypeptide comprises a deletion of the TSL or C-terminal domain, or any combination thereof. In some embodiments, the effector domain comprises a sequence having at least 80% sequence identity to SEQ ID NO: 45, or a variant thereof, and comprises a mutation of at least one of the residues in SEQ ID NO: 46 that are mutated relative to SEQ ID NO: 45.

[0004] In some aspects, the disclosure provides a polynucleotide or nucleic acid encoding any of the fusion proteins described herein. In some embodiments, the nucleic acid is codon-optimized for expression in a mammalian or human cell.

[0005] In some aspects, the disclosure provides a system comprising: (a) any fusion protein described herein or a polynucleotide encoding any fusion protein described herein; (b) an engineered guide nucleic acid or a template polynucleotide encoding the engineered guide nucleic acid, wherein the engineered guide nucleic acid is configured to form a complex with the nuclease-deficient Class 2, Type V Cas polypeptide, the engineered guide nucleic acid comprising: (i) an anti-target nucleic acid sequence configured to hybridize to a target nucleic acid sequence; (ii) a scaffold nucleic acid sequence configured to bind to the nuclease-deficient Class 2, Type V Cas polypeptide. In some embodiments, the engineered guide nucleic acid comprises (ii) and (i) in 5’ to 3’ order. In some embodiments, the engineered guide nucleic acid comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 6-12 or 47-52, or a variant thereof. In some embodiments, the Class 2, Type V Cas polypeptide sequence is a Class 2, Type-E polypeptide sequence. In some embodiments, the anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

[0006] In some aspects, the present disclosure provides a polypeptide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 1, 35, 36, or a variant thereof, or comprising a deletion of a TSL or C-terminal domain relative to SEQ ID NO: 1, or any combination thereof. In some embodiments, the nuclease-deficient Class 2, Type V Cas polypeptide comprises a substitution of residue 925 or 929 relative to SEQ ID NO: 1 with a non-wild type residue. In some embodiments, the non-wild type residue comprises an alanine, glycine, lysine, or arginine. In some embodiments, the nuclease-deficient Class 2, Type V Cas polypeptide comprises an A925K or I929K mutation relative to SEQ ID NO: 1, or any combination thereof. In some embodiments, the nuclease-deficient Class 2, Type V Cas polypeptide comprises the A925K mutation and the I929K mutation relative to SEQ ID NO: 1.

[0007] In some aspects, the present disclosure provides a polynucleotide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 47-52, or a variant thereof.

[0008] In some aspects, the present disclosure provides a polynucleotide comprising or encoding a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to any one of SEQ ID NOs: 47-52, or a variant thereof.

[0009] In some aspects, the disclosure provides a system comprising: (a) a polypeptide comprising a Class 2, Type V Cas polypeptide sequence or a polynucleotide encoding the Class 2, Type V Cas polypeptide sequence; and (b) a polynucleotide of any of the engineered guide nucleic acids described herein or a template polynucleotide encoding any of the engineered guide nucleic acids described herein. In some embodiments, the Class 2, Type V Cas polypeptide comprises any of the polypeptides described herein. In some embodiments, the Class 2, Type V Cas polypeptide sequence is a Class 2, Type V-E polypeptide sequence. In some embodiments, the anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

[0010] In some aspects, the disclosure provides a method of binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, the method comprising contacting the cell with any of the systems described herein. In some embodiments, the DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to 5'-TCN-3' of the target nucleic acid sequence. In some embodiments, the contacting comprises transfecting the cell with the system. In some embodiments, the contacting comprises virally transducing the cell with the system. In some embodiments, the contacting comprises virally transducing the cell with one of (a) or (b) and transfecting the cell with the other of (a) or (b). In some embodiments, (a) further comprises an effector domain comprising a transcription repression domain, further comprising repressing transcription of the target nucleic acid sequence. In some embodiments, (a) further comprises an effector domain comprising a transcription activation domain, further comprising activating transcription of the target nucleic acid sequence.

[0011] In some aspects, the disclosure provides a cell comprising any of the systems described herein, any of the polynucleotides described herein, or any of the polypeptides described herein. In some embodiments, the cell is a eukaryotic cell.

[0012] In some aspects, the disclosure provides a vector comprising or encoding any of the polynucleotides or nucleic acids described herein. In some embodiments, the vector is a plasmid or a viral vector. In some embodiments, the vector is a viral vector, wherein the viral vector is an adeno-associated virus (AAV) vector.

[0013] In some aspects, the present disclosure provides a fusion protein comprising: (a) a nuclease-deficient Class 2, Type V Cas polypeptide comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity, or 100% sequence identity to SEQ ID NO: 1, or a variant thereof; (b) an effector domain comprising a transcriptional activation domain or a transcriptional repression domain. In some embodiments, the effector domain comprises a transcriptional repression domain. In some embodiments, the effector domain comprises a transcriptional repression domain, wherein the transcriptional repression domain comprises a Krüppel-associated box (KRAB) domain, a chromo domain, a chromo-shadow domain, a CBX domain, a methyl-CpG binding protein 2 (MECP2) domain, a SIN3 Transcriptional Regulator Family Member A (SIN3A) domain, a histone deacetylase HDT1 (HDT1) domain, a methyl-CpG binding domain protein 2 (MBD2) domain, a protein phosphatase 1 regulatory subunit 8 (PPP1R8) domain, or a Chromobox 5 (CBX5) domain, a DNA cytosine methyltransferase (DCM) domain, a DNA cytosine methyltransferase-like (DCML) domain, or any combination thereof. In some embodiments, the effector domain comprises a transcriptional activation domain. In some embodiments, the effector domain comprises a transcriptional activation domain, wherein the transcriptional activation domain comprises a Krüppel-associated box (KRAB) domain, a transcriptional / transactivation domain (TAD), a RelA / p65 gene activation domain (p65), a FOXO3 domain, or a viral gene activation domain, or any combination thereof. In some embodiments, the transcriptional activation domain comprises a viral gene activation domain, wherein the viral gene activation protein / domain comprises a Tegument protein VP16 (VP16) or a Replication and Transcription Activation Factor (RTA) domain. In some embodiments, the fusion protein comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-5, 13-19, or a variant thereof. In some embodiments, the fusion protein further comprises a nuclear localization signal (NLS). In some embodiments, the NLS is located between (a) and (b). In some embodiments, the NLS is proximal to the N-terminus or C-terminus of the fusion protein.In some embodiments, (a) is located near the N-terminus of the fusion protein, and (b) is located near the C-terminus of the fusion protein. In some embodiments, (a) is located near the C-terminus of the fusion protein, and (b) is located near the N-terminus of the fusion protein. In some embodiments, (a) and (b) are arranged sequentially from the N-terminus to the C-terminus. In some embodiments, (b) and (a) are arranged sequentially from the N-terminus to the C-terminus. In some embodiments, the nuclease-deficient class 2 type V Cas polypeptide comprises at least one residue corresponding to residues 659, 756, 922 or any combination thereof of SEQ ID NO: 1, substituted with a non-wild-type residue. In some embodiments, the non-wild-type residue is a glycine or alanine residue. In some embodiments, the nuclease-deficient class 2 type V Cas polypeptide comprises at least one mutation corresponding to the D659A, D756A or D922A mutation, variant thereof, or any combination thereof of SEQ ID NO: 1.

[0014] In some respects, this disclosure provides a polynucleotide encoding any of the fusion proteins described herein.

[0015] In some aspects, this disclosure provides a system comprising: (a) a fusion protein according to any aspect or embodiment described herein, or a polynucleotide encoding a fusion protein according to any aspect or embodiment described herein; (b) an engineered guiding nucleic acid or a template polynucleotide encoding said engineered guiding nucleic acid, wherein said engineered guiding nucleic acid is configured to form a complex with said nuclease-deficient class 2 type V Cas polypeptide, said engineered guiding nucleic acid comprising: (i) an anti-target nucleic acid sequence configured to hybridize with a target nucleic acid sequence; and (ii) a scaffold nucleic acid sequence configured to bind with said nuclease-deficient class 2 type V Cas polypeptide. In some embodiments, the engineered guiding nucleic acid comprises (ii) and (i) in the order from 5' to 3'. In some embodiments, the engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 6-12, or a variant thereof. In some embodiments, the engineered guiding nucleic acid comprises a sequence or a variant thereof having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity or 100% sequence identity with SEQ ID NO: 6, and further comprises at least one mutation relative to SEQ ID NO: 6 according to Table 2. In some embodiments, the anti-target sequence comprises at least 16-22 nucleotides. In some embodiments, the fusion protein comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity or 100% sequence identity with SEQ ID NO: 14, and further comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, or at least about 99% sequence identity with SEQ ID NO: 14. 15. A polypeptide having at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity or 100% sequence identity.

[0016] In some aspects, this disclosure provides a polynucleotide comprising a sequence or a variant thereof having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity or 100% with SEQ ID NO: 6, and further comprising at least one mutation relative to SEQ ID NO: 6 according to Table 2. In some embodiments, the polynucleotide comprises a sequence or a variant thereof having at least 90% sequence identity with any one of SEQ ID NO: 7-12. In some embodiments, the polynucleotide further comprises an anti-target nucleic acid sequence configured to hybridize with a target nucleic acid sequence, the anti-target sequence being located at the 3' end of the sequence having at least 80% sequence identity with SEQ ID NO: 6 or at least 90% sequence identity with any of SEQ ID NO: 7-12, or a variant thereof. In some embodiments, the anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

[0017] In some respects, this disclosure provides a polynucleotide encoding any of the polynucleotides described herein.

[0018] In some aspects, this disclosure provides a system comprising: (a) a polypeptide comprising a type 2 V Cas polypeptide sequence or a polynucleotide encoding said type 2 V Cas polypeptide sequence; and (b) a polynucleotide according to any aspect or embodiment described herein, or a template polynucleotide encoding said polynucleotide according to any aspect or embodiment described herein. In some embodiments, the type 2 V Cas polypeptide sequence is a type 2 VE polypeptide sequence. In some embodiments, the polypeptide further comprises an effector domain comprising a transcriptional activation domain or a transcriptional repression domain. In some embodiments, the polypeptide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 1-5, 13-19, or a variant thereof. In some embodiments, the fusion protein comprises a sequence having at least 80% sequence identity with SEQ ID NO: 14, and also comprises a polypeptide having at least 80% sequence identity with SEQ ID NO: 15. In some embodiments, the polypeptide is nuclease-deficient. In some embodiments, the polypeptide comprises at least one residue corresponding to residues 659, 756, 922, or any combination thereof of SEQ ID NO: 1, substituted with a non-wild-type residue. In some embodiments, the non-wild-type residue is a glycine or alanine residue. In some embodiments, the polypeptide comprises at least one mutation corresponding to the D659A, D756A, or D922A mutation, variant thereof, or any combination thereof of SEQ ID NO: 1.

[0019] In some aspects, this disclosure provides a polypeptide comprising a sequence having at least 80% identity with SEQ ID NO: 1, wherein at least two residues corresponding to residues 659, 756, or 922 of SEQ ID NO: 1 are substituted with non-wild-type residues. In some embodiments, the non-wild-type residues are glycine or alanine residues. In some embodiments, the polypeptide comprises at least one mutation corresponding to the D659A, D756A, or D922A mutation, variants thereof, or any combination thereof of SEQ ID NO: 1. In some embodiments, the polypeptide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 2-5, 13-19, or a variant thereof. In some embodiments, the polypeptide further comprises an effector domain comprising a transcriptional activation domain or a transcriptional repression domain. In some embodiments, the polypeptide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 1-5, 12-19, or a variant thereof.

[0020] In some respects, this disclosure provides a polynucleotide encoding any of the polypeptides described herein.

[0021] In some aspects, this disclosure provides a system comprising: (a) a polypeptide according to any aspect or embodiment described herein, or a polynucleotide encoding a polypeptide according to any aspect or embodiment described herein; (b) an engineered guiding nucleic acid or a template polynucleotide encoding said engineered guiding nucleic acid, wherein said engineered guiding nucleic acid is configured to form a complex with said polypeptide, said engineered guiding nucleic acid comprising: (i) an anti-target nucleic acid sequence configured to hybridize with a target nucleic acid sequence; and (ii) a scaffold nucleic acid sequence configured to bind with said polypeptide. In some embodiments, said engineered guiding nucleic acid comprises (ii) and (i) in the order from 5' to 3'. In some embodiments, said engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 6-12 or a variant thereof. In some embodiments, said engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with SEQ ID NO: 6 or a variant thereof, and further comprises at least one mutation relative to SEQ ID NO: 6 according to Table 2. In some embodiments, said anti-target sequence comprises at least 16-22 nucleotides. In some embodiments, the polypeptide comprises a sequence having at least 80% sequence identity with SEQ ID NO: 14, and also comprises a polypeptide having at least 80% sequence identity with SEQ ID NO: 15.

[0022] In some aspects, this disclosure provides a method for binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, the method comprising contacting the cell with any of the systems described herein. In some embodiments, the DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to 5'-TCN-3'- of the target nucleic acid sequence. In some embodiments, the contact comprises transfecting the cell with the system. In some embodiments, the contact comprises transducing the cell with a virus of the system. In some embodiments, the contact comprises transducing the cell with a virus of (a) or (b) and transfecting the cell with the other of (a) or (b). In some embodiments, (a) further comprises an effector domain comprising a transcriptional repression domain, further comprising repressing transcription of the target nucleic acid sequence. In some embodiments, (a) further comprises an effector domain comprising a transcriptional activation domain, further comprising activating transcription of the target nucleic acid sequence.

[0023] In some aspects, this disclosure provides a cell comprising any system, fusion protein, or polynucleotide described herein. In some embodiments, the cell is a eukaryotic cell.

[0024] In some aspects, this disclosure provides a vector comprising any of the polynucleotides described herein. In some embodiments, the vector is a plasmid or a viral vector. In some embodiments, the vector is a viral vector, wherein the viral vector is an adeno-associated virus (AAV) vector.

[0025] Other aspects and advantages of this disclosure will be readily apparent to those skilled in the art from the following detailed description, in which only exemplary embodiments of this disclosure are shown and described. It should be understood that this disclosure may take many other different forms, and that several details thereof may be modified in various obvious respects without departing from this disclosure. Therefore, the drawings and descriptions should be considered illustrative rather than restrictive. Incorporation

[0026] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the same extent that each individual publication, patent or patent application is specifically and individually cited by reference. Attached Figure Description

[0027] The novel features of the present invention are described in detail in the appended claims. A better understanding of the features and advantages of the invention can be achieved by referring to the following detailed description of exemplary embodiments utilizing the principles of the invention, along with the accompanying drawings, in which: Figure 1 A schematic diagram of the genetic unit used in the targeted regulation of mammalian gene expression using dPlmCas12e protein fusion is depicted. To regulate mammalian gene expression patterns, a nuclease-inactivated (dead) Cas12e enzyme is fused with a mammalian nuclear localization signal (NLS) and an effector domain via a variable-length linker region. The effector domain includes documented gene regulatory domains (e.g., from vertebrate and viral genomes) that repress or activate gene expression. Expression of the fusion protein is driven by a robust constitutive RNA polymerase II promoter. Simultaneously, a single Cas12e guide RNA molecule with a variable and programmable “target or spacer sequence” is expressed by a constitutive RNA polymerase III promoter. This ribonucleoprotein complex is guided by the target / spacer sequence to a site in the human genome that simultaneously contains the target sequence and an upstream (5') nucleotide motif composed of or containing a TCN.

[0028] Figure 2The nuclease-inactivating mutation in PlmCas12e is shown. The structure shown in the figure is based on Tsuchida et al. (Molecular Cell 82, Vol. 6 (March 2022): 1199-1209.e6, the full text of which is incorporated herein by reference; PDB: 7WB1), which shows the binding of Cas12e to gRNA and double-stranded DNA; the structure shows three amino acids (D659, E756, and D922; shown as spheres) clustered near the non-target DNA strand, promoting DNA cleavage through DNA binding and hydrolysis. These residues are mutated to alanine, resulting in PlmCas12e. 3A Nuclease-inactivated variant.

[0029] Figure 3 A putative model of the PlmCas12e gRNA variant (SEQ ID NO: 6-11) was depicted. The putative folding pattern of the Cas12e gRNA variant was based on electron microscopy observations of the homologous ribonucleoprotein complex (Liu et al., Nature 566, Vol. 7743 (February 2019): 218–23). The proposed Cas12e gRNA variants include substitution mutations to improve transcription in mammalian cells, thereby replacing the RNA loop sequence to adopt a more energy-favorable geometry, and elongating the stem of the hairpin structure. Potential mutation sites and possible substitution residues (e.g., the AK mutation in Table 2) are shown in bold and indicated by *.

[0030] Figure 4 A hypothetical diagram of the Cas12e gRNA variant within the ribonucleoprotein complex was drawn, revealing the motivation behind the mutation. Based on the finding that isothymidine repeat sequences lead to premature transcriptional termination, a substitution was made (top right) to improve transcription in mammalian cells (Bogenhagen and Brown, Cell 24, Vol. 1 (April 1981): 261-70, the full text of which is incorporated herein by reference). Based on structural data, the selected site is unlikely to negatively impact complex formation. Replacing the trinucleotide RNA loop with a tetranucleotide motif common in mammalian RNA sequences (bottom left) may result in more stable gRNA and / or complex formation. Cas9 gRNA exhibited enhanced activity after elongating the stem-loop (bottom right), which may have increased the binding affinity of the gRNA-Cas complex.

[0031] Figure 5 A schematic diagram depicts an experiment using transient plasmid expression to test gene regulation based on dPlmCas12e. The diagram shows the assay used to test PlmCas12e. 3AA schematic diagram illustrating the experimental method for gene repression activity of HA-2xNLS-BFP-KRAB in human cells. Lipid particles are used to deliver multiple copies of a DNA plasmid encoding a Cas12e gRNA that targets a non-natural GFP gene, one of four different natural genes, or a gene that does not target any location in the human genome. These plasmids are delivered to cells containing PlmCas12. 3A Engineered cells containing non-natural genome copies of HA-2xNLS-BFP-KRAB and GFP. All RNA molecules were extracted from cells transfected with the plasmid, and the relative transcriptional levels of the genes were determined by qPCR.

[0032] Figure 6 ( Figure 6 This demonstrates the effect of transiently expressing gRNA on PlmCas12e. 3A -KRAB targets five different genomic sites in human cells to cause gene silencing. Transient transfection with plasmids encoding matching gRNAs expressing dSpCas9-HA-2xNLS-BFP-KRAB or PlmCas12e was performed. 3A -HA-2xNLS-BFP-KRAB gRNA targets the transcription start sites of ectopic GFP reporter genes or endogenous human genes (RPA1, STK33, GUCY1A2, and CD164) in human tumor cells. Cells were harvested 48 hours later, and the amount of target RNA was determined by quantitative PCR and normalized for cells transfected with plasmids encoding non-target control gRNAs (which do not bind to the human genome). The results show the effects of PlmCas12e. 3A - Bar graph showing the effect of the KRAB construct or similar dCas9-based construct on the expression of GFP, RPA1, STK33, GUCY1A2, or CD164 RNA relative to the non-targeted control.

[0033] Figure 7 A schematic diagram depicts an experiment testing dPlmCas12e-based gene regulation via viral transduction of the host genome. The diagram shows the assay used to test PlmCas12e. 3A -NLS-KRAB ZNF10 A schematic diagram of the experimental method for gene repression activity in human cells. Non-replicating viral particles were used to insert multiple genes into the human cell genome: (a) genes conferring resistance to the toxin puromycin; and (b) Cas12e gRNA targeting the human RPA1 gene or not targeting any location in the human genome. Viral particles were formed and secreted in HEK293T cells, and then culture medium containing non-replicating viral particles was collected. Infection with PlmCas12-containing cells was performed using viral particles. 3A -NLS-KRAB ZNF10Engineered cells containing non-natural copies of the GFP genome were exposed to the toxin puromycin to isolate pure cell populations expressing gRNA. RNA was extracted from these pure populations, and relative RPA1 transcription levels were determined by qPCR.

[0034] Figure 8 The virus transduced gRNA showed that it would convert PlmCas12e 3A -KRAB targets the RPA1 locus in human cells, leading to gene silencing. Expression of dSpCas9-HA-2xNLS-BFP-KRAB or PlmCas12e 3A Human cells with the HA-2xNLS-BFP-KRAB gene were infected with lentiviral particles encoding matching gRNAs that target the transcription start site of the endogenous RPA1 gene. Cells were collected 48 hours later, and the amount of RPA1 RNA was determined by quantitative PCR and normalized for cells infected with lentiviral particles encoding a non-targeting control gRNA (which does not bind to the human genome). A bar graph showing RPA1 RNA expression relative to the non-targeting control is presented.

[0035] Figure 9 A schematic diagram depicts an experiment using transient plasmid expression to test gene regulation based on dPlmCas12e. The diagram shows the assay used to test PlmCas12e. 3A A schematic diagram illustrating the experimental method for gene activation activity of -2xNLS-VP64-p65-Rta (dPlmCas12e-VPR) in human cells. Lipid particles are used to deliver DNA plasmids encoding Cas12e gRNA, which targets six different natural genes or any location in the human genome. These plasmids are delivered to cells containing PlmCas12e. 3A Engineered cells containing 2xNLS-VP64-p65-Rta non-natural genome copies. All RNA molecules were extracted from cells transfected with the plasmid, and the relative transcriptional levels of the genes were determined by qPCR.

[0036] Figure 10This study illustrates how transient expression of gRNA targeting dPlmCas12e-VPR to six different genomic sites in human cells leads to gene activation. Human tumor cells expressing dSpCas9-2xNLS-VP64-p65-Rta or dPlmCas12e-2xNLS-VP64-p65-Rta were transiently transfected with plasmids encoding matching gRNAs that target transcription initiation sites of endogenous human genes (GCK, CD164, GUCY1A2, KDSR, PPAT, and STK17A). Cells were harvested 48 hours later, and the amount of target RNA was determined by quantitative PCR and normalized for cells transfected with plasmids encoding non-target control gRNAs (which do not bind to the human genome). Bar graphs showing the expression of GCK, CD164, GUCY1A2, KDSR, PPAT, or STK17A RNA relative to non-target control RNAs when affected by the dPlmCas12e-VPR construct or similar dCas9-based constructs are presented.

[0037] Figure 11 This diagram illustrates the PlmCas12e amino acid profile required for nuclease activity identification based on sequence homology principles. (Top) The PlmCas12e domain structure in two-dimensional sequence space. This includes the non-target DNA strand loading (NTSL) domain, the oligonucleotide binding domain (OBD), and the RuVC nuclease domain, which are sequentially broken down by the target DNA strand loading domain. (Bottom) Sequence alignment of eight different type V Cas nuclease (RuVC) domains. Highly conserved amino acid positions are highlighted, and Cas12e is specifically labeled. 3A The catalytic residues D659, E765, and D922 were mutated. Additionally, two highly conserved residues, A925 and I929, were labeled, and they are expected to also contribute to nuclease activity.

[0038] Figure 12 This illustrates the principle of ribonucleoprotein structure used to identify the amino acids required for nuclease activity in PlmCas12e. (Top) Evolutionarily restricted residues A925 and I929 were specifically identified because modeling based on cryo-electron microscopy data indicated that substitution of these positions with lysine residues could induce spatial and electrostatic interference with DNA cleavage. (Middle) RNP structural analysis also revealed the C-terminal target DNA strand loading (TSL) domain and a portion of the nuclease domain at the distal end of the PlmCas12e:gRNA interface and the gRNA:DNA hybrid, suggesting that their removal could inactivate DNA cleavage without negatively impacting DNA binding. (Bottom) Sequence illustration of potential nuclease-degrading PlmCas12e variants.

[0039] Figure 13The study demonstrated the successful transduction of cells containing chlorofluoroprotein (GFP), the neomycin resistance gene (NeoR), and ectopic genomic copies of the PlmCas12e variant using lentiviral particles encoding PlmCas12e gRNA targeting the GFP or NeoR coding region. Successful viral transduction relied on the delivery of the puromycin resistance gene along with the gRNA, which enabled cell survival after puromycin treatment. Nuclease activity was then measured either by direct DNA sequencing of the target genomic DNA or by loss of NeoR gene function due to editing.

[0040] Figure 14 Experimental results depicting the nuclease activity of the PlmCas12e variant. (Left) NeoR knockout assays show that wild-type PlmCas12e, compared to nuclease-silenced PlmCas12e, exhibits higher nuclease activity. 3A This resulted in a 40% loss of resistance. The PlmCas12e variant genome showed the entire range of nuclease activity. Dots represent the average of four experiments using a single gRNA, and bars represent the average of two NeoR gRNAs. (Right) PCR amplification of gRNA target sites in the genome, followed by deconvolutional capillary sequencing, confirmed the results of the NeoR knockout assay. Values ​​represent the average aberration readings observed in two experiments.

[0041] Figure 15 A schematic diagram depicts the experiment measuring gene regulatory activity of plmCas12e variants and effector fusions. Human tumor-derived cells were transduced with dPlmCas12e variants fused to different effector domains, and co-transduced blastomycin resistance genes were positively screened. A plasmid encoding non-targeting plmCas12e gRNA was transiently transfected into a cell subpopulation, while another cell subpopulation was transfected with a plasmid encoding gRNA targeting an endogenous human gene. RNA was extracted, converted to cDNA, and the relative amounts of mRNA molecules were measured by qPCR to determine both gene repression and activation.

[0042] Figure 16 It shows when with KRAB ZNF10 Upon fusion, the PlmCas12e nuclease-inactivating variant provides more robust transcriptional repression than dSpCas9. Human tumor-derived cells were fused with the dPlmCas12e variant or with KRAB. ZNF10Fusion dSpCas9 transduction was performed using 12 gRNAs targeting 6 endogenous human genes, and mRNA levels were determined by quantitative real-time PCR after 72 hours. (Left) The bar chart represents the average mRNA levels of two cell populations transfected with different gRNAs targeting the same genes. RNA levels are reported as a fraction relative to the same cell type transfected with gRNAs not targeting any location in the human genome. (Right) Individual gene values ​​are organized and presented as a box-and-whisker plot, where boxes represent the median and interquartile range, and whiskers indicate 1.5 times the interquartile range. Figures were compared using a paired t-test with the dSpCas9 values ​​to the dPlmCas12e variant. All values ​​were determined by three replicates.

[0043] Figure 17 This study demonstrates that the PlmCas12e nuclease-inactivating variant, when fused with VPR, provides more robust transcriptional activation than dSpCas9. Human tumor-derived cells were transduced with either the dPlmCas12e variant or dSpCas9 fused with VPR, followed by transduction with 12 gRNAs targeting six endogenous human genes, and mRNA content was determined by quantitative real-time PCR after 72 hours. (Left) Bars represent the mean mRNA content of two cell populations transfected with different gRNAs targeting the same genes. RNA content is reported as a fraction relative to the same cell type transfected with gRNAs not targeting anywhere in the human genome. (Right) Individual gene values ​​are organized and presented as a box-and-whisker plot, where boxes represent the median and interquartile range, while whiskers mark 1.5 times the interquartile range, and any points outside this range are plotted separately. The figures are compared with paired t-tests of dSpCas9 values ​​to those of the dPlmCas12e variant. All values ​​were determined by replication from two experiments.

[0044] Figure 18 The putative secondary structure and principles of mutagenesis of the PlmCas12e gRNA scaffold are illustrated. A single guide RNA (v1.0) derived from a bacterial sequence was modified in several ways to improve the molecule's transcription and stability. A modified variant (v2.1) was generated, which increased base pairing at key positions, removed polypyrimidine sequences that terminate RNA polymerase III, and added a phosphodiester backbone for contact with the Cas12e protein. Further iterations (v2.25) incorporated thermodynamically more favorable geometric features to ensure more consistent transcription initiation and removed more polypyrimidine sequences. Finally, a set of third-generation variants (v3.0) was designed to enable high-throughput next-generation sequencing of the Cas12e gRNA.

[0045] Figure 19 This shows the effect when combined with dCas12e-KRAB ZNF10Upon fusion, modification with Cas12e gRNA (v2.1) provided stronger transcriptional silencing than the bacterial-derived sequence (v1.0). Human tumor-derived cells were used with KRAB. ZNF10 Fusion PlmCas12e 3A Transduction was performed and the cells were transfected with 12 v1.0 or v2.1 gRNAs encoding the same spacer sequences targeting six endogenous human genes. mRNA levels in the populations were determined by quantitative real-time PCR after 72 hours. (Left) Bars represent the mean mRNA levels of two cell populations transfected with different gRNAs targeting the same genes. RNA levels are reported as a fraction relative to the same cell type transfected with gRNAs not targeting anywhere in the human genome. (Right) Individual gene values ​​are rounded up and presented as a box-and-whisker plot, where boxes represent the median and interquartile range, and whiskers represent 1.5 times the interquartile range. The figures shown are compared using a paired t-test between the two gRNA variants. Most v1.0 values ​​are compared with... Figure 7 The results were repeated, and all values ​​were determined by repeating the experiment three times.

[0046] Figure 20 This study demonstrates that modification of the Cas12e gRNA (v2.1) with fusion with dCas12e-NLS-VPR provides more potent transcriptional activation than the bacterial-derived sequence (v1.0). Human tumor-derived cells were treated with PlmCas12e fused with VPR. 3A Transduction was performed and the cells were transfected with 12 v1.0 or v2.1 gRNAs encoding the same spacer sequences targeting six endogenous human genes. mRNA content in the populations was determined by quantitative real-time PCR after 72 hours. (Left) The bar chart represents the average mRNA content of two cell populations transfected with different gRNAs targeting the same genes. RNA content is reported as a fraction relative to the same cell type transfected with gRNAs not targeting any location in the human genome. (Right) Individual gene values ​​are rounded up and represented as a box-and-whisker plot, where boxes represent the median and interquartile range, and whiskers mark 1.5 times the interquartile range; any points outside this range are plotted separately. The figures shown are compared using a paired t-test between the two gRNA variants. (Partial v1.0 values ​​are shown.) Figure 8 The results were repeated, and all values ​​were determined by repeating the experiment twice.

[0047] Figure 21 This demonstrates the effect of PPP1R8 transcriptional repression domain (TReD) when fused with dPlmCas12e. PPP1R8 Different truncated forms of KRAB (KRAB-ZNF10) have similar or better gene expression silencing effects than KRAB-ZNF10. Human tumor-derived cells were used with KRAB-ZNF10. ZNF10 Fusion PlmCas12e 3A Or two PPP1R8 fragments (TReD) PPP1R8or mini-TReD PPP1R8 Transduction was performed using 12 gRNAs targeting six endogenous human genes. mRNA levels in the populations were determined by quantitative real-time PCR 72 hours later. (Left) Bars represent the average mRNA levels of two cell populations transfected with different gRNAs targeting the same genes. RNA levels are reported as a fraction relative to the same cell type transfected with gRNAs not targeting anywhere in the human genome. (Right) Individual gene values ​​were rounded up and plotted as a box-and-whisker plot, where boxes represent the median and interquartile range, and whiskers represent 1.5 times the interquartile range; any points outside this range are plotted separately. The numbers shown are for comparison of TReD. PPP1R8 Variations with KRAB ZNF10 Paired t-test. Most KRAB ZNF10 Values ​​are all from Figure 7 All values ​​were determined by repeating the experiment three times.

[0048] Figure 22 This demonstrates the effect of FOXO3 transcriptional activation domain (TAD) when fused with dPlmCas12e. FOXO3 It can activate gene expression in a similar or superior manner to VPR. Human tumor-derived cells were used with PlmCas12e fused to VPR. 3A Or TAD FOXO3 Transduction was performed using 12 gRNAs targeting six endogenous human genes. mRNA levels in the populations were determined by quantitative real-time PCR 72 hours later. (Left) Bars represent the average mRNA levels of two cell populations transfected with different gRNAs targeting the same genes. RNA levels are reported as a fraction relative to the same cell type transfected with gRNAs not targeting anywhere in the human genome. (Right) Individual gene values ​​are organized and presented as a box-and-whisker plot, where boxes represent the median and interquartile range. The numbers shown are compared using TAD. FOXO3 Paired t-test for VPR. Some VPR values ​​are from... Figure 8 The values ​​were copied from the original data, and all values ​​were determined by repeating the experiment twice.

[0049] Figure 23 The use of newly described dPlmCas12e, gRNA scaffold, and repressor domain variants demonstrates, on average, more robust transcriptional silencing. Three gene silencing systems were compared in human tumor-derived cells. The Cas9 system comprises dSpCas9-KRAB. ZNF10 The Cas12e system comprises a fusion and 12 Cas9 sgRNAs targeting six endogenous human genes. The Cas12e system includes PlmCas12e. 3A -KRAB ZNF10The fusion and 12 Cas12ev1.0 sgRNAs from 6 identical human genes targeting the Cas9 target site were present, but shifted due to unique PAM restrictions. The Algen Modification System (Algen Mods) includes PlmCas12e... ΔC -TReD PPP1R8 Fusion and 12 Cas12e v2.1 sgRNAs, which contain the exact same spacer sequence as v1.0. mRNA content in the populations was determined by quantitative real-time PCR 72 hours after gRNA transfection. (Left) Bars represent the mean mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with gRNAs not targeting anywhere in the human genome. (Right) Individual gene values ​​are rounded up and presented as a box-and-whisker plot, where boxes represent the median and interquartile range, and whiskers indicate 1.5 times the interquartile range. Parentheses indicate populations used for paired t-tests for comparison. Most Cas9 and Cas12e values ​​are compared with... Figure 7 The results were repeated, and all values ​​were determined by repeating the experiment three times.

[0050] Figure 24 The use of the newly described dPlmCas12e, gRNA scaffold, and activation domain variants demonstrates, on average, more robust transcriptional activation. Three gene silencing systems were compared in human tumor-derived cells. The Cas9 system comprises a dSpCas9-VPR fusion and 12 Cas9 sgRNAs targeting six endogenous human genes. The Cas12e system includes PlmCas12e... 3A -VPR fusion and 12 Cas12e v1.0sgRNAs from 6 identical human genes targeting the Cas9 target site, but with a shift due to unique PAM restrictions. Algen Modification Systems (Algen Mods) include PlmCas12e... ΔC -TAD FOXO3 Fusion and 12 Cas12e v2.1 sgRNAs, which contain the exact same spacer sequence as v1.0. mRNA content in cell populations was determined by quantitative real-time PCR 72 hours after gRNA transfection. (Left) Bars represent the average mRNA content of two cell populations transfected with different gRNAs targeting the same gene. RNA content is reported as a fraction relative to the same cell type transfected with gRNAs not targeting anywhere in the human genome. (Right) Individual gene values ​​are rounded up and presented as a box-and-whisker plot, where boxes represent the median and interquartile range, and whiskers mark 1.5 times the interquartile range; any points outside this range are plotted separately. Parentheses indicate populations used for paired t-tests for comparison. Most Cas9 and Cas12e values ​​are compared with... Figure 8Repeat, all values ​​are determined by repeating the experiment twice. Detailed Implementation

[0051] Overview In some aspects of this paper, the CRISPR / Cas ribonucleoprotein complex (PlmCas12e) identified in the phylum Planicillium was modified to alter gene expression behavior in mammals in a targeted manner. Figure 1 By fusing two RNA components of the Plm CRISPR / Cas system and mutating the "spacer" or "targeting" sequence of this ribonucleoprotein complex, it is modified to target any site in the human genome. Furthermore, amino acids within the PlmCas12e sequence are altered to allow it to bind to DNA but prevent DNA cleavage, termed "dead" PlmCas12e or "dPlmCas12e". Figure 2 , 11 (and 12), to promote the complex to perform transcription-based functions, rather than DNA editing, homologous recombination, or DNA repair. We combined this precise genome-targeting behavior with various vertebrate (e.g., mammalian) and viral proteins / domains that have been shown to regulate nuclear localization and transcriptional activity in human cells. Figure 1 In summary, this allows for precise regulation of mammalian gene expression based on PlmCas12e.

[0052] This targeted gene regulation based on PlmCas12e relies on a synthetic RNA molecule called a guide RNA or single guide RNA (gRNA or sgRNA). Initial molecules based on naturally occurring bacterial sequences have been used to experimentally validate the technology; however, key modifications to this molecule have been tested or remain to be tested to meet the unique needs of mammalian gene regulation and laboratory experiments. Figure 3 , 4 (and 18). These modifications are proposed to optimize the biosynthesis and stability of Cas12e gRNA, improve complex formation, and enable high-throughput nucleic acid experiments in mammalian cells, and some information on these modifications is derived from cryo-electron microscopy studies.

[0053] By repressing the expression of non-natural GFP genes and endogenous genes (RPA1, STK33, GUCY1A2, CD164), a basic mammalian gene regulation tool based on PlmCas12e was tested in human tumor-derived cells. Figure 5 , 67 and 8). The dPlmCas12e protein was fused with the gene-silencing "KRAB" domain from the human ZNF10 / Kox1 protein and inserted into the genome of human tumor cells. Subsequently, the Cas12e gRNA molecule was delivered to these cells using two methods: lipid-based plasmid transfection ( Figure 5 and 6 ) or viral transduction into the host cell genome ( Figure 7 and 8 These gRNAs target non-natural GFP genes or multiple endogenous human genes (RPA1, STK33, GUCY1A2, CD164) and were compared with control gRNAs that do not bind to the human genome at any location. When the amount of mRNA transcripts of the targeted genes was compared with the non-targeted controls, we observed robust gene repression (e.g., a 25-99% reduction in mRNA expression). We also compared this behavior with the dSpCas9 system and found that in some cases, dPlmCas12e exhibited gene repression similar to or more robust than dSpCas9. Figure 6 and 8 ).

[0054] Definitions While various embodiments of the invention have been shown and described herein, those skilled in the art will understand that these embodiments are provided by way of example only. Various modifications, alterations, and substitutions can be made by those skilled in the art without departing from the spirit of the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be employed.

[0055] Unless otherwise stated, some of the methods disclosed herein have been practiced using immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA technologies. For example, see Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th ed. (2012); the series Current Protocols in Molecular Biology (edited by FM Ausubel et al.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (edited by MJ MacPherson, BD Hames, and GR Taylor (1995)); and Harlow and Lane (edited by 1988), Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th ed. (edited by RI Freshney (2010)) (the entire contents of which are incorporated herein by reference).

[0056] The singular forms “a,” “an,” and “the” used herein also include the plural forms, unless the context clearly indicates otherwise. Furthermore, when “comprising,” “having,” “having,” “with,” or variations thereof are used in the detailed description and / or claims, these terms have a similar meaning to the term “comprising” and are intended to be open-ended.

[0057] The term “about” or “approximately” means that a particular value, as determined by a person skilled in the art, is within an acceptable margin of error, depending in part on how the value is measured or determined, such as limitations of the measurement system. For example, “about” may mean within one or more standard deviations in practice in the art. Alternatively, “about” may mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0058] The term "or" as used in this article is intended to indicate an open "or".

[0059] As used herein, “cell” generally refers to a biological cell. A cell can be the basic structural, functional, or biological unit of a living organism. Cells can originate from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, single-celled eukaryotic cells, protozoan cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells (e.g., fish, amphibians, reptiles, birds, mammals), or mammalian cells (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.).

[0060] As used herein, the term "nucleotide" generally refers to a combination of a base-sugar-phosphate ester. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides can be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term "nucleotide" may include ribonucleoside triphosphates, adenosine triphosphates (ATP), uridine triphosphates (UTP), cytosine triphosphates (CTP), guanosine triphosphates (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. Nucleotides may also be chemically modified or labeled. Chemically modified mononucleotides may include biotin-dNTPs. Some non-limiting examples of biotinylated dNTPs may include biotin-dATP (e.g., biotin-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP). In some cases, the nucleotide may contain modifications to the base moiety or the phosphate moiety (e.g., modifications intended to affect the stability of the nucleotide polymer).

[0061] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably and generally refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or their analogues, and can be single-stranded, double-stranded, or multi-stranded. Polynucleotides can be exogenous or endogenous. Polynucleotides can exist in cell-free environments. Polynucleotides can be genes or segments thereof. Polynucleotides can be DNA. Polynucleotides can be RNA. Polynucleotides may contain one or more analogues (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure can be made before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acids, heteronucleic acids, morpholino, locked nucleic acids, glycol nucleic acids, threonine nucleic acids, dideoxynucleotides, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, brassinoside, and wyomin glycoside. Non-restrictive examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci identified by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes, and primers. Nucleotide sequences may be broken down by non-nucleotide components.

[0062] The terms “transfection” or “transfected” generally refer to the introduction of nucleic acids into cells via non-viral methods. Nucleic acid molecules can be gene sequences encoding complete proteins or functional parts thereof. See, for example, Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1–18.88.

[0063] The term "transduction" or "transduced" generally refers to the introduction of nucleic acids into cells via virus-based methods. Nucleic acid molecules can be gene sequences that encode complete proteins or their functional parts.

[0064] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein and generally refer to polymers consisting of at least two amino acid residues linked by peptide bonds. This term does not refer to a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or is naturally occurring. These terms apply to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer may be broken down into non-amino acid chains. These terms include amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structures (e.g., domains). These terms also cover modified amino acid polymers, for example, through disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation (e.g., conjugation with labeled components). The term “amino acid” as used herein generally refers to natural and non-natural amino acids, including but not limited to modified amino acids and amino acid analogs. Modified amino acids can include both natural and non-natural amino acids, the latter being chemically modified to include naturally occurring groups or chemical moieties on the amino acid. Amino acid analogs can refer to amino acid derivatives. The term “amino acid” includes both D-amino acids and L-amino acids.

[0065] As used in this article, the term "expression" generally refers to the process of transcribing a nucleic acid sequence or polynucleotide from a DNA template (e.g., into mRNA or other RNA transcripts), or the subsequent translation of transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotide originates from genomic DNA, expression may include the splicing of mRNA in eukaryotic cells.

[0066] As used in this article, "vector" generally refers to a macromolecule or macromolecular complex containing or associated with polynucleotides that can be used to mediate the delivery of polynucleotides into cells. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vectors. Vectors typically contain genetic elements (e.g., regulatory elements) operatively linked to a gene to promote the expression of that gene at a target.

[0067] As used in this article, "guide nucleic acid" generally refers to a nucleic acid that can hybridize with another nucleic acid. Guide nucleic acid can be RNA. Guide nucleic acid can be DNA. Guide nucleic acid can be programmed to bind to a nucleic acid sequence at a site-specific location. The target nucleic acid, or target nucleic acid, may contain nucleotides. Guide nucleic acid may contain nucleotides. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. The strand in a double-stranded target polynucleotide that is complementary to and hybridizes with the guide nucleic acid is called the complementary strand. The strand in a double-stranded target polynucleotide that is complementary to the complementary strand and therefore may not be complementary to the guide nucleic acid is called the non-complementary strand. Guide nucleic acid may contain one polynucleotide chain and may be called a "single guide nucleic acid". Guide nucleic acid may contain two polynucleotide chains and may be called a "double guide nucleic acid". Unless otherwise specified, the term "guide nucleic acid" can be open-ended, referring to both single and double guide nucleic acids. Guide nucleic acid may contain a segment that may be called an "anti-target nucleic acid sequence" or "nucleic acid targeting sequence". Guide nucleic acid may contain a sub-segment that may be called a "protein-binding segment", "protein-binding sequence", or "scaffold nucleic acid sequence".

[0068] When used in the context of engineered guide nucleic acids, the term "scaffold nucleic acid sequence" can generally refer to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity or similarity to a wild-type example sequence, configured to form a complex with a Cas polypeptide (e.g., a direct repeat sequence). "Scaffold nucleic acid sequence" can also refer to modified forms of a scaffold nucleic acid sequence (e.g., a direct repeat sequence), which can include nucleotide changes (e.g., deletions, insertions, or substitutions), variants, mutations, or chimeras.

[0069] In the context of two or more nucleic acid or polypeptide sequences, the term "sequence identity" or "percentage identity" generally refers to the fact that two (e.g., in paired alignments) or more (e.g., in multiple sequence alignments) sequences are identical, or have a specified percentage of identical amino acid residues or nucleotides, when compared and aligned within a local or global comparison window using sequence comparison algorithms to achieve maximum correspondence. Sequence comparison algorithms suitable for polypeptide sequences include, for example, BLASTP, using parameters of a word length (W) of 3 and an expected value I of 10; and the BLOSUM62 scoring matrix, setting a gap penalty of 11, an extension penalty of 1, and using conditional composition scoring matrix adjustments for polypeptide sequences longer than 30 residues.

[0070] This disclosure includes proteins, nucleic acids, or variants thereof. In some cases, such variants have at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, to about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity or 100% sequence identity with a parent protein or nucleic acid or the SEQ ID NO specifying the parent protein or nucleic acid.

[0071] This disclosure includes variants of any polypeptide described herein that have one or more conserved amino acid substitutions. Such conserved substitutions can be made in the amino acid sequence of the polypeptide without disrupting its three-dimensional structure or function. Conservative substitutions can be achieved by substituting amino acids of similar hydrophobicity, polarity, and R-chain length with each other. Additionally or alternatively, conserved substitutions can be identified by locating amino acid residues that are mutated (e.g., non-conserved residues) between species but do not alter the essential function of the encoded protein, by comparing aligned sequences of homologous proteins from different species. Such conserved substitution variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity or 100% sequence identity with any of the polypeptide sequences described herein.

[0072] In some embodiments, such conserved substitution variants are functional variants. These functional variants may contain a substituted sequence such that the activity of the key active site residues of the endonuclease is not compromised. In some embodiments, any functional variant of the protein described herein lacks the substitution of at least one conserved or functional residue.

[0073] Conserved substitutions of functionally similar amino acids are available from a variety of references (e.g., see Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993), which is incorporated herein by reference in its entirety). The following eight groups contain amino acids that are conserved in their substitutions for each other: 1) Alanine (A), glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine ​​(C), Methionine (M).

[0074] In some cases, the Cas endonuclease In described herein may comprise variants having one or more nuclear localization sequences (NLS). The NLS may be located near the N-terminus or C-terminus of the endonuclease. The NLS may be appended to the N-terminus or C-terminus of any Cas endonuclease sequence described herein, or appended to the N-terminus or C-terminus of a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any Cas endonuclease sequence. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may be the nucleolar phosphatase NLS. The NLS may be a bipartite fusion of two different NLS sequences. The NLS may contain a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity with any of SEQ ID NO: 65-80. The NLS may contain a sequence substantially identical to any of SEQ ID NO: 65-80.

[0075] In some aspects, this disclosure includes any polypeptide or nucleic acid described in Table 1. In some cases, this disclosure includes nucleic acids encoding any polypeptide or nucleic acid described in Table 1. In some cases, this disclosure includes any polypeptide domain identified as a single entity in Table 1, or any combination thereof.

[0076] Table 1: Sequences of the genes and components described in this article

[0077] Table 2: Description of single or combined mutations in PlmCas12e gRNA as described in this article

[0078] Table 3: Sequences of the guiding RNA targeting sequences (e.g., anti-target sequences) described in this article

[0079]

[0080] Table 4: Example NLS sequences that can be used with the Cas effector according to this disclosure

[0081] Example Example 1. Design and generation of nucleic acid reagents Based on experiments on Cas12e isolated from *Proteus delta* (Nature 566, Vol. 7743 (February 2019): 218–23, which is incorporated herein by reference for all purposes), three amino acids (D659, E756, and D922) in Cas12e (PlmCas12e, SEQ ID NO: 1) isolated from *Plomyces cerevisiae* were mutated to alanine (SEQ ID NO: 2). These mutations inactivate or “die” the nuclease, hence the name “PlmCas12e”. 3A “Cas12e” 3A The fusion protein is named “dPlmCas12e”, “dPlmCas12e”, or “dCas12e”. The corresponding DNA sequence is codon-optimized for expression in human cells and inserted into a third-generation lentiviral packaging plasmid, resulting in a protein fusion containing the following components: dPlmCas12e, a hemagglutinin epitope tag (HA), two nuclear localization signals (NLS), blue fluorescent protein (BFP), and a KRAB domain from the human protein ZNF10 / Kox1. This fusion protein is called “dPlmCas12e-HA-2xNLS-BFP-KRAB” or “dPlmCas12e-KRAB”. Although dPlmCas12e… The -HA-2xNLS-BFP-KRAB fusion was the primary unit tested experimentally, but the fusion of dPlmCas12e, two nuclear localization signals, and effector domains (such as KRAB) was sufficient for gene regulatory units. Subsequently, for this purpose, lentiviral packaging plasmid DNA encoding a separately expressed blastcin S resistance gene was inserted to allow enrichment of successfully transduced cells. As a positive control, an equivalent lentiviral packaging plasmid containing a nuclease-deficient or "disappointing" mutant (dSpCas9) of Cas9 from Streptococcus pyogenes was generated. This protein product was named "dSpCas9-HA-2xNLS-BFP-KRAB" or "dSpCas9-KRAB".

[0082] Lentiviral packaging plasmids were also generated to express PlmCas12e and SpCas9 guide RNA (gRNA), sometimes referred to as single guide RNA (sgRNA). The specific DNA-binding activity of PlmCas12e and SpCas9 naturally depends on a pair of annealed RNA molecules, tracrRNA and crRNA. To target the human genome, the sequences of SpCas9 or PlmCas12e tracrRNA and crRNA are fused with a synthetic RNA sequence to generate an sgRNA inserted downstream of the human U6 promoter. The PlmCas12e sgRNA contains a constant 101-nucleotide scaffold followed by a variable 20-nucleotide targeting or spacer sequence at the 3' end of the molecule. Like other Cas enzymes, the DNA targeting site of dPlmCas12e-KRAB is specified by the sequence of the protospacer adjacent motif (PAM) and the targeting sequence of the closely adjacent gRNA. To design targets for the PlmCas12e gRNA, we identified a 23-nucleotide sequence beginning with a 5' PAM sequence TCN followed by 20 nucleotides that constitute a unique sequence in the human genome (e.g., TCNNNNNNNNNNNNNNNNNNNNNNNN). The last 20 nucleotides, excluding the PAM (5' TCN), contain a target or spacer sequence inserted into the 3' end of the Cas12e sgRNA. For both SpCas9 and PlmCas12e sgRNA sequences, universal lentiviral packaging plasmids were generated that produce gRNAs that do not target any location in the human genome. These vectors serve as non-target controls and also as scaffolds where different target sequences can be easily interchanged. In addition to the gRNA, these vectors also produce puromycin N-acetyltransferase (puroR), thus conferring resistance to puromycin toxin. Using this method, we generated SpCas9 gRNA plasmids targeting four human genes (RPA1, STK33, GUCY1A2, and CD164) and the GFP gene. We also generated PlmCas12e gRNA plasmids that target these same genes.

[0083] Example 2. Generation of cell reagents Replication-deficient lentiviral particles were generated in Hek293T cells by co-transfecting the Cas-effector fusion donor plasmid with plasmids encoding the env gene from vesicular stomatitis virus type G and the gag / pol gene from human immunodeficiency virus type 1. These viral particles were used to insert DNA encoding these fusion proteins into the genome of human tumor-derived cells. In some cases, these cells also contained non-endogenous copies of green fluorescent protein (GFP) and the neomycin resistance gene (NeoR). Successfully transduced cells were isolated by treating the cell population with blastomycin S. Replication-deficient lentiviral particles targeting gRNA vectors for GFP and NeoR were also generated using the same method. Human tumor-derived cells expressing the above-described PlmCas12e variant were exposed to viral particles containing gRNA, and then successfully modified cells were screened by puromycin treatment.

[0084] Example 3. Testing dPlmCas12e-mediated gene regulation using transient plasmid expression Following the manufacturer's instructions, Lipofectamine 3000 was used to transiently transfect human tumor cells expressing dSpCas9-KRAB and non-endogenous GFP genes with SpCas9 gRNA encoding GFP, RPA1, STK33, GUCY1A2, or CD164, or with a lentiviral packaging plasmid encoding a non-targeting control. Simultaneously, Lipofectamine 3000 was used to transiently transfect human tumor cells expressing dPlmCas12e-KRAB and non-endogenous GFP genes with PlmCas12e gRNA encoding GFP, RPA1, STK33, GUCY1A2, or CD164, or with a lentiviral packaging plasmid encoding a non-targeting control. Cells were incubated for 48 hours, and then all cellular RNA was collected, and polyadenylated mRNA was converted to DNA. Subsequently, quantitative PCR was used to measure the relative abundance of mRNA molecules transcribed from the GFP gene in cells transfected with the plasmid encoding the non-targeting control and cells transfected with the plasmid encoding the GFP-targeting gRNA. Figure 6 The same experiment was performed by quantitatively measuring the relative abundance of mRNA molecules transcribed from four other target genes (RPA1, STK33, GUCY1A2, CD164) in cells transfected with plasmids encoding non-target control gRNA and cells transfected with plasmids encoding gene-specific target gRNA. Figure 6 ).

[0085] Example 4 - Testing dPlmCas12e-mediated gene regulation using viral transduction into the host genome Replication-deficient lentiviral particles encoding gRNA were generated in HEK293T cells by co-transfecting gRNA plasmids with plasmids encoding the env gene from vesicular stomatitis virus type G and the gag / pol gene from human immunodeficiency virus type 1. Viral particles were generated using gRNAs targeting both SpCas9 and RPA1. Similarly, viral particles were generated using gRNAs targeting both PlmCas12e and RPA1.

[0086] These replication-defective viral particles were used to insert DNA encoding one of these gRNAs and puroR into the genome of tumor cells expressing non-endogenous GFP genes and dSpCas9-KRAB or dPlmCas12e-KRAB. Cell populations containing gRNAs in their genomes were enriched by exposing cells to puromycin 48 hours after exposure to the viral particles. Five days later, all cellular RNA was collected, and polyadenylated mRNA was converted to DNA. The relative abundance of RPA1 mRNA molecules was then measured using quantitative PCR in cells containing genome copies of the non-targeted control gRNA and cells containing genome copies of the RPA1-targeting gRNA. Figure 8 ).

[0087] Example 5. Testing dPlmCas12e-mediated gene activation using transient plasmid expression. A mammalian gene regulation tool based on PlmCas12e was tested in human tumor-derived cells by activating the expression of six endogenous genes: CD164, GCK, GUCY1A2, STK17A, KDSR, and PPAT. The dPlmCas12e protein was fused with three gene activation domains: VP64 (a 13-amino acid quadruple repeat of the herpes simplex virus protein VP16, which has been documented as a recruitment transcription initiation factor for RNA polymerase II), p65 / RelA (a 201-amino acid sequence derived from the human gene RELA, which heterodimerizes with the transcription factor NF-κB to activate gene expression), and Rta (a 190-amino acid sequence isolated from the Rta protein encoded by EBV R, which has also been documented as a recruitment transcription factor specific to RNA polymerase II) (SEQ ID NO: 20). The fused PlmCas12e gRNA molecule was then delivered to these cells using lipid-based plasmid transfection. Figure 10These gRNAs target endogenous human genes (CD164, GCK, GUCY1A2, STK17A, KDSR, and PPAT, SEQ ID NO: 21-26) and are compared with control gRNAs that do not bind to the human genome at any location. Robust gene activation (e.g., mRNA molecule increases ranging from 2.6 to 1800-fold) was observed when mRNA transcripts from the gene targets were compared with non-targeted controls. When this level of activation was compared with a similar dSpCas9 system (gRNA SEQ ID NO: 27-32) targeting the same genes fusions of VP64, p65 / RelA, and Rta, the dPlmCas12e module was found to be a more effective activator. In five of the six genes, dPlmCas12e stimulated transcriptional levels 6-270-fold higher than dSpCas9. Figure 10 In summary, we found that dPlmCas12e-mediated gene activation was successful in all test cases, and that dPlmCas12e was superior to dSpCas9 in transcriptional amplification for many targets.

[0088] Example 6 - Detection of PlmCas12e nuclease activity (general procedure) Genetically modified cells containing GFP, NeoR, and PlmCas12e variants were infected with lentiviral particles containing one of two PlmCas12e gRNAs targeting the coding region of NeoR. Successful integration of these gRNAs was ensured by puromycin treatment. After puromycin selection, cells were incubated for 3 days to allow NeoR editing. Cells were divided into two populations: one exposed to the neomycin analog G418, and the other untreated. After 3 days, the effect of G418 on cell proliferation was determined by cell count. Nuclease-active variants were configured to target and edit, sometimes even delete, the NeoR gene, making it sensitized to the drug; while nuclease-dead variants were configured to retain resistance and be unaffected by the drug.

[0089] Untreated populations were also collected, and their genomic DNA was extracted. Using deconvolution of PCR amplification and capillary sequencing reads, we were able to determine the fraction of genomic DNA containing aberrant sequences due to nuclease editing. This same procedure was repeated for cells expressing PlmCas12e gRNA targeting GFP.

[0090] Example 7 - Testing dPlmCas12e-mediated gene regulation using transient plasmid expression (general procedure) Using Lipofectamine 3000, lentiviral packaging plasmids encoding SpCas9 or PlmCas12 egRNAs targeting 10 endogenous human genes, or non-target controls, were transiently transfected into tumor cells expressing Cas-effector fusion proteins. Cells were incubated for 72 hours, and all cellular RNA was collected as crude lysates. Polyadenylated mRNA was specifically converted to complementary DNA (cDNA). The amount of cDNA was quantified using fluorescence-based qPCR in an Agilent AriaMX real-time PCR system. This procedure was performed on cells transfected with plasmids encoding non-target controls and cells transfected with plasmids encoding targeting gRNAs to determine the relative abundance of target mRNA molecules. This procedure was repeated for consecutive Cas-effector fusions and gRNA variants.

[0091] Example 8 Design of mutants for other nuclease defects in .-PlmCas12e (e.g., nuclease death, nuclease inactivation) The above text identified three negatively charged amino acids (D659, E756, and D922) in PlmCas12e as coordinating Mg. 2+ Ions and residues essential for promoting DNA hydrolysis. Consistent with its proven function, the sequences of various V-type Cas nuclease domains were compared ( Figure 11 The results show that these three residues remain unchanged in orthologs, and interestingly, five other positions also show similar conservation, suggesting they regulate type V Cas nuclease activity. We are particularly interested in A925 and I929 because their positions in documented protein structures (see, for example, T Tsuchida et al., 2022, Molecular Cell 82, Vol. 6 (March 2022): 1199-1209.e6. doi.org / 10.1016 / j.molcel.2022.02.002. This literature is incorporated herein by reference for all purposes) suggest that lysine substitution at any position can disrupt the coordination of Mg. 2+ The electrostatic pocket required by ions and may prevent DNA cleavage. Figure 12 Analysis of the structure of the PlmCas12e ribonucleic acid-protein complex also showed that the truncated C-terminus of the protein can impair nuclease activity without interfering with PlmCas12e:gRNA interaction or affecting the ability of gRNA to hybridize with its DNA target. Figure 12 Therefore, the nuclease activities of seven PlmCas12e variants were determined: PlmCas12 WT ,PlmCas12 3A (D569A, E756A, D922A), PlmCas12ΔTSL (Delete amino acids 643-816), PlmCas12 ΔC (Delete amino acids 643-980), PlmCas12 A925K ,PlmCas12 I929K and PlmCas12 2K (A925K, I929K). The PlmCas12e variant targets the ectopic neomycin resistance gene in human cells, causing an enzyme with nuclease activity to destroy the resistance gene, and the cells die after treatment. Figure 13 This functional assay reveals that PlmCas12 A925K ,PlmCas12 2K and PlmCas12 ΔC It retains the death enzyme PlmCas12 3A Similar drug resistance, and part of PlmCas12 WT and PlmCas12 I929 Cells lose drug resistance due to DNA damage and die after treatment with the neomycin analog G418. Figure 14 Finally, PlmCas12 ΔTSL This resulted in an intermediate phenotype, indicating impaired but not completely eliminated nuclease activity. This functional assay was validated by directly sequencing the DNA targeted by the enzyme, and the percentage of aberrant sequencing reads resulting from erroneous DNA damage repair following nuclease activity was quantified. Figure 14 ).

[0092] Example 9 - Generate additional CRISPRi structure fields PlmCas12 3A ,PlmCas12 ΔC and PlmCas12 2K With transcriptional repressor KRAB ZNF10 Alternatively, VPR fusions of transcriptional activators were used to measure their ability to regulate gene expression relative to the well-characterized SpCas9 system. In this experiment, gRNAs not targeting any location in the genome or six different endogenous human genes were transiently transfected into human tumor cells stably expressing one of the dPlmCas12e fusions. RNA was harvested from the cell population 48–72 hours later and measured by quantitative PCR. Figure 15 This indicates that all three dPlmCas12e variants have stronger gene regulatory capabilities than dSpCas9. Three days post-transfection, with KRAB... ZNF10 The fused dSpCas9 reduced mRNA levels by an average of about 35%, while the PlmCas12e variant with the same fusion reduced mRNA levels by nearly 65%.Figure 16 Similarly, when dSpCas9 is fused with VPR, the cells contain approximately 3 times more target mRNA, while in cells containing the PlmCas12e variant fused with VPR, mRNA expression increases to approximately 11 times. Figure 17 However, PlmCas12 2K The dPlmCas12 variant exhibited relatively poor performance. ΔC Its success as a gene regulator is particularly interesting because it exhibits higher activity than SpCas9 and compared to PlmCas12. WT It is 20% narrower, making it a more easily delivered product for use in genetic manipulation.

[0093] Example 10 - Generate additional guide RNA scaffold mutants PlmCas12e-based gene regulation relies on a synthetic RNA molecule called a guide RNA or single guide RNA (gRNA or sgRNA). Initial molecules based on naturally occurring bacterial sequences were used to experimentally validate the technology; however, to meet the unique needs of mammalian gene regulation, this molecule has been and will continue to undergo crucial modifications. Figure 18 These modifications are proposed to optimize the biosynthesis and stability of Cas12e gRNA, improve RNP complex formation, and enable high-throughput next-generation sequencing in mammalian cells. Information on these modifications is partly derived from cryo-electron microscopy studies. (Bogenhagen and Brown 1981, Antao et al. 1991, Repogle et al. 2020, Tsuchida et al. 2022) The first round of rationally designed gRNA modification enhanced gene repression within the dPlmCas12e gene regulatory system. Figure 19 ) and gene activation ( Figure 20 The efficacy of both. Transfection of either version 1.0 gRNA based on bacterial sequence or version 2.1 gRNA with unique modifications into cells expressing KRAB... ZNF10 Or VPR fusion of PlmCas12 3A In cells. Studies of 12 different gRNAs targeting 6 genes showed that some version 1 gRNAs had high activity, while version 2.1 modifications did not improve their activity; however, for the poorly performing version 1 gRNAs, modifications in version 2.1 significantly improved gene repression (…). Figure 19 ) and gene activation ( Figure 20 )both.

[0094] Example 11 - Generate additional activation and effector domains for CRISPRi. We will next use PlmCas12e 3AFusion with the transcriptional repression domain (TReD) of human protein PPP1R8 / Nipp1 or the transcriptional activation domain (TAD) of human protein FOXO3 was used to test whether the PlmCas12e gene regulatory system is compatible with multiple effector domains. The full-length TReD was then fused. PPP1R8 and truncated mini-TReD PPP1R8 The variant's transcriptional repression ability is comparable to industry-standard KRAB. ZNF10 The repressors were compared across six different human genes. Both variants provided similarities to KRAB. ZNF10 At least the same gene silencing effect, as observed through similar median reductions in mRNA levels, however, when observing specific genes, the PPP1R8 domain appears to play a more robust role at certain genomic sites. Figure 21 This trend was also observed in gene activation. VPR and TAD FOXO3 The effectors are all successful gene amplicones, and their effects on different genes vary slightly. Figure 22 For some targets, VPR is more effective, while for others, TAD is more effective. FOXO3 Stronger. For this part of the gene, we observed that mRNA increased nearly 3-fold when using VPR, while using TAD... FOXO3 mRNA levels increased to approximately 6-fold. Figure 22 ).

[0095] Finally, we showed that the components of these gene regulatory systems are modular and can be combined to create an improved technology. PlmCas12e ΔC With TReD PPP1R8 The components were fused and expressed in human cancer cells, which were then transduced using a version 2.1 gRNA plasmid. This combination of three novel components significantly reduced target mRNA levels, resulting in a reduction of more than 50% across all targets. Figure 23 The same applies to gene activation, specifically PlmCas12e. ΔC With TAD FOXO3 Binding to version 2.1 gRNA led to a significant increase in mRNA levels. As mentioned above, dSpCas9-VPR fusion activated approximately 3 times the target mRNA, while PlmCas12e with version 1.0 gRNA... 3A -VPR stimulated approximately 11-fold increase in mRNA expression, while PlmCas12e ΔC -TAD FOXO3 Version 2.1 of gRNA led to a nearly 50-fold increase in target mRNA levels. Figure 24 ).

[0096] While preferred embodiments of the invention have been shown and described herein, it will be understood by those skilled in the art that these embodiments are provided by way of example only. Many variations, modifications, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The following claims are intended to define the scope of the invention and cover methods and structures within the scope of these claims and their equivalents.

Claims

1. A fusion protein, said fusion protein comprising: (a) A nuclease-deficient class 2 type V Cas polypeptide, or a variant thereof, containing a sequence having at least 80% sequence identity with SEQ ID NO: 1; (b) an effector domain, which includes a transcriptional activation domain or a transcriptional repression domain.

2. The fusion protein according to claim 1, wherein the effector domain comprises a transcriptional repression domain.

3. The fusion protein according to claim 1, wherein the effector domain comprises a transcriptional repression domain, wherein the transcriptional repression domain comprises a Krüppel-related box (KRAB) domain, a cromo domain, a cromo shadow domain, a CBX domain, a methyl-CpG binding protein 2 (MECP2) domain, a SIN3 transcription regulator family member A (SIN3A) domain, a histone deacetylase HDT1 (HDT1) domain, a methyl-CpG binding domain protein 2 (MBD2) domain, a protein phosphatase 1 regulatory subunit 8 (PPP1R8) domain, or a Chromobox 5 (CBX5) domain, a DNA cytosine methyltransferase (DCM) domain, a DNA cytosine methyltransferase-like (DCML) domain, or any combination thereof.

4. The fusion protein according to claim 1, wherein the effector domain comprises a transcriptional activation domain.

5. The fusion protein according to claim 1, wherein the effector domain comprises a transcriptional activation domain, wherein the transcriptional activation domain comprises a Krüppel-associated box (KRAB) domain, a transcriptional / trans-activation domain (TAD), a VP64 domain, a mesothelial protein VP16 (VP16) domain, a RelA / p65 gene activation domain (p65), an Rta domain, a FOXO3 domain, a viral gene activation domain, or any combination thereof.

6. The fusion protein according to claim 5, wherein the transcription activation domain comprises a viral gene activation domain, wherein the viral gene activation protein / domain comprises a VP64 domain, a RelA / p65 gene activation domain (p65), or a replication and transcription activator (RTA) domain.

7. The fusion protein according to any one of claims 1-6, wherein the fusion protein comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 3-5, 13-19, 20, or a variant thereof.

8. The fusion protein according to any one of claims 1-7 further comprises a nuclear localization signal (NLS).

9. The fusion protein according to claim 8, wherein the NLS is located between (a) and (b).

10. The fusion protein of claim 8, wherein the NLS is located near the N-terminus or C-terminus of the fusion protein.

11. The fusion protein according to any one of claims 1-10, wherein (a) is located near the N-terminus of the fusion protein, and (b) is located near the C-terminus of the fusion protein.

12. The fusion protein according to any one of claims 1-10, wherein (a) is located near the C-terminus of the fusion protein, and (b) is located near the N-terminus of the fusion protein.

13. The fusion protein according to any one of claims 1-10, wherein (a) and (b) are arranged sequentially from the N-terminus to the C-terminus.

14. The fusion protein according to any one of claims 1-10, wherein (b) and (a) are arranged sequentially from the N-terminus to the C-terminus.

15. The fusion protein according to any one of claims 1-14, wherein the nuclease-deficient type 2 V Cas polypeptide comprises at least one residue corresponding to residues 659, 756, 922 or any combination thereof of SEQ ID NO: 1, which is substituted with a non-wild-type residue.

16. The fusion protein of claim 15, wherein the non-wild-type residue is a glycine or alanine residue.

17. The fusion protein according to claim 15 or 16, wherein the nuclease-deficient type 2 V Cas polypeptide comprises at least one mutation corresponding to the D659A, D756A or D922A mutation of SEQ ID NO: 1, its variants or any combination thereof.

18. A polynucleotide encoding a fusion protein according to any one of claims 1-17.

19. A system comprising: (a) A fusion protein according to any one of claims 1-17 or a polynucleotide encoding a fusion protein according to any one of claims 1-17; (b) an engineered guiding nucleic acid or a template polynucleotide encoding the engineered guiding nucleic acid, wherein the engineered guiding nucleic acid is configured to form a complex with the nuclease-deficient class 2 type V Cas polypeptide, the engineered guiding nucleic acid comprising: (i) an anti-target nucleic acid sequence, wherein the anti-target nucleic acid sequence is configured to hybridize with a target nucleic acid sequence; (ii) A scaffold nucleic acid sequence configured to bind to the nuclease-deficient type 2 V Cas polypeptide.

20. The system of claim 19, wherein the engineered guiding nucleic acid comprises (ii) and (i) sequentially from 5' to 3'.

21. The system of claim 19 or 20, wherein the engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 6-12, or a variant thereof.

22. The system of claim 19 or 20, wherein the engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with SEQ ID NO: 6, or a variant thereof, and further comprises at least one mutation relative to SEQ ID NO: 6 according to Table 2.

23. The system according to any one of claims 19-22, wherein the anti-target sequence comprises at least 16-22 nucleotides.

24. The system according to any one of claims 19-22, wherein the fusion protein comprises a sequence having at least 80% sequence identity with SEQ ID NO: 14, and further comprises a polypeptide having at least 80% sequence identity with SEQ ID NO:

15.

25. A polynucleotide comprising a sequence having at least 80% sequence identity with SEQ ID NO: 6, or a variant thereof, and further comprising at least one mutation relative to SEQ ID NO: 6 according to Table 2.

26. The polynucleotide of claim 25, wherein the polynucleotide comprises a sequence having at least 90% sequence identity with any one of SEQ ID NO: 7-12, or a variant thereof.

27. The polynucleotide of claim 25 or 26 further comprises an anti-target nucleic acid sequence configured to hybridize with a target nucleic acid sequence, the anti-target sequence being located on the 3' side of the sequence having at least 80% sequence identity with SEQ ID NO: 6 or at least 90% sequence identity with any one of SEQ ID NO: 7-12 or a variant thereof.

28. The polynucleotide of claim 27, wherein the anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

29. A polynucleotide encoding a polynucleotide according to any one of claims 25-28.

30. A system comprising: (a) A polypeptide containing the type V Cas polypeptide sequence or a polynucleotide encoding the type V Cas polypeptide sequence; (b) The polynucleotide according to any one of claims 25-28, or the template polynucleotide encoding the polynucleotide according to any one of claims 25-28.

31. The system according to claim 30, wherein the type 2 V-type Cas polypeptide sequence is a type 2 VE-type polypeptide sequence.

32. The system according to any one of claims 30-31, wherein the polypeptide further comprises an effector domain, the effector domain comprising a transcriptional activation domain or a transcriptional repression domain.

33. The system according to any one of claims 30-32, wherein the polypeptide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 1-5, 13-19, 20, or a variant thereof.

34. The system of claim 33, wherein the fusion protein comprises a sequence having at least 80% sequence identity with SEQ ID NO: 14, and further comprises a polypeptide having at least 80% sequence identity with SEQ ID NO:

15.

35. The system according to any one of claims 30-34, wherein the polypeptide is nuclease-deficient.

36. The system according to any one of claims 30-35, wherein the polypeptide comprises at least one residue corresponding to residues 659, 756, 922 or any combination thereof of SEQ ID NO: 1 being substituted with a non-wild-type residue.

37. The system of claim 36, wherein the non-wild-type residue is a glycine or alanine residue.

38. The system according to any one of claims 36-37, wherein the polypeptide comprises at least one mutation corresponding to the D659A, D756A or D922A mutation, variant thereof or any combination thereof of SEQ ID NO:

1.

39. A polypeptide comprising a sequence having at least 80% identity with SEQ ID NO: 1, comprising at least two residues corresponding to residues 659, 756 or 922 of SEQ ID NO: 1 being substituted with non-wild-type residues.

40. The polypeptide of claim 39, wherein the non-wild-type residue is a glycine or alanine residue.

41. The polypeptide of claim 40, wherein the polypeptide comprises at least one mutation corresponding to the D659A, D756A or D922A mutation, variant thereof or any combination thereof of SEQ ID NO:

1.

42. The polypeptide according to any one of claims 39-41, wherein the polypeptide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 2-5, 13-19, or a variant thereof.

43. The polypeptide according to any one of claims 39-42, wherein the polypeptide further comprises an effector domain, the effector domain comprising a transcriptional activation domain or a transcriptional repression domain.

44. The polypeptide according to any one of claims 39-43, wherein the polypeptide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 1-5, 12-19, or a variant thereof.

45. A polynucleotide encoding a polypeptide according to any one of claims 39-44.

46. ​​A system comprising: (a) The polypeptide according to any one of claims 39-44, or the polynucleotide encoding the polypeptide according to any one of claims 39-44; (b) an engineered guiding nucleic acid or a template polynucleotide encoding the engineered guiding nucleic acid, wherein the engineered guiding nucleic acid is configured to form a complex with the polypeptide, the engineered guiding nucleic acid comprising: (i) an anti-target nucleic acid sequence, wherein the anti-target nucleic acid sequence is configured to hybridize with a target nucleic acid sequence; (ii) A scaffold nucleic acid sequence configured to bind to the polypeptide.

47. The system of claim 46, wherein the engineered guiding nucleic acid comprises (ii) and (i) sequentially from 5' to 3'.

48. The system according to any one of claims 46-47, wherein the engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with any one of SEQ ID NO: 6-12, or a variant thereof.

49. The system according to any one of claims 46-47, wherein the engineered guiding nucleic acid comprises a sequence having at least 80% sequence identity with SEQ ID NO: 6, or a variant thereof, and further comprises at least one mutation relative to SEQ ID NO: 6 according to Table 2.

50. The system according to any one of claims 46-49, wherein the anti-target sequence comprises at least 16-22 nucleotides.

51. The system according to any one of claims 46-50, wherein the polypeptide comprises a sequence having at least 80% sequence identity with SEQ ID NO: 14, and further comprises a polypeptide having at least 80% sequence identity with SEQ ID NO:

15.

52. A method for binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, the method comprising contacting the cell with a system according to any one of claims 18-24, 30-38 or 46-51.

53. The method of claim 52, wherein the DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to the target nucleic acid sequence of 5'-TCN-3'.

54. The method of claim 52 or 53, wherein the contact comprises transfecting the cells with the system.

55. The method of claim 52 or 53, wherein the contact comprises transducing the cell with the system virus.

56. The method according to claim 52 or 53, wherein the contact comprises transducing the cell with one of (a) or (b) viruses and transfecting the cell with the other of (a) or (b).

57. The method according to any one of claims 52-56, wherein (a) further comprises an effector domain containing a transcriptional repression domain, and the method further comprises inhibiting the transcription of the target nucleic acid sequence.

58. The method according to any one of claims 52-56, wherein (a) further comprises an effector domain containing a transcription activation domain, and the method further comprises activating transcription of the target nucleic acid sequence.

59. A cell comprising the system according to any one of claims 19-24, 30-38 or 46-51, the fusion protein according to any one of claims 1-17, or the polynucleotide according to any one of claims 18, 25-29 or 45.

60. The cell of claim 59, wherein the cell is a eukaryotic cell.

61. A vector comprising a polynucleotide according to any one of claims 18, 25-29 or 45.

62. The vector according to claim 61, wherein the vector is a plasmid or a viral vector.

63. The vector according to claim 61, wherein the vector is a viral vector, and wherein the viral vector is an adeno-associated virus (AAV) vector.

64. A fusion protein, said fusion protein comprising: (a) Nuclease-deficient type 2 V Cas polypeptide, the polypeptide (i) Contains a sequence having at least 80% sequence identity with any one of SEQ ID NO: 1, 35, 36, or a variant thereof; or (ii) Contains a missing TSL or C-terminal domain relative to SEQ ID NO: 1, or any combination thereof; and (b) an effector domain comprising a sequence having at least 80% sequence identity with any one of SEQ ID NO: 40-45 or a variant thereof, or consisting of a sequence having at least 80% sequence identity with any one of SEQ ID NO: 40-45 or a variant thereof.

65. The fusion protein of claim 64 further comprises a nuclear localization signal (NLS).

66. The fusion protein of claim 64, wherein the NLS is located between (a) and (b).

67. The fusion protein of claim 64, wherein the NLS is located near the N-terminus or C-terminus of the fusion protein.

68. The fusion protein according to any one of claims 64-67, wherein (a) is located near the N-terminus of the fusion protein, and (b) is located near the C-terminus of the fusion protein.

69. The fusion protein according to any one of claims 64-67, wherein (a) is adjacent to the C-terminus of the fusion protein, and (b) is adjacent to the N-terminus of the fusion protein.

70. The fusion protein according to any one of claims 64-67, wherein (a) and (b) are arranged sequentially from the N-terminus to the C-terminus.

71. The fusion protein according to any one of claims 64-67, wherein (b) and (a) are arranged sequentially from the N-terminus to the C-terminus.

72. The fusion protein according to any one of claims 64-71, wherein the nuclease-deficient type 2 V Cas polypeptide comprises residue 925 or 929 of SEQ ID NO: 1 being substituted with non-wild-type residues.

73. The fusion protein of claim 72, wherein the non-wild-type residue comprises alanine, glycine, lysine, or arginine.

74. The fusion protein of claim 73, wherein the nuclease-deficient type 2 V Cas polypeptide comprises an A925K or I929K mutation relative to SEQ ID NO: 1, or any combination thereof.

75. The fusion protein of claim 74, wherein the nuclease-deficient type 2 V Cas polypeptide comprises the A925K mutation and the I929K mutation relative to SEQ ID NO:

1.

76. The fusion protein according to any one of claims 64-75, wherein the nuclease-deficient type 2 V Cas polypeptide comprises at least one residue corresponding to residues 659, 756, 922 or any combination thereof of SEQ ID NO: 1 being substituted with a non-wild-type residue.

77. The fusion protein of claim 76, wherein the non-wild-type residue is a glycine or alanine residue.

78. The fusion protein of claim 77, wherein the nuclease-deficient class 2 V Cas polypeptide comprises at least one mutation corresponding to the D659A, D756A or D922A mutation of SEQ ID NO: 1, its variants or any combination thereof.

79. The fusion protein according to any one of claims 64-78, wherein the nuclease-deficient class 2 V Cas polypeptide comprises the sequence having at least 80% sequence identity with any one of SEQ ID NO: 35, 36, or a variant thereof.

80. The fusion protein according to any one of claims 64-78, wherein the nuclease-deficient type 2 V Cas polypeptide comprises the TSL or C-terminal domain deletion, or any combination thereof.

81. The fusion protein according to any one of claims 64-78, wherein the effector domain comprises a sequence having at least 80% sequence identity with SEQ ID NO: 45, and comprises a mutation of at least one of the residues in SEQ ID NO: 46 that are mutated relative to SEQ ID NO:

45.

82. A polynucleotide encoding a fusion protein according to any one of claims 64-81.

83. A system comprising: (a) A fusion protein according to any one of claims 1-17 or 64-81, or a polynucleotide encoding a fusion protein according to any one of claims 1-17 or 64-81; (b) an engineered guiding nucleic acid or a template polynucleotide encoding the engineered guiding nucleic acid, wherein the engineered guiding nucleic acid is configured to form a complex with the nuclease-deficient class 2 type V Cas polypeptide, the engineered guiding nucleic acid comprising: (i) an anti-target nucleic acid sequence, wherein the anti-target nucleic acid sequence is configured to hybridize with a target nucleic acid sequence; (ii) A scaffold nucleic acid sequence configured to bind to the nuclease-deficient type 2 V Cas polypeptide.

84. The system of claim 83, wherein the engineered guiding nucleic acid comprises (ii) and (i) sequentially from 5' to 3'.

85. The system according to claim 83 or 84, wherein the engineered guiding nucleic acid comprises a variant of a sequence having at least 80% sequence identity with any one of SEQ ID NO: 6-12 or 47-52.

86. The system according to any one of claims 83-85, wherein the type 2 V Cas polypeptide sequence is a type 2 VE polypeptide sequence.

87. The system according to any one of claims 83-86, wherein the anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

88. A polypeptide comprising a sequence having at least 80% sequence identity with any one of SEQ ID NO: 1, 35, 36, or a variant thereof, or comprising a deletion of the TSL or C-terminal domain relative to SEQ ID NO: 1, or any combination thereof.

89. The polypeptide of claim 88, wherein the nuclease-deficient type 2 V Cas polypeptide comprises residue 925 or 929 of SEQ ID NO: 1 being substituted with non-wild-type residues.

90. The polypeptide according to claim 88 or 89, wherein the non-wild-type residue comprises alanine, glycine, lysine, or arginine.

91. The polypeptide of claim 90, wherein the nuclease-deficient type 2 V Cas polypeptide comprises an A925K or I929K mutation relative to SEQ ID NO: 1, or any combination thereof.

92. The polypeptide of claim 91, wherein the nuclease-deficient type 2 V Cas polypeptide comprises the A925K mutation and the I929K mutation relative to SEQ ID NO:

1.

93. A polynucleotide comprising a variant thereof having at least 80% sequence identity with any one of SEQ ID NO: 47-52.

94. A polynucleotide comprising or encoding a sequence having at least 80% sequence identity with any one of SEQ ID NO: 47-52, or a variant thereof.

95. A system comprising: (a) A polypeptide containing two types of V-type Cas polypeptide sequences, or a polynucleotide encoding said two types of V-type Cas polypeptide sequences; (b) The polynucleotide according to any one of claims 25-28 or 94, or the template polynucleotide encoding the polynucleotide according to any one of claims 25-28 or 94.

96. The system of claim 95, wherein the type 2 V Cas polypeptide comprises the polypeptide of any one of claims 88-92.

97. The system according to any one of claims 95-96, wherein the type 2 V Cas polypeptide sequence is a type 2 VE polypeptide sequence.

98. The system according to any one of claims 95-97, wherein the anti-target nucleic acid sequence comprises at least 16-22 nucleotides.

99. A method for binding a target nucleic acid sequence in a deoxyribonucleic acid (DNA) molecule in a cell, the method comprising contacting the cell with a system according to any one of claims 83-87 or 95-98.

100. The method of claim 99, wherein the DNA molecule comprises a protospacer adjacent motif (PAM) sequence according to the 5'-TCN-3'-target nucleic acid sequence.

101. The method of claim 99 or 100, wherein the contact comprises transfecting the cells with the system.

102. The method according to any one of claims 99-101, wherein the contact comprises transducing the cells with the system virus.

103. The method according to any one of claims 99-102, wherein the contact comprises transducing the cell with one of (a) or (b) viruses and transfecting the cell with the other of (a) or (b).

104. The method according to any one of claims 99-103, wherein (a) further comprises an effector domain containing a transcriptional repression domain, and the method further comprises inhibiting the transcription of the target nucleic acid sequence.

105. The method according to any one of claims 99-103, wherein (a) further comprises an effector domain containing a transcription activation domain, and the method further comprises activating transcription of the target nucleic acid sequence.

106. A cell comprising the system according to any one of claims 83-87 or 95-98.

107. The cell of claim 106, wherein the cell is a eukaryotic cell.

108. A vector comprising the polynucleotide according to claim 94.

109. The vector according to claim 108, wherein the vector is a plasmid or a viral vector.

110. The vector of claim 109, wherein the vector is a viral vector, and wherein the viral vector is an adeno-associated virus (AAV) vector.

Citation Information

Cited By

  • Cas protein, fusion protein, corresponding gene editing system and application thereof

    CN121780486A