Engineered Cas12b effector proteins and methods of use thereof

JP2025501473A5Pending Publication Date: 2025-12-17BEIJING INST FOR STEM CELL & REGENERATIVE MEDICINE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024534426
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-09
Filing Date
2022-12-09
Publication Date
2025-12-17

AI Technical Summary

Technical Problem

Traditional CRISPR-Cas systems have limitations such as limited gene editing efficiency, which hinders their effectiveness in genome editing across multiple loci.

Method used

Engineered Cas12b nucleases with specific mutations, including substitutions of amino acid residues to enhance interactions with PAM, DNA duplex, and single-stranded DNA substrates, along with engineered gRNAs, improve the enzymatic activity and specificity of genome editing.

Benefits of technology

The engineered Cas12b nucleases and gRNAs enhance gene editing efficiency, enabling targeted and efficient modification of nucleic acids in various cell types, including bacterial, yeast, plant, and animal cells, with applications in treating diseases like cancer and viral infections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000125_0000
    Figure 00000125_0000
  • Figure 00000125_0001
    Figure 00000125_0001
  • Figure 00000125_0002
    Figure 00000125_0002
Patent Text Reader

Abstract

The present application provides engineered Cas12b nucleases or derivatives thereof that contain one or more types of mutations to improve activity (e.g., gene editing activity) or eliminate nuclease activity, as well as engineered Cas12b effector proteins, engineered gRNAs (e.g., sgRNAs or tracrRNAs), engineered CRISPR-Cas12b systems, and methods of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [CROSS REFERENCE TO RELATED APPLICATIONS] This application claims priority to International Patent Application No. PCT / CN2021 / 136761, filed December 9, 2021, the contents of which are incorporated by reference in their entirety herein.

[0002] About the Electronic Sequence Listing The contents of the electronic sequence listing (253112000541SEQLIST.xml; size: 111,583 bytes; and creation date: November 22, 2022) are incorporated herein by reference in their entirety.

[0003] The present application generally relates to the field of biotechnology. More specifically, the present application relates to engineered Cas12b effector proteins and engineered gRNA scaffolds with improved activity (e.g., gene editing activity) or eliminated nuclease activity, and methods of using them. [Background technology]

[0004] Genome editing is an important and useful technology in genome research and various applications. Various systems are available for genome editing, including clustered regularly interspaced short palindromic repeats (CRISPR)-Cas systems, transcription activator-like effector nucleases (TALEN) systems, and zinc finger nuclease (ZFN) systems.

[0005] The CRISPR-Cas system is an efficient and economical genome editing technology that has been widely used in various eukaryotic organisms, ranging from yeast and plants to zebrafish and humans (reviewed in Van der Oost 2013, Science 339: 768-770, and Charpentier and Doudna, 2013, Nature 495: 50-51). The CRISPR-Cas system provides adaptive immunity to archaea and bacteria by combining Cas effector proteins with CRISPR RNA (crRNA). To date, two classes (class 1 and class 2) of CRISPR-Cas systems, including six types (types I-VI), have been characterized based on the superior functionality and evolved modularity of the system. Among the class 2 CRISPR-Cas systems, the type II Cas9 system and the VA / B / E / J type Cas12a / Cas12b / Cas12e / Cas12f / Cas12j systems have been used for genome editing and offer broad prospects for biomedical research. Summary of the Invention

[0006] Conventional CRISPR-Cas systems have various limitations, such as limited efficiency of gene editing. The present application provides an improved method and system for efficient genome editing across multiple loci. Specifically, the present application provides an engineered Cas12b nuclease with improved enzymatic activity, an engineered Cas12b effector protein, an engineered gRNA (e.g., sgRNA or tracrRNA) comprising an engineered scaffold, and a method of using the engineered Cas12b effector protein and / or the engineered gRNA for gene editing. In one aspect, the present application provides an engineered Cas12b nuclease comprising one, two or three types of mutations relative to a reference Cas12b nuclease, the mutations comprising: (1) substitution of one or more amino acid residues that interact with a protospacer adjacent motif (PAM) in the reference Cas12b nuclease with positively charged amino acid residues (e.g., R, H, K); and / or (2) substitution of one or more amino acid residues that are involved in DNA duplex (dsDNA) cleavage in the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W); and / or (3) substitution of one or more amino acid residues that interact with a single-stranded DNA (ssDNA) substrate in the RuvC domain of the Cas12b nuclease with positively charged amino acid residues (e.g., R, H, K) or hydrophobic amino acid residues (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease is a wild-type Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a Cas12b nuclease from Alicyclobacillus acidophilus (AaCas12b). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1.

[0007] In some embodiments of the engineered Cas12b nuclease described above, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues that interact with the PAM in a reference Cas12b nuclease with a positively charged amino acid residue. In some embodiments, the one or more amino acid residues that interact with the PAM are within 10 Å (e.g., 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) of the PAM in the three-dimensional structure. The one or more amino acid residues that interact with the PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475. In some embodiments, the one or more amino acid residues that interact with the PAM include one or more of the following amino acid residues: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and E475. In some embodiments, the one or more amino acid residues that interact with the PAM include one or more of the following amino acid residues: D116 and E475. In some embodiments, the positively charged amino acid residue is R or K. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D116R and E475R. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:2 or 3.

[0008] In some embodiments of the engineered Cas12b nuclease described above, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues involved in DNA duplex cleavage in a reference Cas12b nuclease with an amino acid residue having an aromatic ring. In some embodiments, the one or more amino acid residues involved in DNA duplex cleavage interact with the last base pair in the PAM to the 3' end of the target strand. In some embodiments, the one or more amino acid residues involved in DNA duplex cleavage are located at one or more of positions 118 and 119. In some embodiments, the amino acid residue having an aromatic ring is Y, F, or W. In some embodiments, the substitution of one or more amino acid residues involved in DNA duplex cleavage in a reference Cas12b nuclease with an amino acid residue having an aromatic ring is Q119Y, Q119F, or Q119W. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of any of SEQ ID NOs:4-6.

[0009] In some embodiments of the engineered Cas12b nuclease described above, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues located in the RuvC domain in a reference Cas12b nuclease that interacts with a single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue. In some embodiments, the one or more amino acid residues located in the RuvC domain that interacts with a single-stranded DNA substrate are within 10 Å (e.g., 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) of the single-stranded DNA substrate in the three-dimensional structure. In some embodiments, the one or more amino acid residues located within the RuvC domain and that interact with the single-stranded DNA substrate are located at one or more of positions 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097. In some embodiments, the one or more amino acid residues located within the RuvC domain and that interact with the single-stranded DNA substrate include one or more of the following amino acid residues: D300, K301, E304, N329, E636, Q639, T647, Q682, I757, E758, E761, E764, K768, E852, Q854, N856, N857, D858, P860, S862, E863, N865, Q866, L867, Q869, E938, E956, G957, E958, I994, Q1093, and W1097. In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more of the following amino acid residues with a positively charged amino acid residue: E636, Q639, T647, Q682, I757, E758, E761, K768, Q854, N857, D858, N865, Q866, I994, Q1093, and W1097. In some embodiments, the positively charged amino acid residue is R or K.In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R, and W1097R. In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more of the following amino acid residues: E758, E761, E863, N865, Q866, Q869, Q956, and Q1093 with a hydrophobic amino acid residue. In some embodiments, the hydrophobic amino acid residue is W, Y, F, or M, e.g., W, Y, or M. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W, and Q1093Y. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of any of SEQ ID NOs: 7-19.

[0010] In some embodiments of the engineered Cas12b nuclease described in any of the above, the engineered Cas12b nuclease is selected from the group consisting of (1) D116R, (2) E475R, (3) Q119F+E475R, (4) Q119F+E475R+E758R, (5) Q119Y, (6) Q119F, (7) Q119W, (8) I757R, (9) E758R, (10) E761R, (11) K768R, (12) I757R+E758R, (13) I757R+E761R, (14) I757R+K768R, (15) E758R+E761R, (16) E758R+K768R, (17) E761R+K768R, (18) I757R+E758R+E761R, (19) I757R+E758R+K768R, (20) I757R+E761R+K768R, (21) E758R+E761R+K768R, (22) I757R+E758R+E761R+K768R, (23) Q866M, (24) Q869M, (25) Q866M+Q869M, (26) E636R, (27) Q854R, (28) N857K, (29) N865W, (30) N865Y, (31) Q1093W, (32) Q1093Y, and (33) In some embodiments, the engineered Cas12b nuclease comprises any of the following substitutions: (1) Q866M+Q869M, (2) Q119F+E475R, and (3) Q119F+E475R+E758R, or a combination thereof, where the amino acid residues are numbered according to SEQ ID NO: 1. ... amino acid sequences of SEQ ID NOs: 20-22.

[0011] In some embodiments of the engineered Cas12b nuclease described above, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% (e.g., at least about any of 88%, 90%, 95%, 96%, 97%, 98%, 99% or more) sequence identity to any of SEQ ID NOs: 1-22. In some embodiments, the engineered Cas12b nuclease comprises (or consists of, or consists essentially of) the amino acid sequence of any of SEQ ID NOs: 2-22.

[0012] In some embodiments of the engineered Cas12b nuclease described above, the engineered Cas12b nuclease further comprises one or more mutations of a reference Cas12b nuclease that increase flexibility of a flexible region comprising amino acid residues 855-859. In some embodiments, the one or more mutations that increase flexibility comprise N856G. In some embodiments, the amino acid position numbering refers to SEQ ID NO:1.

[0013] In one embodiment of the present application, (1) D116R, (2) E475R, (3) Q119F+E475R, (4) Q119F+E475R+E758R, (5) Q119Y, (6) Q119F, (7) Q119W, (8) I757R, (9) E758R, (10) E761R, (11) K768R, (12) I757R+E758R, (13) I757R+E761R, (14) I757R+K768R, (15) E758R+E761R, (16) E758R+K768R, (17) E761R+K768R, (18) I757R+E758R+E761R , (19) I757R+E758R+K768R, (20) I757R+E761R+K768R, (21) E758R+E761R+K768R, (22) I757R+E758R+E761R+K768R, (23) Q866M, (24) Q869M, (25) Q866M+Q869M, (26) E636R, (27) Q854R, (28) N857K, (29) N865W, (30) N865Y, (31) Q1093W, (32) Q1093Y, and (33) D858R mutations, and the amino acid position numbers are SEQ ID NO: 1. The present invention provides an engineered Cas12b nuclease, which refers to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises any one or combination of the following substitutions: (1) Q866M+Q869M, (2) Q119F+E475R, and (3) Q119F+E475R+E758R, wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises a substitution that is Q119F+E475R+E758R, wherein the amino acid residues are numbered according to SEQ ID NO: 1.

[0014] One aspect of the present application provides an engineered Cas12b nuclease having at least about 85% (e.g., at least about any of 88%, 90%, 95%, 96%, 97%, 98%, 99% or more) sequence identity to any of SEQ ID NOs:2-22, or comprising (or consisting of, or consisting essentially of) the amino acid sequence of any of SEQ ID NOs:2-22.

[0015] An aspect of the present application provides an engineered Cas12b effector protein comprising any of the engineered Cas12b nucleases, variants, or functional derivatives thereof described above. In some embodiments, the engineered Cas12b nuclease or functional derivatives thereof has enzymatic activity. In some embodiments, the engineered Cas12b effector protein may induce a double-stranded break in a DNA molecule. In some embodiments, the engineered Cas12b effector protein may induce a single-stranded break in a DNA molecule. In some embodiments, the engineered Cas12b effector protein comprises an enzymatically inactive mutant of the engineered Cas12b nuclease. In some embodiments, the enzymatically inactive mutant of the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues selected from the group consisting of D570A, E848A, R785A, E848A, R911A, and D977A, wherein the amino acid residues are numbered according to SEQ ID NO:1. In some embodiments, an enzymatically inactive mutant of an engineered Cas12b nuclease comprises (or consists of, or consists essentially of) the amino acid sequence of any of SEQ ID NOs:79-81, or a variant having at least about 85% (e.g., at least about any of 88%, 90%, 95%, 96%, 97%, 98%, 99% or more) sequence identity to any of SEQ ID NOs:79-81.

[0016] In some embodiments of the engineered Cas12b effector protein described above, the engineered Cas12b effector protein further comprises a functional domain fused to the engineered Cas12b nuclease or a functional derivative thereof. In some embodiments, the functional domain is selected from the group consisting of a translation initiation domain, a transcriptional repressor domain, a transactivation domain, an epigenetic modification domain, a nucleobase editing domain, a reverse transcriptase domain, a reporter domain, and a nuclease domain. In some embodiments, the transcriptional repression domain is a Kruppel-associated box (KRAB) domain, such as comprising the amino acid sequence of SEQ ID NO:72.

[0017] In some embodiments of the engineered Cas12b effector protein described above, the engineered Cas12b effector protein may comprise a first polypeptide comprising an N-terminal portion of an engineered Cas nuclease or a functional derivative thereof, and a second polypeptide comprising a C-terminal portion of an engineered Cas nuclease or a functional derivative thereof, wherein the first and second polypeptides may associate with each other in the presence of a guide RNA comprising a guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence. In some embodiments, the engineered Cas12b effector protein may comprise a first polypeptide comprising amino acid residues 1 to X at the N-terminus of an engineered Cas12b nuclease or a functional derivative thereof, and a second polypeptide comprising the X+1 residue to the C-terminus of an engineered Cas12b nuclease or a functional derivative thereof, wherein the first and second polypeptides may associate with each other in the presence of a guide RNA comprising a guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence. In some embodiments, the first and second polypeptides each comprise a dimerization domain. In some embodiments, the first and second dimerization domains associate with each other in the presence of an inducing agent. In some embodiments, the first and second polypeptides do not comprise a dimerization domain.

[0018] Another aspect of the present application provides a single guide RNA (sgRNA) comprising any of the sequences of SEQ ID NOs: 25 to 53.

[0019] Another aspect of the present application provides an engineered CRISPR-Cas12b system comprising: (a) an engineered Cas12b nuclease according to any of the above-described engineered Cas12b nucleases, an engineered Cas12b effector protein according to any of the above-described engineered Cas12b effector proteins, or nucleic acids encoding the same; and (b) a guide RNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the guide RNA, wherein the engineered Cas12b nuclease or engineered Cas12b effector protein and the guide RNA form a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence, and induces modification of the target nucleic acid. In some embodiments, the guide RNA comprises a crRNA and a tracrRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises an array of precursor guide RNAs encoding a plurality of crRNAs. In some embodiments, the guide RNA is a single guide RNA (sgRNA). In some embodiments, the sgRNA comprises any of the sequences of SEQ ID NOs: 23-53. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding an engineered Cas12b nuclease or an engineered Cas12b effector protein. In some embodiments, the one or more vectors are adeno-associated virus (AAV) vectors. In some embodiments, the AAV vector further encodes a guide RNA.

[0020] Another aspect of the present application provides an engineered CRISPR-Cas12b system, comprising: (a) an engineered Cas12b nuclease according to any of the above-described engineered Cas12b nucleases, an engineered Cas12b effector protein according to any of the above-described engineered Cas12b effector proteins, a Cas12b nuclease or its effector protein comprising any of the amino acid sequences of SEQ ID NOs: 1-22 and 79-81, or a nucleic acid encoding the same; and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding a gRNA, wherein the gRNA comprises an engineered scaffold comprising any of the sequences of SEQ ID NOs: 25-53, wherein the Cas12b nuclease (e.g., engineered) or its effector protein and the gRNA form a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid. In some embodiments, the gRNA comprises a crRNA and a tracrRNA, and the tracrRNA comprises an engineered scaffold or a portion thereof. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor gRNA array encoding multiple crRNAs. In some embodiments, the gRNA is an sgRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding an engineered Cas12b nuclease or its effector proteins, or a Cas12b nuclease or its effector proteins. In some embodiments, the one or more vectors are AAV vectors. In some embodiments, the one or more vectors further encode a gRNA.

[0021] One aspect of the present application provides a method for detecting a target nucleic acid in a sample, comprising: (a) contacting the sample with any of the engineered CRISPR-Cas12b system described above and a labeled detection nucleic acid, wherein the gRNA comprises a guide sequence complementary to a target sequence of the target nucleic acid, and the labeled detection nucleic acid is single-stranded and does not hybridize to the guide sequence of the guide gRNA; and (b) detecting the target nucleic acid by measuring a detectable signal resulting from cleavage of the labeled detection nucleic acid by an engineered Cas12b nuclease or an effector protein thereof.

[0022] An aspect of the present application provides a method of modifying a target nucleic acid comprising a target sequence, the method comprising contacting the target nucleic acid with any of the engineered CRISPR-Cas12b systems described above. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is present in a cell. In some embodiments, the cell is a bacterial cell, a yeast cell, a plant cell, or an animal cell (e.g., a mammalian cell). In some embodiments, the method is performed ex vivo. In some embodiments, the method is performed in vivo. In some embodiments, the target nucleic acid is degraded. In some embodiments, a target sequence in the target nucleic acid is altered by the engineered CRISPR-Cas12b system. In some embodiments, expression of the target nucleic acid is altered by the engineered CRISPR-Cas12b system. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target sequence is associated with a disease or condition. In some embodiments, the engineered CRISPR-Cas12b system comprises an array of precursor guide RNAs encoding multiple crRNAs, each crRNA comprising a different guide sequence.

[0023] Another aspect of the present application provides a method of treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising modifying a target nucleic acid in a cell of the individual using an engineered CRISPR-Cas12b system according to any of the above-described engineered CRISPR-Cas12b systems to treat the disease or condition. In some embodiments, the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.

[0024] Also provided is an engineered cell comprising a modified target nucleic acid, the target nucleic acid being modified by any of the methods described above, and also an engineered non-human animal comprising one or more of the engineered cells.

[0025] Also provided are compositions, kits and articles of manufacture for use in any of the above methods.

[0026] It should be understood that, for clarity, certain features of the present disclosure described in the context of individual embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, multiple features of the present disclosure described in the context of a single embodiment may also be provided alone or in any suitable subcombination. The present disclosure specifically encompasses all combinations of embodiments relating to specific method steps, reagents or conditions, or composition components, and are disclosed herein as if each combination was individually and explicitly disclosed. [Brief description of the drawings]

[0027] [Figure 1]Figure 1 shows the gene editing efficiency (% insertion / deletion) of exemplary AaCas12b variants in which amino acid residues that interact with the PAM in wild-type AaCas12b were replaced with R. AaCas12b variants with D116R or E475R substitutions showed improved editing efficiency compared to wild-type (WT) AaCas12b. [Diagram 2] 1 shows the gene editing efficiency of exemplary AaCas12b variants in which amino acid residues involved in DNA double strand cleavage in wild-type AaCas12b are replaced with aromatic amino acid residues. AaCas12b variants with Q119Y, Q119F, or Q119W substitutions show improved gene editing efficiency compared to wild-type AaCas12b. [Diagram 3] 1 shows the gene editing efficiency of an exemplary AaCas12b variant in which an amino acid residue located within the RuvC domain in wild-type AaCas12b that interacts with single-stranded DNA substrates is replaced with R. [Figure 4A] Figure 4 shows the gene editing efficiency of exemplary AaCas12b variants in which the amino acid residues located in the RuvC domain in wild-type AaCas12b and interacting with single-stranded DNA are replaced with lysine (K) or arginine (R) residues. Figure 4A shows the editing efficiency at genomic site CCR5-3, and Figure 4B shows the editing efficiency at genomic site RNF2-5. AaCas12b variants with E636R, I757R, E758R, E761R, Q854R, D858R, E758K, I994R, N857K, or D858K substitutions show the highest improvement in gene editing efficiency compared to WT AaCas12b. [Figure 4B]Figure 4 shows the gene editing efficiency of exemplary AaCas12b variants in which the amino acid residues located in the RuvC domain in wild-type AaCas12b and interacting with single-stranded DNA are replaced with lysine (K) or arginine (R) residues. Figure 4A shows the editing efficiency at genomic site CCR5-3, and Figure 4B shows the editing efficiency at genomic site RNF2-5. AaCas12b variants with E636R, I757R, E758R, E761R, Q854R, D858R, E758K, I994R, N857K, or D858K substitutions show the highest improvement in gene editing efficiency compared to WT AaCas12b. [Diagram 5] 1 shows the gene editing efficiency of exemplary AaCas12b variants in which amino acid residues located within the RuvC domain in wild-type AaCas12b that interact with single-stranded DNA substrates are substituted with hydrophobic amino acid residues W, Y, F, or M. AaCas12b variants with N865W, N865Y, Q866M, Q869M, Q1093W, or Q1093Y substitutions show the highest improvement in gene editing efficiency compared to WT AaCas12b. [Figure 6] 1 shows the gene editing efficiency of exemplary AaCas12b variants with combined mutations compared to WT AaCas12b. [Figure 7] The AaCas12b variant Q119F+E475R+E758R shows significantly improved gene editing efficiency compared to WT AaCas12b and the corresponding single mutants. [Figure 8]Alicyclobacillus acidiphilus Cas12b (AaCas12b) (SEQ ID NO: 1), Alicyclobacillus kakegawensis Cas12b (AkCas12b) (SEQ ID NO: 54), Alicyclobacillus macrosporangiidus Cas12b (AmCas12b) (SEQ ID NO: 55), Bacillus sp. V3-13 Cas12b (Bs3Cas12b) (SEQ ID NO: 56), Bacillus Cas12b (BsCas12b) (SEQ ID NO: 57), Laceyella sediminis 1 shows an amino acid sequence alignment of Cas12b proteins, including Cas12b (LsCas12b) (SEQ ID NO: 58), Bacillus hisashii Cas12b (BhCas12b) (SEQ ID NO: 59), and Spirochaetes bacterium Cas12b (SbCas12b) (SEQ ID NO: 60). Substitutions based on AaCas12b described herein can be made at the corresponding amino acid positions in any of the Cas12b orthologs described herein. [Figure 9] We show that sgRNAs with engineered scaffolds significantly improved the gene editing efficiency of AaCas12b variant Q119F+E475R+E758R. As a control, sgRNAs with AaCas12b Aa-sg scaffold or AacCas12b sgRNA scaffold (V0) are used. [Figure 10A] FIG. 1 is a schematic diagram of an exemplary construct encoding AaCas12b variants Q119F+E475R+E758R+D570A under the control of a CMV promoter and an sgRNA under the control of a U6 promoter. [Figure 10B]Figure 1 shows T7EI assay results as a measure of nuclease activity of AaCas12b(Q119F+E475R+E758R) and AaCas12b(Q119F+E475R+E758R+D570A). sgRNA1 and sgRNA2 specifically recognize the target site in HBG1 / 2. As a negative control, a control sgRNA that does not target the HBG1 / 2 sequence is used. [Figure 11A] FIG. 1 is a schematic diagram of an exemplary construct encoding AaCas12b variants Q119F+E475R+E758R+D570A+E848A or Q119F+E475R+E758R+D570A+D977A under the control of a CMV promoter and an sgRNA under the control of a U6 promoter. [Figure 11B] Figure 1 shows T7EI measurements as a measure of nuclease activity of AaCas12b (Q119F+E475R+E758R), AaCas12b (Q119F+E475R+E758R+D570A+E848A) and AaCas12b (Q119F+E475R+E758R+D570A+D977A) mediated by sgRNA1 and sgRNA2, which specifically recognize the target site in HBG1 / 2. As a negative control, a control sgRNA that does not target the HBG1 / 2 sequence is used. [Figure 12A] FIG. 1 is a schematic diagram of an exemplary construct encoding AaCas12b (Q119F+E475R+E758R+D570A+D977A) fused to KRAB under the control of a CMV promoter and an sgRNA under the control of a U6 promoter. [Figure 12B] Figure 2 shows the relative mRNA levels of mouse Nav1.7 in mouse N2a cells transfected with AaCas12b(Q119F+E475R+E758R+D570A+D977A)-KRAB fusion protein mediated by different sgRNAs targeting different sites of SCN9A gene. As a control, no sgRNA transfection was performed. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] The present application provides an engineered Cas12b nuclease that has improved enzymatic activity (e.g., gene editing activity) by introducing one, two, or three types of mutations relative to a reference Cas12b nuclease. Also provided is an engineered Cas12b nuclease or its effector protein (e.g., dCas12b) that has reduced or eliminated nuclease activity. Also provided is an engineered guide RNA (gRNA) with an engineered scaffold sequence that can improve Cas12b enzymatic activity (e.g., gene editing activity) when used with a (wild-type or engineered) Cas12b nuclease. Also provided is an engineered Cas12b effector protein, a method of using the engineered Cas12b nuclease or the engineered Cas12b effector protein, and / or an engineered gRNA.

[0029] I. Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0030] As used herein, the term "Cas12b protein" is used in the broadest sense and includes a parent or reference Cas12b protein (e.g., including the AaCas12b protein of SEQ ID NO:1), a derivative or variant thereof (e.g., an engineered Cas12b, dCas12b, or an engineered Cas12b effector protein), and a functional fragment thereof, such as an oligonucleotide-binding fragment.

[0031] As used herein, "effector protein" means a protein having an activity such as site-specific binding activity, single-stranded DNA cleavage or editing activity, double-stranded DNA cleavage or editing activity, single-stranded RNA cleavage or editing activity, or transcriptional regulation activity.

[0032] As used herein, "guide RNA" and "gRNA" can be used interchangeably and refer to an RNA that can form a complex with Cas12b nuclease or effector protein and a target nucleic acid (e.g., double-stranded DNA). A guide RNA can include a single RNA molecule or can include two or more RNA molecules associated with each other by hybridization of complementary regions in two or more RNA molecules. When used in combination with a dual RNA-guided Cas nuclease (e.g., Cas12b), a guide RNA includes a crRNA and a tracrRNA, or a single guide RNA (sgRNA). A "crRNA" or "CRISPR RNA" includes a guide sequence that has sufficient complementarity with a target sequence of a target nucleic acid (e.g., double-stranded DNA) to guide the CRISPR complex to specifically bind to a sequence of the target nucleic acid. The "tracrRNA" or "trans-activating CRISPR RNA" can be partially complementary to the crRNA, base-pair, and play a role in the maturation process of the crRNA. A "single guide RNA" or "sgRNA" is an engineered guide RNA that is a single-molecule fusion of crRNA and tracrRNA.

[0033] As used herein, the term "CRISPR array" refers to a nucleic acid (e.g., DNA) fragment that includes a CRISPR repeat sequence and a spacer region, beginning with the first nucleotide of the first CRISPR repeat sequence and ending with the last nucleotide of the last (terminal) CRISPR repeat sequence. Generally, each spacer region in a CRISPR array is located between two repeat sequences. As used herein, "CRISPR repeat sequence" or "CRISPR direct repeat sequence" or "direct repeat sequence" refers to multiple short direct repeat sequences with little or no sequence variation in a CRISPR array. Advantageously, the VI direct repeat sequence may form a stem-loop structure.

[0034] As used herein, "donor template nucleic acid" or "donor template" can be used interchangeably and refer to a nucleic acid molecule that can change the structure of a target nucleic acid by one or more cellular proteins after a CRISPR enzyme described herein changes the target nucleic acid. In some examples, the donor template nucleic acid is a double-stranded nucleic acid. In some examples, the donor template nucleic acid is a single-stranded nucleic acid. In some examples, the donor template nucleic acid is linear. In some examples, the donor template nucleic acid is circular (e.g., a plasmid). In some examples, the donor template nucleic acid is an exogenous nucleic acid molecule. In some examples, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome).

[0035] The terms "nucleic acid," "polynucleotide," and "nucleotide sequence" can be used interchangeably and refer to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and the like. "Oligonucleotide" and "oligo" can be used interchangeably and refer to short polynucleotides up to about 50 nucleotides in length.

[0036] As used herein, "complementarity" refers to the ability of one nucleic acid to form hydrogen bonds with another nucleic acid by conventional Watson-Crick base pairing. Complementarity percentage refers to the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (i.e., Watson-Crick base pairing) with a second nucleic acid (e.g., about 5, 6, 7, 8, 9, 10 out of 10, about 50%, 60%, 70%, 80%, 90% and 100% complementary, respectively). "Fully complementary" means that all contiguous residues of one nucleic acid sequence form hydrogen bonds with the same number of contiguous residues of another nucleic acid sequence. As used herein, "substantially complementary" means that the degree of complementarity is at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of about 40, 50, 60, 70, 80, 100, 150, 200, 250 or more nucleotides, or that the two nucleic acids hybridize under stringent conditions.

[0037] As used herein, "stringency conditions" for hybridization refer to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily to the target sequence and to a lesser extent to non-target sequences. Stringency conditions are generally sequence-dependent and vary depending on a variety of factors. In general, the longer the sequence, the higher the temperature at which it will specifically hybridize to a target sequence. Non-limiting examples of stringency conditions are described in detail in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology-Hybridization With Nucleic Acid Probes Part I, Second Chapter "Overview of principles of hybridization and the strategy of nucleic acid probe assay," Elsevier, N,Y.

[0038] "Hybridize" refers to a reaction in which one or more polynucleotides react to form a complex stabilized by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonding can occur by Watson-Crick base pairing, Hoogstein binding, or other sequence-specific methods. A sequence that is capable of hybridizing to a given sequence is called the "complement" of the given sequence.

[0039] "Percent sequence identity" with respect to nucleic acid sequences is defined as the percentage of nucleotides in a candidate sequence that are identical to the nucleotides in a particular nucleic acid sequence, after alignment of the sequences, allowing gaps (if necessary) to achieve the maximum percent sequence identity. In the case of peptide, polypeptide or protein sequences, "percent sequence identity" means the percentage of amino acid residues in a candidate sequence that are identically substituted with amino acid residues in a particular peptide or amino acid sequence, after alignment of the sequences, allowing gaps (if necessary) to achieve the maximum percent sequence identity. Alignment to determine percent amino acid sequence identity can be achieved in a variety of ways known to those skilled in the art, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN or MEGALIGN® (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms required to achieve maximum alignment over the entire length of the sequences being compared.

[0040] The terms "polypeptide" and "peptide" can be used interchangeably and refer to amino acid polymers of any length. The polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. A protein may have one or more polypeptides. The term also encompasses modified amino acid polymers, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component.

[0041] As used herein, a "variant" refers to a polynucleotide or polypeptide that differs from a reference polynucleotide or polypeptide, respectively, but retains essential properties. A typical variant of a polynucleotide differs in nucleic acid sequence from another reference polynucleotide. A change in the nucleic acid sequence of a variant may or may not change the amino acid sequence of a polypeptide encoded by a reference polynucleotide. Nucleotide changes can lead to amino acid substitutions, additions, deletions, fusions, and truncations in the polypeptide encoded by the reference sequence, as described below. A typical variant of a polypeptide differs in amino acid sequence from another reference polypeptide. Generally, the differences are limited, so that the sequences of the reference polypeptide and the variant are overall very similar and identical in many regions. The variant and the reference polypeptide may differ in amino acid sequence by one or more substitutions, additions, deletions, and any combination thereof. The substituted or inserted amino acid residues may or may not be encoded by the genetic code. A variant of a polynucleotide or polypeptide may be a naturally occurring one, such as an allelic variant, or a naturally occurring unknown variant. Non-naturally occurring variants of polynucleotides and polypeptides can be made by mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.

[0042] As used herein, the term "wild-type" has the meaning commonly understood by one of ordinary skill in the art and refers to a typical form of an organism, strain or gene, or a characteristic that distinguishes it from a mutant or variant when present in nature, that can be isolated from a natural source and has not been intentionally modified.

[0043] As used herein, the terms "non-naturally occurring" or "engineered" can be used interchangeably and refer to artificial involvement. When these terms are used to describe a nucleic acid molecule or polypeptide, it is meant that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated or associated as found in nature.

[0044] As used herein, the terms "orthologue" or "ortholog" have the meaning commonly understood by one of ordinary skill in the art. Furthermore, an "ortholog" of a protein as referred to herein means a protein that belongs to a different species and has the same or similar function as the protein that is its ortholog.

[0045] As used herein, the term "identity" is used to mean sequence matching between two polypeptides or two nucleic acids. If a position in two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., if the position in each of the two DNA molecules is occupied by adenine, or if the position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions being compared times 100. For example, if 6 positions out of 10 positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 positions out of a total of 6 positions are matched). Generally, such comparisons are performed when the two sequences are aligned to produce maximum identity. Such alignments can be accomplished, for example, by the method of Needleman et al. (1970) J. Mol. Biol. 48: 443-453, which can be readily performed by computer programs such as, for example, the Align program (DNAstar, Inc.). Alternatively, the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4: 11-17 (1988)) integrated into the ALIGN program (version 2.0) may be used, using a PAM 120 weighted residue table. Percent identity between two amino acid sequences is determined using a gap length penalty of 12 and a gap penalty of 4. Additionally, the Needleman and Wunsch (J MoI Biol. 48: 444-453 (1970)) algorithm of the GAP program integrated into the GCG software package (available at www.gcg.com) can be used to determine percent identity between two amino acid sequences using a Blossum 62 matrix or a PAM250 matrix with gap weights of 16, 14, 12, 10, 8, 6, or 4 and length weights of 1, 2, 3, 4, 5, or 6.

[0046] As used herein, "cell" refers not only to a particular individual cell but also to the progeny or potential progeny of that cell. Because certain modifications can occur in successive generations due to mutations or environmental influences, such progeny may not actually be identical to the parent cell, but are still within the scope of the term as used herein.

[0047] As used herein, the terms "transduction" and "transfection" include all methods known in the art for introducing DNA into cells using infectious agents (e.g., viruses) or other means to express a protein or molecule of interest. In addition to viruses or virus-like agents, chemical transfection methods such as those using calcium phosphate, dendrimers, liposomes or cationic polymers (e.g., DEAE-dextran or polyethyleneimine); non-chemical methods such as electroporation, cell squeezing, sonoporation, phototransfection, puncture infection, protoplast fusion, plasmid delivery, transposons; particle-based methods such as gene guns, magnetic or magnetically assisted transfection, particle bombardment; hybridization methods such as nucleofection, etc.

[0048] The terms "transfection" or "transformation" or "transduction" as used herein refer to the process of transferring or introducing exogenous nucleic acid into a host cell. A "transfected" or "transformed" or "transduced" cell is one that has been transfected, transformed, or transduced with exogenous nucleic acid.

[0049] "In vivo" means within the body of the organism from which the cells were obtained. "Ex vivo" or "in vitro" means outside the organism from which the cells were obtained.

[0050] As used herein, "treatment" or "treatment" is a method of obtaining a beneficial or desired result (including a clinical result). For purposes of this application, beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms caused by the disease, reducing the extent of the disease, stabilizing the disease (e.g., preventing or slowing the progression of the disease), preventing or slowing the spread of the disease (e.g., metastasis), preventing or slowing the recurrence of the disease, reducing the recurrence rate of the disease, slowing or reducing the progression of the disease, improving the condition of the disease, providing palliation (partial or total) of the disease, reducing the dose of one or more other drugs required to treat the disease, slowing the progression of the disease, improving the quality of life, and / or increasing survival time. "Treatment" also includes reducing the pathological consequences of cancer. The methods of this application contemplate any one or more of these aspects of treatment.

[0051] As used herein, the term "effective amount" refers to an amount of a compound or composition sufficient to treat (e.g., ameliorate, alleviate, attenuate, and / or delay one or more symptoms of) a particular disease, condition, or disorder. As is understood in the art, an "effective amount" may be one or more doses, i.e., single or multiple doses may be required to achieve a desired treatment endpoint.

[0052] When used for therapeutic purposes, the terms "subject," "individual," or "patient" may be used interchangeably herein and refer to any animal, such as mammals (humans; domestic and farm animals; including zoo, sports, or pet animals, such as dogs, horses, cats, hamsters, guinea pigs, rabbits, monkeys, sheep, cows, etc.), birds, reptiles, fish, etc. In some embodiments, the individual is a human individual.

[0053] It should be understood that embodiments of the present application described herein include embodiments that "consist of" and / or "consist essentially of".

[0054] As used herein, reference to "about" a value or parameter includes (and describes) a variation on that value or parameter itself. For example, a description of "about X" also includes the description of "X."

[0055] As used herein, a reference to a value or parameter "is not" generally refers to "other than" the value or parameter. For example, the method is not used to treat cancer type X means that the method is used to treat cancers other than type X.

[0056] As used herein, "about X to Y" has the same meaning as "about X to about Y."

[0057] As used in this specification and the appended claims, the singular forms "a," "one," and "the" include plural references unless the context clearly dictates otherwise. It should also be noted that the claims may be written to exclude any element. Accordingly, this statement is intended as a predicate for using specialized language such as "only," "only one," or "negative" limitations in describing elements of the claims.

[0058] As used herein, the term "and / or" in "A and / or B" is intended to include A and B, A or B, A (single), and B (single). Similarly, as used herein, the term "and / or" in "A, B, and / or C" is intended to encompass each of the following embodiments: A, B, and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A (single); B (single); and C (single).

[0059] Those skilled in the art will understand that instead of representing uracil with "u" and thymine with "t," both uracil and thymine can be represented by "t." In the context of ribonucleic acids, unless otherwise specified, "t" should be understood to represent uracil.

[0060] II. Cas12b Nuclease and Effector Proteins The present application provides engineered Cas12b nucleases and effector proteins with improved activities, such as target binding, double-strand cleavage activity, nickase activity, and / or gene editing activity. Engineered Cas12b nucleases (dCas12b) are also provided that have reduced or eliminated nuclease activity. In some embodiments, engineered Cas12b effector proteins (e.g., Cas12b nucleases, Cas12b nickases, Cas12b fusion effector proteins, or truncated Cas12b effector proteins) are provided that include any of the engineered Cas12b nucleases or functional derivatives thereof described herein.

[0061] Engineered Cas12b nuclease In one aspect, the present application provides engineered Cas12b effector proteins with improved activity (e.g., target binding, double-strand cleavage activity, nickase activity, and / or gene editing activity).

[0062] In some embodiments, engineered Cas12b nucleases are provided that include one, two, or three types of mutations relative to a reference Cas12b nuclease, the mutations including (1) substitutions of one or more amino acid residues that interact with a protospacer adjacent motif (PAM) in the reference Cas12b nuclease with positively charged amino acid residues (e.g., R, H, K), and / or (2) substitutions of one or more amino acid residues that are involved in DNA double-stranded (dsDNA) cleavage in the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (3) substitutions of one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with positively charged amino acid residues (e.g., R, H, K) or hydrophobic amino acid residues (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease is a naturally occurring wild-type Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a naturally occurring variant Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a Cas12b nuclease from Alicyclobacillus acidophilus (AaCas12b). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease has improved activity (e.g., at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 1.2-fold, 1.5-fold, 2-fold, 5-fold, 10-fold, 20-fold, 50-fold, 100-fold or more improved) compared to a reference Cas12b nuclease (e.g., target binding, double-strand cleavage activity, nickase activity, and / or gene editing activity).

[0063] An engineered Cas12b nuclease may include one or more of the mutations described below in Sections A-C. In some embodiments, one or more of the mutations herein may be combined with any of the known Cas12b mutations (e.g., those described below in Section D) to generate an engineered Cas12b nuclease with improved activity.

[0064] In some embodiments, engineered Cas12b nucleases are provided that include one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising a substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R or K). In some embodiments, engineered Cas12b nucleases are provided that include one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising a substitution of one or more amino acid residues involved in DNA duplex cleavage in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., W, Y or F). In some embodiments, engineered Cas12b nucleases are provided that include one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising a substitution of a positively charged amino acid residue (e.g., R or K) from one or more amino acid residues that interact with single-stranded DNA substrates in the RuvC domain of the reference Cas12b nuclease. In some embodiments, engineered Cas12b nucleases are provided that include one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising a substitution of a hydrophobic amino acid residue (e.g., W, Y, F or M) from one or more amino acid residues that interact with single-stranded DNA substrates in the RuvC domain of the reference Cas12b nuclease. In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1.

[0065] In some embodiments, engineered Cas12b nucleases are provided that comprise one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising: 1) a substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K); and 2) a substitution of one or more amino acid residues that are involved in DNA duplex cleavage in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W). In some embodiments, engineered Cas12b nucleases are provided that comprise one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising: 1) a substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K); and 2) a substitution of one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M). In some embodiments, an engineered Cas12b nuclease is provided that comprises one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising: 1) substitution of one or more amino acid residues involved in DNA double-strand cleavage in the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and 2) substitution of one or more amino acid residues interacting with single-stranded DNA substrates in the RuvC domain of the reference Cas12b nuclease with positively charged amino acid residues (e.g., R, H, K) or hydrophobic amino acid residues (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1.

[0066] In some embodiments, an engineered Cas12b nuclease is provided that comprises one or more mutations relative to a reference Cas12b nuclease, the one or more mutations comprising: 1) a substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K), 2) a substitution of one or more amino acid residues that participate in DNA double-strand cleavage in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W), and 3) a substitution of one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1.

[0067] In some embodiments, engineered Cas12b nucleases are provided that include one or more mutations relative to a reference Cas12b nuclease, the one or more mutations including: 1) substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease with positively charged amino acid residues (e.g., R, H, K), 2) substitution of one or more amino acid residues that participate in DNA double-strand cleavage in the reference Cas12b nuclease with amino acid residues having aromatic rings (e.g., F, Y, W), and 3) substitution of one or more amino acid residues that cross-reference with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with hydrophobic amino acid residues (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1.

[0068] The mutations described herein can be designed based on the structure of the reference Cas12b nuclease. Yang H., et al. Cell 167:1814-1828(2016) and Liu L. et al. Mol. Cell 65:310-322(2017) describe the binding of Alicyclobacillus acidoterrestris Cas12b and sgRNA to obtain a binary complex, which is then bound to target DNA to obtain a ternary complex crystal structure. Briefly, the crystal structure shows two discontinuous REC (recognition, residues 15-386, 658-783) and NUC (nuclease, residues 1-14, 387-658, and 784-1129) lobes, each of which is composed of several domains. The crRNA (or single guide RNA, sgRNA) is bound to the central passage between the two lobes. PAM recognition is sequence-specific and occurs primarily through interactions with the REC1 (helical-1) and WED-II (OBD-II) domains. The sgRNA-target DNA heteroduplex is bound to the REC lobe in a primarily sequence-independent manner.

[0069] It should be understood that other Cas12b orthologues, such as BhCas12b (SEQ ID NO:59), Bs3Cas12b (SEQ ID NO:56), LsCas12b (SEQ ID NO:58), SbCas12b (SEQ ID NO:60), AkCas12b (SEQ ID NO:54), AmCas12b (SEQ ID NO:55), BsCas12b (SEQ ID NO:57), and DiCas12b, have similar domain structures to AaCas12b (SEQ ID NO:1) and other exemplary reference Cas12b proteins described herein, and engineered AaCas12b proteins may be designed based on any orthologue using cleavage positions corresponding to the exemplary engineered AaCas12b proteins described herein. Corresponding positions refer to positions in two polypeptides that align with each other when the amino acid sequences of the two polypeptides are aligned with each other. See Figure 8 herein. Additionally, Figure S2 in Teng F. et al., Cell Discovery (2019) 5:23 provides an alignment of AaCas12b, AkCas12b, AmCas12b, Bs3Cas12b, BsCas12b, LsCas12b, BhCas12b, and SbCas12b, which is incorporated herein by reference in its entirety.

[0070] A. Substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b with a positively charged amino acid residue. In some embodiments, the engineered Cas12b nuclease comprises a substitution of a positively charged amino acid residue (e.g., R, H, K) from one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease. In some embodiments, the engineered Cas12b nuclease comprises 1, 2, 3, 4, 5, or 6 amino acid substitutions.

[0071] In some embodiments, the one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are within 15 (e.g., within any of 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) Å of the PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are within 10 Å of the PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are within 9 Å of the PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with the PAM are located at one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475. In some embodiments, the one or more amino acid residues that interact with PAM include one or more of the following amino acid residues: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and E475. In some embodiments, the one or more amino acid residues that interact with PAM include one or more of the following amino acid residues: D116 and E475. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0072] In the context of this application, D116 refers to the 116th amino acid in the cited amino acid sequence, D (aspartic acid). Commonly used three-letter and one-letter abbreviations of amino acids are as follows:

[0073] JPEG2025501473000001.jpg43142

[0074] As used herein, "an amino acid at position X, where the amino acid is numbered according to SEQ ID NO:1" refers to an amino acid residue at a position in the reference enzyme Cas12b that corresponds to position X in SEQ ID NO:1 when the amino acid sequences of the reference enzyme Cas12b and SEQ ID NO:1 are aligned based on sequence homology. For example, FIG. 8 shows a homology alignment of the amino acid sequences of Cas12b orthologues (SEQ ID NO:1 and 54-60). One skilled in the art can use known software such as Clustal Omega to align the amino acid sequence of any reference Cas12b nuclease with SEQ ID NO:1 and determine the amino acid position that corresponds to position X in SEQ ID NO:1.

[0075] In some embodiments, the positively charged amino acid residue is R, H or K. In some embodiments, the positively charged amino acid residue is R. In some embodiments, the positively charged amino acid residue is K.

[0076] In some embodiments, the substitutions of positively charged amino acid residues from one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are one or more of the following substitutions: D116R, K123R, D130R, D132R, N144R, K145R, E153R, D173R, Q222R, D395R, N400R, and E475R. In some embodiments, the substitutions of positively charged amino acid residues from one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are one or more of the following substitutions: D116R and E475R. In some embodiments, the engineered Cas12b nuclease comprises a D116R mutation. In some embodiments, the engineered Cas12b nuclease comprises an E475R mutation. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0077] In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 87%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity) to the amino acid sequence of SEQ ID NO:2 or 3. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:2 or 3.

[0078] B. Substitution of one or more amino acid residues involved in DNA duplex cleavage in the reference Cas12b nuclease with amino acid residues having an aromatic ring In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues involved in DNA duplex cleavage in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W). In some embodiments, the engineered Cas12b nuclease comprises a substitution of 1, 2, 3, 4, 5, or 6 amino acid residues.

[0079] In some embodiments, one or more amino acid residues involved in DNA duplex cleavage interact with the last base pair to the 3' end of the target strand in PAM.For example, the PAM sequence recognized by AaCas12b is 5'-TTN-3' base pair.The last base pair to the 3' end of the target strand in PAM is the base pair formed by the N base at the 3' end of the PAM sequence, followed by the sequence of the target site.

[0080] In some embodiments, the one or more amino acid residues involved in the cleavage of the DNA duplex are located at one or more of positions 118 and 119, such as Q118 and Q119. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0081] In some embodiments, the amino acid residue having an aromatic ring is Y, F, or W. In some embodiments, the amino acid residue involved in DNA duplex cleavage is substituted with F, Y, or W. In some embodiments, the engineered Cas12b nuclease comprises any of i) Q118Y, Q118F, or Q118W; and / or ii) Q119Y, Q119F, or Q119W. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0082] In some embodiments, the substitution of one or more amino acid residues involved in DNA duplex cleavage in the reference Cas12b nuclease with an amino acid having an aromatic ring is Q119Y, Q119F, or Q119W. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0083] In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity) to the amino acid sequence of SEQ ID NO:4, 5 or 6. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:4, 5 or 6.

[0084] C. Substitution of one or more amino acid residues that interact with single-stranded DNA substrates in the RuvC domain of the reference Cas12b nuclease with positively charged or hydrophobic amino acid residues. In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease that interact with single-stranded DNA substrates with positively charged amino acid residues (e.g., R, H, K). In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease that interact with single-stranded DNA substrates with hydrophobic amino acid residues (e.g., F, Y, W, M). In some embodiments, the engineered Cas12b nuclease comprises a substitution of 1, 2, 3, 4, 5, or 6 amino acid residues.

[0085] In some embodiments, one or more amino acid residues located in the RuvC domain that interact with a single-stranded DNA substrate are within 15 Å (e.g., within any of 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) of the single-stranded DNA substrate in the three-dimensional structure. In some embodiments, one or more amino acid residues located in the RuvC domain that interact with a single-stranded DNA substrate are within 10 Å of the single-stranded DNA substrate in the three-dimensional structure. In some embodiments, one or more amino acid residues located in the RuvC domain that interact with a single-stranded DNA substrate are within 9 Å of the single-stranded DNA substrate in the three-dimensional structure.

[0086] The RuvC domain is the active domain of the Cas12b protein and is responsible for cleaving single-stranded or double-stranded DNA. In the primary sequence of the protein, the RuvC domain includes the first RuvC domain (RuvC-1), the second RuvC domain (RuvC-II) and the third RuvC domain (RuvC-III).

[0087] In some embodiments, the one or more amino acid residues located within the RuvC domain and that interact with a single-stranded DNA substrate are located at one or more of positions 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097. In some embodiments, the one or more amino acid residues located within the RuvC domain and that interact with a single-stranded DNA substrate include one or more of the following amino acid residues: D300, K301, E304, N329, E636, Q639, T647, Q682, I757, E758, E761, E764, K768, E852, Q854, N856, N857, D858, P860, S862, E863, N865, Q866, L867, Q869, E938, E956, G957, E958, I994, Q1093, and W1097. In some embodiments, the one or more amino acid residues located within the RuvC domain and interacting with a single-stranded DNA substrate comprise one or more of the following amino acid residues: D300, K301, E636, Q639, T647, Q682, I757, E758, E761, K768, Q854, N857, D858, N865, Q866, Q869, I994, Q1093, and W1097. In some embodiments, the one or more amino acid residues located within the RuvC domain and interacting with a single-stranded DNA substrate comprise one or more of the following amino acid residues: E636, I757, E758, E761, Q854, N857, D858, N865, Q866, Q869, and Q1093. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0088] In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease that interact with single-stranded DNA substrates with a positively charged amino acid residue (e.g., R, H, K). In some embodiments, the positively charged amino acid residue is R. In some embodiments, the positively charged amino acid residue is K. In some embodiments, the engineered Cas12b nuclease is selected from the group consisting of D300R, K301R, E304R, N329R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, E764R, K768R, E852R, Q854R, N856R, N857R, D858R, P860R, S862R, E863R, N865R, Q866R , L867R, Q869R, E938R, E956R, G957R, E958R, I994R, Q1093R, W1097R, E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, Q866K, I994K, Q1093K, and W1097K substitutions, wherein the amino acid residues are numbered according to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D300R, K301R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, K768R, Q854R, N857R, D858R, N865R, Q866R, I994R, Q1093R, W1097R, E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, I994K, Q1093K, and W1097K, wherein the amino acid residues are numbered according to SEQ ID NO:1.In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R, D858R, E636K, I757K, E758K, E761K, Q854K, N857K, and D858K, where amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R, and D858R, where amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636K, I757K, E758K, E761K, Q854K, N857K, and D858K, where the amino acid residues are numbered according to SEQ ID NO:1.

[0089] In some embodiments, the substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease and interacting with single-stranded DNA substrates is one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R, N857K, and D858R, where the amino acid residues are numbered according to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity) with any of the amino acid sequences of SEQ ID NO:7-13. In some embodiments, the engineered Cas12b nuclease comprises any of the amino acid sequences of SEQ ID NO:7-13.

[0090] In some embodiments, the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues located in the RuvC domain in a reference Cas12b nuclease that interacts with a single-stranded DNA substrate with a hydrophobic amino acid residue. In some embodiments, the hydrophobic amino acid residue is A, M, L, I, V, C, Y, F, or W. In some embodiments, the hydrophobic amino acid residue is W, Y, F, or M. In some embodiments, the hydrophobic amino acid residue is W, Y, or M. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: i) E758W, E758Y, E758F or E758M; ii) E761W, E761Y, E761F or E761M; iii) E863W, E863Y, E863F or E863M; iv) N865W, N865Y, N865F or N865M; v) Q866W, Q866F, Q866Y or Q866M; vi) Q869W, Q869Y, Q869F or Q869M; vii) E956W, E956Y, E956F or E956M; and viii) Q1093W, Q1093F, Q1093Y or Q1093M; In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions, where the amino acid residues are numbered according to SEQ ID NO:1: i) E758W, E758Y or E758M, ii) E761Y, iii) N865W, N865F or N865Y, iv) Q866M, v) Q869M, and vi) Q1093W, Q1093F, Q1093Y or Q1093M, and the amino acid residues are numbered according to SEQ ID NO:1: In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions, where the amino acid residues are numbered according to SEQ ID NO:1: i) N865W or N865Y, ii) Q866M, iii) Q869M, and iv) Q1093W or Q1093Y, and the amino acid residues are numbered according to SEQ ID NO:1: In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: 865W, 865Y, 866M, 869M, 1093W, and 1093Y.In some embodiments, the substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease and interacting with single-stranded DNA substrates is one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W, and Q1093Y. In some embodiments, the engineered Cas12b nuclease comprises a Q866M and a Q869M substitution. In some embodiments, the amino acid residues are numbered according to SEQ ID NO:1.

[0091] In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity) to any of the amino acid sequences of SEQ ID NOs: 14-20. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of any of SEQ ID NOs: 14-20.

[0092] D. Other mutations Any one or more of the mutations described in Sections A-C above may be combined with any one or more of the known mutations that improve Cas12b activity, such as target binding, target specificity, double-stranded cleavage activity, nickase activity, and / or gene editing activity. Exemplary mutations are described in the following documents: WO2022120520, WO2022040909, WO2022042557, CN113308451A, and CN112195164A, the contents of which are incorporated herein by reference in their entirety.

[0093] In some embodiments, the reference Cas12b protein comprises, from N-terminus to C-terminus, one or more of the following: a first WED domain (WED-I), a first REC domain (REC1), a second WED domain (WED-II), a first RuvC domain (RuvC-I), a BH domain, a second REC domain (REC2), a second RuvC domain (RuvC-II), a first Nuc domain (Nuc-I), a third RuvC domain (RuvC-III), and a second Nuc domain (Nuc-II). In some embodiments, one or more other mutations (e.g., insertions, deletions, substitutions) may be present in one or more such domains.

[0094] In some embodiments, the engineered Cas12b nuclease further comprises one or more flexibility region mutations that improve the flexibility of the flexible region in the reference Cas12b nuclease. The flexibility region in the reference Cas12b nuclease can be determined using any method known in the art. In some embodiments, the plurality of flexibility regions is determined based only on the amino acid sequence of the reference enzyme. In some embodiments, the plurality of flexibility regions is determined based on structural information of the reference enzyme, including, for example, secondary structure, crystal structure, NMR structure, etc.

[0095] In some embodiments, the plurality of flexible regions are determined using a program selected from the group of PredyFlexy, FoldUnfold, PROFbval, Flexserv, FlexPred, DynaMine, and Disomine. In some embodiments, the plurality of flexible regions are located in random crimps. In some embodiments, the plurality of flexible regions are located in the DNA and / or RNA interaction domain of the reference Cas12b nuclease. In some embodiments, the length of the flexible region is at least about 5 (e.g., 5) amino acids.

[0096] In some embodiments, the engineered Cas12b nuclease comprises one or more mutations that increase flexibility in a flexible region of amino acid residues 855-859, the amino acid residues being numbered according to SEQ ID NO: 1, and has improved activity (e.g., target binding, double-strand cleavage activity, nickase activity, and / or gene editing activity) (at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 1.2-fold, 1.5-fold, 2-fold, 5-fold, 10-fold, 20-fold, 50-fold, 100-fold or more) compared to a reference Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is AaCas12b. In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the one or more mutations comprise inserting one or more (e.g., two) G residues in the flexibility region. In some embodiments, the one or more G residues are inserted N-terminal to a flexible amino acid residue in the flexibility region, the flexible amino acid residue being selected from the group consisting of G, S, N, D, H, M, T, E, Q, K, R, A, and P. In some embodiments, the flexible amino acid residues are selected according to the following preference: G>S>N>D>H>M>T>E>Q>K>R>A>P. In some embodiments, the one or more mutations comprise substituting a hydrophobic amino acid residue in the flexibility region with a G residue, the hydrophobic amino acid residue being selected from the group consisting of L, I, V, C, Y, F, and W. In some embodiments, the one or more mutations that increase flexibility comprise N856G.

[0097] E. Combination of Mutations Any engineered enzyme obtained by combining multiple amino acid substitutions in the exemplary sequence listing using the methods described in Sections A-D herein is within the scope of the present application. In some embodiments, the engineered Cas12b nuclease comprises one or more mutations (e.g., substitutions) described in Sections A-D above.

[0098] In some embodiments, the engineered Cas12b nuclease is selected from the group consisting of: (1) 116, (2) 475, (3) 119 and 475, (4) 119, 475, and 758, (5) 119, (6) 636, (7) 757, (8) 758, (9) 761, (10) 768, (11) 858, (12) 854, (13) 857, (14) 119, 475, and 758, (15) 768, (16) 757 and 758, (17) 757 and 761, (18) 757 and 768, (19) 7 (20) 758 and 768, (21) 761 and 768, (22) 757, 758, and 761, (23) 757, 758, and 768, (24) 757, 761 and 768, (25) 758, 761, and 768, (26) 757, 758, 761, and 768, (27) 865, (28) 866, (29) 869, (30) 1093, and (31) 866 and 869 amino acid residue positions, where the amino acid position numbers are relative to SEQ ID NO:1.

[0099] In some embodiments, the engineered Cas12b nuclease is selected from the group consisting of: (1) D116, (2) E475, (3) Q119 and E475, (4) Q119, E475, and E758, (5) Q119, (6) E636, (7) I757, (8) E758, (9) E761, (10) K768, (11) D858, (12) Q854, (13) N857, (14) Q119, E475, and E758, (15) K768, (16) I757 and E758, (17) I757 and E761, (18) I757 and K768, (19) E7 (20) E758 and K768, (21) E761 and K768, (22) I757, E758, and E761, (23) I757, E758, and K768, (24) I757, E761 and K768, (25) E758, E761, and K768, (26) I757, E758, E761, and K768, (27) N865, (28) Q866, (29) Q869, (30) Q1093, and (31) Q866 and Q869, where the amino acid position numbers are relative to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises a substitution or combination of substitutions at any of the amino acid residues at (1) Q866+Q869, (2) Q119+E475, and (3) Q119+E475+E758, where the amino acid residues are numbered according to SEQ ID NO:1. In some embodiments, the substitution at amino acid position D116 and / or E475 is with a positively charged amino acid residue (e.g., R or K). In some embodiments, the substitution at amino acid position Q119 is with an amino acid residue having an aromatic side chain (e.g., Y, F, or W). In some embodiments, the substitution at amino acid position E636, I757, E758, E761, K768, Q854, D858, and / or N857 is with a positively charged amino acid residue (e.g., R or K). In some embodiments, the substitutions at amino acid positions N865, Q866, Q869 and / or Q1093 are with a hydrophobic amino acid residue (eg, W, Y or M).

[0100] In some embodiments, the engineered Cas12b nuclease is selected from the group consisting of: (1) 116R, (2) 475R, (3) 119F and 475R, (4) 119F, 475R, and 758R, (5) 119Y, (6) 119F, (7) 119W, (8) 636R, (9) 757R, (10) 758R, (11) 761R, (12) 854R, (13) 857K, (14) 768R, (15) 757R and 758R, (16) 757R and 761R, (17) 757R and 768R, (18) 758R and 761R, (19) 758R and 768R, (20) (20) 761R and 768R, (21) 757R, 758R, and 761R, (22) 757R, 758R, and 768R, (23) 757R, 761R, and 768R, (24) 758R, 761R, and 768R, (25) 757R, 758R, 761R, and 768R, (26) 865W, (27) 865Y, (28) 866M, (29) 869M, (30) 1093W, (31) 1093Y, (32) 866M and 869M, and (33) 858R, or any combination thereof, wherein the amino acid position numbers are relative to SEQ ID NO:1.

[0101] In some embodiments, the engineered Cas12b nuclease is selected from the group consisting of: (1) D116R, (2) E475R, (3) Q119F+E475R, (4) Q119F+E475R+E758R, (5) Q119Y, (6) Q119F, (7) Q119W, (8) I757R, (9) E758R, (10) E761R, (11) K768R, (12) I757R+E758R, (13) I757R+E761R, (14) I757R+K768R, (15) E758R+E761R, (16) E758R+K768R, (17) E761R+K768R, (18) I757 (20) I757R+E758R+K768R, (21) E758R+E761R+K768R, (22) I757R+E758R+E761R+K768R, (23) Q866M, (24) Q869M, (25) Q866M+Q869M, (26) E636R, (27) Q854R, (28) N857K, (29) N865W, (30) N865Y, (31) Q1093W, (32) Q1093Y, and (33) D858R substitutions or combinations thereof, and the amino acid position numbers are ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises any or a combination of the following substitutions: (1) Q866M+Q869M, (2) Q119F+E475R, and (3) Q119F+E475R+E758R, where the amino acid residues are numbered according to SEQ ID NO:1.

[0102] In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D116R, K123R, D130R, D132R, N144R, K145R, E153R, D173R, Q222R, D395R, N400R, and E475R. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: Q118Y, Q118F, Q118W, Q119Y, Q119F, and Q119W. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D300R, K301R, E304R, N329R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, E764R, K768R, E852R, Q854R, N856R, N857R, D858R, P860R, S862R, E863R, N865R, Q866R, L867R, Q869R, E938R, E956R, G957R, E958R, I944R, Q1093R, and W1097R. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, Q866K, I994K, Q1093K, and W1097K. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E758W, E758Y, E758F, E758M, E761W, E761Y, E761F, E761M, E863W, E863Y, E863F, E863M, N865W, N865Y, N865F, N865M, Q866W, Q866Y, Q866F, Q866M, Q869W, Q869Y, Q869F, Q869M, E956W, E956Y, E956F, E956M, Q1093W, Q1093Y, Q1093F, and Q1093M. In some embodiments, the amino acid position numbers are relative to SEQ ID NO:1.

[0103] In some embodiments, the engineered Cas12b nuclease comprises amino acid substitutions at Q866 and Q869. In some embodiments, the engineered Cas12b nuclease comprises amino acid substitutions at Q866M and Q869M. In some embodiments, the amino acid position numbers are relative to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity) to the amino acid sequence of SEQ ID NO:20. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of SEQ ID NO:20.

[0104] In some embodiments, the engineered Cas12b nuclease comprises amino acid substitutions at Q119 and E475. In some embodiments, the engineered Cas12b nuclease comprises amino acid substitutions at Q119F and E475R. In some embodiments, the amino acid position numbers are relative to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity) to the amino acid sequence of SEQ ID NO:21. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:21.

[0105] In some embodiments, the engineered Cas12b nuclease comprises amino acid substitutions at Q119, E475, and E758. In some embodiments, the engineered Cas12b nuclease comprises amino acid substitutions at Q119F, E475R, and E758R. In some embodiments, the amino acid position numbers are relative to SEQ ID NO:1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity) to the amino acid sequence of SEQ ID NO:22. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of SEQ ID NO:22.

[0106] Reference Cas12b nuclease In some embodiments, the reference Cas12 nuclease is AaCas12b, or an orthologue thereof. In some embodiments, the reference Cas12b nuclease is a naturally occurring Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a wild-type Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is an engineered Cas12b nuclease.

[0107] To provide the engineered Cas12b nuclease and effector protein of the present application, Cas12b nucleases from various organisms can be used as reference Cas12b nucleases. In some embodiments, the reference Cas12b nuclease has enzymatic activity. In some embodiments, the reference Cas12b is a nuclease that cleaves both strands of a target double-stranded nucleic acid (e.g., double-stranded DNA). In some embodiments, the reference Cas12b is a nickase that cleaves a single strand of a target double-stranded nucleic acid (e.g., double-stranded DNA). In some embodiments, the reference Cas12b nuclease is enzymatically inactive (e.g., dCas12b). An orthologue having a certain sequence identity (e.g., at least about 60%, 70%, 80%, 85%, 90%, 95%, 98% or more) with Cas12b or its functional derivative can be used as a basis for designing the engineered Cas12b nuclease or effector protein of the present application. In some embodiments, the reference Cas12b nuclease is a mutant Cas12b, but does not contain any of the mutations described in sections A-E above.

[0108] In some embodiments, the engineered Cas12b nuclease is based on a functional variant of a naturally occurring Cas12b nuclease. In some embodiments, the functional variant has one or more mutations, such as amino acid substitutions, insertions and / or deletions. For example, the functional variant may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid substitutions compared to a naturally occurring wild-type Cas12b nuclease. In some embodiments, the one or more substitutions are conservative substitutions. In some embodiments, the functional variant has all the domains of a naturally occurring Cas12b nuclease. In some embodiments, the functional variant does not have one or more domains of a naturally occurring Cas12b nuclease.

[0109] The VB-type CRISPR-Cas12b (also called C2c1) system has been identified as a dual RNA-guided (i.e., crRNA and tracrRNA) DNA endonuclease system with characteristics distinct from Cas9 and Cas12a (Shmakov, S. et al. Mol. Cell 60, 385-397 (2015)). First, it was reported that Cas12b generates cross-ends at the distal end of the PAM site in vitro when reconstituted with a crRNA / tracrRNA duplex. Second, the RuvC domain of Cas12b is similar to the RuvC domains of Cas9 and Cas12a, but the hypothetical Nuc domain has no sequence or structural similarity to the HNH domain of Cas9 and the Nuc domain of Cas12a. In addition, Cas12b protein is smaller than the most widely used SpCas9 and Cas12a (e.g., AacCas12b: 1,129 amino acids (aa); SpCas9: 1,369 aa; AsCas12a: 1,353 aa; LbCas12a: 1,228 aa), which makes Cas12b suitable for adeno-associated virus (AAV)-mediated in vivo delivery in gene therapy. Cas12b recognizes simpler PAM sequences (e.g., AacCas12b: 5'-TTN-3') compared to the small size of Cas9 proteins (e.g., SaCas9 and CjCas9), which greatly increases the target range of Cas12b in the genome compared to SaCas9: 5'-NNGRRT-3', CjCas9: 5'-NNNNRYAC-3'. Furthermore, Cas12b has minimal off-target effects, making it a safer option for therapeutic and clinical applications.

[0110] To provide the engineered Cas12b effector protein of the present application, Cas12b (C2c1) nucleases from various organisms can be used as the reference Cas12b nuclease. Exemplary Cas12b nucleases are described, for example, in Shmakov, S. et al. Mol. Cell 60, 385-397 (2015), Shmakov, S. et al. Nat. Rev. Microbiol. 15, 169-182 (2017), WO2016205764, and WO2020087631, the contents of which are incorporated herein by reference in their entirety.

[0111] In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b protein (e.g., a reference Cas12b nuclease), such as a Cas12b protein from Alicyclobacillus acidifilus (AaCas12b), a Cas12b from Alicyclobacillus kakegawensis (AkCas12b), a Cas12b from Alicyclobacillus macrosporangiidus (AmCas12b), a Cas12b from Bacillus hisashii (BhCas12b), a BsCas12b from Bacillus spp., a Bs3Cas12b from Bacillus spp., a BsCas12b from Desulfovibrio inpinatus (Desulfovibrio inpinatus), or a Cas12b from Alicyclobacillus acidophilus (AaCas12b). The Cas12b is selected from Cas12b from Bacillus inopinatus (DiCas12b), Cas12b from Raceella sediminis (LsCas12b), Cas12b from Spirochete bacteria (SbCas12b), Cas12b from Tuberibacillus calidus (TcCas12b), and functional derivatives thereof. Sequences of naturally occurring Cas12b proteins are known, for example, from UniProtKB IDs: T0D7A2, A0A6I3SPI6, and A0A6I7FUC4, which are incorporated herein by reference in their entireties.

[0112] In some embodiments, the reference Cas12b protein is a Cas12b nuclease from Alicyclobacillus acidifilus (AaCas12b) or a functional derivative thereof. In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b protein, which comprises an amino acid sequence having at least about 85% (e.g., at least about any of 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to the amino acid sequence of SEQ ID NO:1. In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b nuclease comprising the amino acid sequence of SEQ ID NO:1.

[0113] It should be noted that orthologues having a certain degree of sequence identity (e.g., at least about any of 60%, 70%, 80%, 85%, 90%, 95%, 98% or more) with a reference Cas12b protein or fragment thereof can be used as a basis for designing the engineered Cas12b effector protein of the present application. Those skilled in the art can determine the percent sequence identity of a Cas12b orthologue or fragment thereof suitable for use in the present application depending on the purpose and application. Methods for determining sequence identity values ​​are described in Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). WO2020 / 087631 describes various Cas12b orthologs, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b protein, which comprises an amino acid sequence having at least about 85% (e.g., at least about any of 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to any of the amino acid sequences of SEQ ID NOs:54-60.

[0114] Engineered Cas12b activity In some embodiments, the engineered Cas12b nuclease has improved activity compared to a reference Cas12b nuclease. In some embodiments, the activity is target DNA binding activity. In some embodiments, the activity is site-specific nuclease activity. In some embodiments, the activity is double-stranded DNA cleavage activity. In some embodiments, the activity is single-stranded DNA cleavage activity, including, for example, site-specific DNA cleavage activity or non-specific DNA cleavage activity. In some embodiments, the activity is single-stranded RNA cleavage activity, such as site-specific RNA cleavage activity or non-specific RNA cleavage activity. In some embodiments, the activity is measured in vitro. In some embodiments, the activity is measured in a cell, such as, for example, a bacterial cell, a plant cell, or a eukaryotic cell. In some embodiments, the activity is measured in a mammalian cell (e.g., a rodent cell or a human cell). In some embodiments, the activity is measured in a human cell (e.g., a 293T cell). In some embodiments, the activity is measured in a mouse cell (e.g., a Hepa1-6 cell). In some embodiments, the engineered Cas12b nuclease has at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 50-fold or more increased activity compared to a reference Cas12b nuclease. Site-specific nuclease activity of the engineered Cas12b nuclease can be measured using methods known in the art, including PCR, sequencing, or gel migration assays, as described in the examples provided herein.

[0115] In some embodiments, the activity is gene editing activity in a cell. In some embodiments, the cell is a bacterial cell, a plant cell, or a eukaryotic cell. In some embodiments, the cell is a mammalian cell, such as a rodent cell or a human cell. In some embodiments, the cell is a 293T cell. In some embodiments, the activity is measured in a mouse cell (e.g., Hepa1-6 cell). In some embodiments, the activity is insertion / deletion formation activity at a target genomic site in a cell, such as site-specific cleavage of a target nucleic acid by an engineered Cas12b nuclease and DNA repair by a non-homologous end joining (NHEJ) mechanism. In some embodiments, the activity is insertion of an exogenous nucleic acid sequence at a target genomic site in a cell, such as site-specific cleavage of a target nucleic acid by an engineered Cas12b nuclease and DNA repair by a homologous recombination (HR) mechanism. In some embodiments, the homologous recombination after cleavage by an engineered Cas12b nuclease further comprises the introduction of a donor template. In some embodiments, the engineered Cas12b nuclease has improved gene editing activity (e.g., insertion / deletion formation) at least about 20% (e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 50-fold or more) at genomic sites in a cell (e.g., a human cell, a 293T cell, or a mouse Hepa1-6 cell) compared to a reference Cas12b nuclease. In some embodiments, the engineered Cas12b nuclease can edit a greater number of genomic sites (e.g., 2, 3, 4, 5, 10, 20, 50, 100 or more) than the reference Cas12b nuclease. In some embodiments, the shared PAM sequence of the engineered Cas12b nuclease is the same as the reference Cas12b nuclease. In some embodiments, the engineered Cas12b nuclease recognizes more (e.g., 1, 2, 3, 4, 5, 10, 20, 50, 100 or more) PAM sequences compared to a reference Cas12b nuclease.

[0116] The cleavage or gene editing efficiency of the engineered Cas12b nuclease in cells can be determined using any method known in the art, including, for example, T7 endonuclease 1 (T7E1) assay, PCR, targeted DNA sequencing (including, for example, Sanger sequencing and second generation sequencing), deletion tracking insertion and deletion (TIDE) assay, or insertion / deletion detection by amplicon analysis (IDAA) assay. See, for example, Sentmanat MF et al., "A survey of validation strategies for CRISPR-Cas9 editing," Scientific Reports, 2018, 8, article number 888, the contents of which are incorporated herein by reference in their entirety. In some embodiments, for example, as described in the Examples herein, the gene editing efficiency of the engineered Cas12b nuclease in cells is measured using targeted next-generation sequencing (NGS). Genomic sites for determining the cleavage or gene editing efficiency of an engineered Cas12b nuclease include, but are not limited to, CCR5, AAVS, CD34, RNF2, SCN9A, HBG1 / 2, and EMX1. In some embodiments, the engineered Cas12b nuclease can cleave or edit at least about 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 100 or more loci compared to the average cleavage or gene editing efficiency of a reference Cas12b nuclease in a human cell genome. In some embodiments, the cleavage or gene editing efficiency (e.g., insertion / deletion rate) of the engineered Cas12b nuclease is at least about 10%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 1-fold, 2-fold, 5-fold, 10-fold, 20-fold, 50-fold or more that of a reference Cas12b nuclease.

[0117] Engineered Cas12b effector proteins The present application also provides engineered Cas12b effector proteins based on any of the engineered Cas12b nucleases, variants (e.g., dCas12b) or functional derivatives described herein. In some embodiments, the engineered Cas12b effector proteins comprise (or consist of, or consist essentially of) any of the engineered Cas12b nucleases, variants, or functional derivatives described herein. In some embodiments, the engineered Cas12b effector proteins comprise a functional derivative of an engineered Cas12b nuclease, such as any of the functional derivatives described in the "Functional Derivatives" section below.

[0118] In some embodiments, the engineered Cas12b effector protein has enzymatic activity. In some embodiments, the engineered Cas12b effector protein is a nuclease that cleaves both strands of a target double-stranded nucleic acid (e.g., double-stranded DNA). In some embodiments, the engineered Cas12b effector protein is a nickase that cleaves a single strand of a target double-stranded nucleic acid (e.g., double-stranded DNA). In some embodiments, the engineered Cas12b effector protein comprises an enzymatically inactive mutant of engineered Cas12b nuclease (dCas12b). Mutation of one or more amino acid residues in the Cas12b nuclease active site can cause a catalytically dead Cas12b (dCas12b). For example, the D570A, E848A, R785A, E848A, R911A, and / or D977A mutants of AaCas12b (SEQ ID NO:1) have significantly reduced (e.g., at least about any of 60%, 70%, 80%, 90%, 95% or more) or no nuclease activity in human cells. See, e.g., Teng F. et al., Cell Discovery, 4, Article number: 63 (2018), the contents of which are incorporated herein by reference in their entirety. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having one or more mutations to D570A, E848A, R785A, E848A, R911A, and D977A of AaCas12b. In some embodiments, one or more mutations selected from the group consisting of D570A, E848A, R785A, E848A, R911A, and D977A are further introduced into AaCas12b comprising the Q119F+E475R+E758R mutations. In some embodiments, the engineered Cas12b nuclease enzyme-inactive mutant comprises any of the amino acid sequences of SEQ ID NOs:79-81.In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the R785A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the R911A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the D977A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the E848A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the D570A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having mutations corresponding to the D570A+E848A mutation of AaCas12b or the D570A+D977A mutation of AaCas12b.

[0119] In some embodiments, an engineered Cas12b nickase is provided. In some embodiments, an engineered Cas12b fusion effector protein is provided that includes an engineered Cas12b nuclease, variant or functional derivative thereof, fused to a functional domain, such as a translation initiation domain, a transcriptional repressor domain (e.g., a Kruppel-associated box (KRAB) domain), a transactivation domain, an epigenetic modification domain, a nucleobase editing domain (e.g., a cytosine base editor (CBE) or an adenine base editor (ABE) domain), a reverse transcriptase domain, a reporter domain (e.g., a fluorescent domain), or a nuclease domain (e.g., a ZFN domain). In some embodiments, an engineered Cas12b base editor is provided that includes any of the catalytically inactive variants of the engineered Cas12b nucleases described herein (e.g., any of SEQ ID NOs:79-81) fused to a cytosine deaminase domain or an adenosine deaminase domain. In some embodiments, an engineered Cas12b base editor is provided that comprises any catalytically inactive variant of an engineered Cas12b nuclease described herein (e.g., any of SEQ ID NOs:79-81) fused to a KRAB domain or a functional fragment thereof (e.g., ZIM3 KRAB domain (SEQ ID NO:72)). In some embodiments, an engineered Cas12b prime editor is provided that comprises any catalytically inactive variant of an engineered Cas12b nuclease described herein (e.g., any of SEQ ID NOs:79-81) fused to a reverse transcriptase domain. In some embodiments, a truncated Cas12b effector protein system is provided.

[0120] Variants / functional derivatives The present application also provides variants and functional derivatives of any of the engineered Cas12b nucleases described herein. In some embodiments, an engineered Cas12b effector protein is provided that comprises (or consists of, or consists essentially of) any of the functional variants of the engineered Cas12b nucleases described herein. In some embodiments, the amino acid sequence of the functional variant has at least one amino acid residue difference (e.g., deletion, insertion, substitution, and / or fusion) compared to the amino acid sequence of the corresponding engineered Cas12b nuclease (e.g., any of SEQ ID NOs: 2-22). In some embodiments, the functional variant has one or more mutations, such as amino acid substitutions, insertions, and / or deletions. For example, the functional variant may include any of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more amino acid substitutions compared to the engineered Cas12b nuclease. In some embodiments, the one or more substitutions are conservative substitutions. In some embodiments, the functional variant comprises all domains of an engineered Cas12b nuclease. In some embodiments, the functional variant comprises none of one or more domains of an engineered Cas12b nuclease.

[0121] For any of the Cas12b variant proteins described herein (e.g., nickase Cas12b proteins, inactive or catalytically inactivated Cas12b (dCas12b), fusion Cas12b), the Cas12b variant may comprise the same parameters (e.g., domains, percent sequence identity, etc.) as any of the Cas12b protein sequences described herein.

[0122] Exemplary mutations in Cas12b functional variants are described in WO2016205764, WO2016205749, and WO2020 / 087631, the contents of which are incorporated by reference in their entireties.

[0123] Catalytic activity In some embodiments, the functional variant has a different catalytic activity compared to a non-mutated form of the engineered Cas12b nuclease. In some embodiments, the mutation (e.g., amino acid substitution, insertion and / or deletion) is located within the catalytic domain (e.g., RuvC domain) of the Cas12b effector protein. In some embodiments, the variant comprises mutations in multiple catalytic domains. A Cas12b effector protein that cleaves one strand of a double-stranded target nucleic acid but not the other strand is referred to herein as a "nickase" (e.g., "nickase Cas"). In some embodiments, the engineered Cas12b effector protein comprises (or consists of, or consists essentially of) a nickase mutant of an engineered Cas12b nuclease. A Cas12b protein that has substantially no nuclease activity is referred to herein as a dead Cas12b protein ("dCas12b") (although when fused to a Cas12b effector protein, nuclease activity may be provided by a heterologous polypeptide (fusion partner) as described below). In some embodiments, a Cas12b effector protein is considered to be substantially devoid of all DNA cleavage activity if the DNA cleavage activity of the mutated Cas12b is about 25%, 20%, 10%, 5%, 1%, 0.1%, 0.01% or less of the DNA cleavage activity of the unmutated form.

[0124] In some embodiments, the engineered Cas12b nuclease is dCas12b. In some embodiments, the engineered Cas12b functional variant comprises a mutation corresponding to D570A of AaCas12b (SEQ ID NO:1). In some embodiments, the engineered Cas12b functional variant comprises a mutation corresponding to E848A of AaCas12b. In some embodiments, the engineered Cas12b functional variant comprises a mutation corresponding to R785A of AaCas12b. In some embodiments, the engineered Cas12b functional variant comprises a mutation corresponding to E848A of AaCas12b. In some embodiments, the engineered Cas12b functional variant comprises a mutation corresponding to R991A of AaCas12b. In some embodiments, the engineered Cas12b functional variant comprises a mutation corresponding to D977A of AaCas12b. In some embodiments, the engineered functional variant of Cas12b comprises a mutation corresponding to D573A in BthCas12b. In some embodiments, the catalytically inactive or substantially inactive variant of AaCas12b (Q119F+E475R+E758R) further comprises one or more substitutions selected from the group consisting of D570A, E848A, and D977A, said amino acid positions corresponding to SEQ ID NO: 22. In some embodiments, dCas12b comprises any of the amino acid sequences of SEQ ID NOs: 79-81.

[0125] Cleaved Cas12b effector protein The CRISPR-Cas12b systems described herein can include any pair of polypeptides comprising a truncated Cas12b portion in this section (also referred to herein as a "truncated Cas12b polypeptide"). Exemplary truncated Cas12b protein systems are described, for example, in PCT / CN2020 / 111057 and PCT / CN2021 / 114339, the contents of which are incorporated herein by reference in their entireties.

[0126] In some embodiments, a truncated Cas12b effector protein is provided, comprising a first polypeptide comprising an N-terminal portion of any of the engineered Cas nucleases, functional variants or derivatives thereof (also referred to in this section as "parent Cas12b proteins") described herein, and a second polypeptide comprising a C-terminal portion of the engineered Cas nuclease, variants or functional derivatives thereof, wherein the first and second polypeptides associate with each other in the presence of a guide RNA comprising a guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence. In some embodiments, the first and second polypeptides each comprise a dimerization domain. In some embodiments, the first and second dimerization domains associate with each other in the presence of an inducer (e.g., rapamycin). In some embodiments, the first and second polypeptides do not comprise any dimerization domain. In some embodiments, the truncated Cas12b effector protein is auto-induced.

[0127] The cleaved Cas12b portion is designed based on an engineered Cas12b nuclease, a variant or a functional variant thereof described herein.

[0128] Cas12b protein has multiple structural domains. In some embodiments, the parent Cas12b protein comprises, from N-terminus to C-terminus, a first WED domain (WED-I; also referred to as OBD-I domain), a first REC domain (REC1), a second WED domain (WED-II; also referred to as OBD-II domain), a first RuvC domain (RuvC-I), a bridge helix (BH) domain, a second RuvC domain (RuvC-II), a first Nuc domain (Nuc-I; also referred to as UK-I domain), a third RuvC domain (RuvC-III), and a second Nuc domain (Nuc-II; also referred to as UK-II domain). Domain boundaries can be determined using methods known in the art based on the crystal structure of a naturally occurring Cas12b protein (e.g., PDB ID numbers for AaCas12b: 5U30, 5U31, 5U33, 5U34, and 5WQE) and / or sequence homology with known functional domains of a parent Cas12b protein. In some embodiments, AaCas12b comprises a WEB-I domain (amino acid residues 1-14), a REC1 domain (amino acid residues 15-386), a WED-II domain (amino acid residues 387-518), a RuvC-I domain (amino acid residues 519-628), a BH domain (amino acid residues 629-658), a REC2 domain (amino acid residues 659-783), a RuvC-II domain (amino acid residues 784-900), a Nuc-I domain (amino acid residues 901-974), a RuvC-III domain (amino acid residues 975-993), and a Nuc-II domain (amino acid residues 994-1129), wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0129] The engineered Cas12b nuclease, its variants, or functional derivatives are cleaved because the two cleavage Cas12b parts contain substantially functional Cas12b. Cas12b can function as a genome editing enzyme (when complexed with target DNA and guide RNA), such as a nuclease that cleaves one or both strands of double-stranded nucleic acid. Alternatively, it can be a catalytically dead Cas12b (dCas12b), a DNA binding protein with little or no catalytic activity, typically due to mutations in the catalytic domain. Mutations of one or more amino acid residues in the reference Cas12b active site can result in a catalytically dead Cas12b, such as the D570A, E848A, R785A, E848A, R911A, and / or D977A mutants of AaCas12b.

[0130] The truncated Cas12b portions described herein are designed such that an engineered Cas12b nuclease, variant or functional derivative thereof (herein referred to as a "parent Cas12b protein"; any of SEQ ID NOs:2-22 and SEQ ID NOs:79-81) (e.g., a full-length Cas12b protein or a functional variant thereof) is split in half at a cleavage site, which is the point where the N-terminal portion of the parent Cas12b protein separates from the C-terminal portion. In some embodiments, the N-terminal portion comprises amino acid residues 1 through X, and the C-terminal portion comprises amino acid residue X+1 through the C-terminus of the parent Cas12b protein. In this example, the numbering is consecutive, but this is not required, as amino acids (or the nucleotides encoding them) may be trimmed from either cleavage terminus and / or mutated (e.g., insertions, deletions, substitutions) in internal regions of the polypeptide chain while retaining sufficient DNA binding activity of the reconstituted Cas12b protein, and optionally retaining at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more DNA nickase or double-stranded cleavage activity compared to the parent Cas12b protein.

[0131] For the engineered Cas12b nucleases described herein, truncated Cas12b portions having several N- and / or C-terminal truncations or deletions and / or internal mutations are also contemplated. One skilled in the art can readily use the information regarding the exemplary truncated Cas12b polypeptides described herein to design corresponding truncated Cas12b polypeptides based on other Cas12b proteins and functional variants, for example using standard sequence alignment tools.

[0132] The cleavage site may be located within a flexible region, such as within a ring. Preferably, the cleavage site occurs where the interruption of the amino acid sequence does not lead to partial or complete destruction of a structural feature (e.g., an alpha helix or a beta sheet). Unstructured regions (regions not shown in the crystal structure, regions not shown in the crystal structure because they are not structured enough to be "frozen" in the crystal) are generally the preferred choice. It is believed that cleavage may occur in unstructured regions exposed on the surface of the parent Cas12b protein.

[0133] In some embodiments, the parent Cas12b protein is not truncated at or near (e.g., within about 10, 8, 6, 5, 4, 3, 2, or 1 amino acid residues) amino acid residues involved in interacting with the guide RNA and / or target RNA. For example, amino acid residues 4-9, 118-122, 143-144, 442-446, 573-574, 742-746, 753-754, 792-796, 800-819, 835-839, 897-900, and 973-978 of the AaCas12b protein are involved in interacting with the single guide RNA and / or target DNA, numbered based on SEQ ID NO:1.

[0134] In some embodiments, the parent Cas12b protein is truncated at amino acid residues within amino acid residues corresponding to amino acid residues 516-793 of the AaCas12b protein, as numbered according to SEQ ID NO:1. In some embodiments, the parent Cas12b protein is truncated at amino acid residues bordering the WED-II domain and the RuvC-I domain. In some embodiments, the parent Cas12b protein is truncated at amino acid residues within amino acid residues corresponding to amino acid residues 516-519 of the AaCas12b protein, as numbered according to SEQ ID NO:1. In some embodiments, the parent Cas12b protein is truncated at amino acid residues bordering the BH domain and the REC2 domain. In some embodiments, the parent Cas12b protein is truncated at amino acid residues within amino acid residues corresponding to amino acid residues 621-627 of the AaCas12b protein, as numbered according to SEQ ID NO:1. In some embodiments, the parent Cas12b protein is truncated at amino acid residues bordering the REC2 domain and the RuvC-II domain. In some embodiments, the parent Cas12b protein is truncated at an amino acid residue within amino acid residues corresponding to amino acid residues 777-793 of the AaCas12b protein, numbering based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is truncated within the RCE2 domain. In some embodiments, the parent Cas12b protein is truncated at an amino acid residue within amino acid residues corresponding to amino acid residues 659-664, 676-684, or 702-706 of the AaCas12b protein, numbering based on SEQ ID NO: 1.

[0135] In some embodiments, the parent Cas12b protein is truncated about 20 or less (e.g., about any of 18, 16, 14, 12, 10, 8, 7, 6, 5, 4, 3, 2, or 1) amino acid residues from the amino acid residue corresponding to amino acid residue 518 of the AaCas12b protein, as numbered based on SEQ ID NO:1. ...658 of the AaCas12b protein, as numbered based on SEQ ID NO:1. In some embodiments, the parent Cas12b protein is truncated at an amino acid residue corresponding to amino acid residue 658 of the AaCas12b protein, where the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is truncated about 20 or less (e.g., about any of 18, 16, 14, 12, 10, 8, 7, 6, 5, 4, 3, 2, or 1) amino acid residues from an amino acid residue corresponding to amino acid residue 783 of the AaCas12b protein, where the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is truncated at an amino acid residue corresponding to amino acid residue 783 of the AaCas12b protein, where the numbering is based on SEQ ID NO: 1.

[0136] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises the WED-I, REC1, WED-II, RuvC-I, and BH domains of the AaCas12b protein, and the C-terminal portion of the parent Cas12b protein comprises the REC2, RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the AaCas12b protein. In some embodiments, the N-terminal portion of the parent Cas12b protein comprises amino acid residues 1 to 658 of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises amino acid residues 659 to 1129 of the parent Cas12b protein, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0137] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises the WED-I, REC1, WED-II, RuvC-I, BH and REC2 domains of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises the RuvC-II, Nuc-I, RuvC-III and Nuc-II domains of the parent Cas12b protein. In some embodiments, the N-terminal portion of the parent Cas12b protein comprises amino acid residues 1 to 783 of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises amino acid residues 784 to 1129 of the parent Cas12b protein, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0138] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises the WED-I, REC1, WED-II, RuvC-I, and BH domains of the parent Cas12b protein, the C-terminal portion of the parent Cas12b protein comprises the RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the parent Cas12b protein, and the REC2 domain of the parent Cas12b protein is truncated between the N-terminal portion of the parent Cas12b protein and the C-terminal portion of the parent Cas12b protein.

[0139] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises the WED-I, REC1, and WED-II domains of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises the RuvC-I, BH, REC2, RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the parent Cas12b protein. In some embodiments, the N-terminal portion of the parent Cas12b protein comprises amino acid residues 1 to 518 of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises amino acid residues 519 to 1129 of the parent Cas12b protein, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0140] The breakpoints are usually designed on a computer and cloned into the construct. The two truncated Cas12b portions, i.e., the N-terminal portion and the C-terminal portion, combine to form a functional Cas12b protein, preferably comprising at least about 70% or more of the amino acid sequence of the parent Cas12b protein, e.g., at least about 75%, 80%, 85%, 90%, 95%, 98%, 99% or more of the amino acid sequence of the parent Cas12b protein. Some trimming and mutation is expected. Non-functional domains can be completely deleted. For all truncated Cas12b systems, the two truncated Cas12b portions can be combined together to restore or reconstitute the desired Cas12b function. The activity of the reconstituted Cas12b protein or CRISPR complex (Cas12b + guide RNA complex) can be evaluated using methods known in the art. For example, the T7 endonuclease I (T7EI) assay can be used to evaluate nuclease activity in cells. Gene editing activity can also be evaluated by DNA sequencing.

[0141] In some embodiments, the parent Cas12b protein is cleaved into two or more parts, such as 3, 4, 5, or 6 parts.

[0142] The truncated Cas12b effector proteins may each comprise one or more dimerization domains. In some embodiments, the first polypeptide comprises a first dimerization domain fused to a first truncated Cas12b effector moiety and the second polypeptide comprises a second dimerization domain fused to a second truncated Cas12b effector moiety. The dimerization domains may be fused to the truncated Cas12b effector moiety via a peptide linker (e.g., a flexible peptide linker such as a GS linker) or a chemical bond. In some embodiments, the dimerization domain is fused to the N-terminus of the truncated Cas12b effector moiety. In some embodiments, the dimerization domain is fused to the C-terminus of the truncated Cas12b effector moiety.

[0143] In some embodiments, the truncated Cas12b effector protein does not comprise a dimerization domain.

[0144] In some embodiments, the dimerization domain promotes association of two truncated Cas12b effector moieties. In some embodiments, the truncated Cas12b effector moieties are induced to associate or dimerize into a functional Cas12b effector protein by an inducer agent. In some embodiments, the truncated Cas12b effector protein comprises an inducible dimerization domain. In some embodiments, the dimerization domain is not an inducible dimerization domain, i.e., the dimerization domain dimerizes without the presence of an inducer agent.

[0145] The inducer may be an induction energy source or an inducer molecule other than a guide RNA (e.g., sgRNA). The inducer acts to reconstitute two cleaved Cas12b effector moieties into a functional Cas12b effector protein by inducing dimerization of the dimerization domain. In some embodiments, the inducer joins two cleaved Cas12b effector moieties by inducible association of the dimerization domain. In some embodiments, in the absence of the inducer, the two cleaved Cas12b effector moieties do not associate with each other to reconstitute as a functional Cas12b effector protein. In some embodiments, in the absence of the inducer, the two cleaved Cas12b effector moieties can be reconstituted into a functional Cas12b effector protein in the presence of a guide RNA (e.g., sgRNA).

[0146] The inducing agent of the present application may be heat, ultrasound, electromagnetic energy, or a chemical compound. In some embodiments, the inducing agent is an antibiotic, a small molecule, a hormone, a hormone derivative, a steroid, or a steroid derivative. In some embodiments, the inducing agent is abscisic acid (ABA), doxycycline (DOX), cumate, rapamycin, 4-hydroxytamoxifen (4OHT), estrogen, or ecdysone. In some embodiments, the truncated Cas12b effector system is an inducer-controlled system selected from the group consisting of an antibiotic-based induction system, an electromagnetic energy-based induction system, a small molecule-based induction system, a nuclear receptor-based induction system, and a hormone-based induction system. In some embodiments, the truncated Cas12b effector system is an inducer-controlled system selected from the group consisting of a tetracycline (Tet) / DOX induction system, a light induction system, an ABA induction system, a cumate repressor / operon system, a 4OHT / estrogen induction system, an ecdysone-based induction system, and an FKBP12 / FRAP (FKBP12-rapamycin complex) induction system. Such inducers are also discussed herein and in PCT / US2013 / 051418, the contents of which are incorporated herein by reference in their entireties. The FRB / FKBP / rapamycin system is described in Paulmurugan and Gambhir, Cancer Res, August 15, 2005 65; 7413; and Crabtree et al., Chemistry & Biology 13, 99-107, Jan 2006, the contents of each of which are incorporated herein by reference in their entireties.

[0147] In some embodiments, the cleaved Cas12b effector protein pair is separated and inactive until dimerization of the dimerization domains (e.g., FRB and FKBP) is induced, thereby leading to reassembly of a functional Cas12b effector nuclease. In some embodiments, a first cleaved Cas12b effector protein comprising a first half of an inducible dimer (e.g., FRB) is delivered and / or localized separately from a second cleaved Cas12b effector protein comprising a second half of an inducible dimer (e.g., FKBP).

[0148] Other exemplary FKBP-based inducible systems that can be used in the inducible agent-controlled split Cas12b effector system described herein include, but are not limited to, an FKBP that dimerizes with calcinillin A (CNA) in the presence of FK506, an FKBP that dimerizes with CyP-Fas in the presence of FKCsA, an FKBP that dimerizes with FRB in the presence of rapamycin, GyrB that dimerizes with GryB in the presence of coumamycin, GAI that dimerizes with GID1 in the presence of gibberellin, and Snap-tag that dimerizes with HaloTag in the presence of HaXS.

[0149] Alternatives to the FKBP family itself are also possible: for example, in the presence of FK1012, FKBPs undergo homodimerization (i.e., one FKBP dimerizes with another FKBP).

[0150] In some embodiments, the dimerization domain is FKBP and the inducer is FK1012. In some embodiments, the dimerization domain is GryB and the inducer is coumamycin. In some embodiments, the dimerization domain is ABA and the inducer is gibberellin.

[0151] In some embodiments, the truncated Cas12b effector moiety can be auto-induced (i.e., auto-activated or self-induced) to associate / dimerize into a functional Cas12b effector protein without the presence of an inducer. Without being bound by any theory or hypothesis, auto-induction of the truncated Cas12b effector moiety can be mediated by binding to a guide RNA (e.g., sgRNA). In some embodiments, the first polypeptide and the second polypeptide do not comprise a dimerization domain. In some embodiments, the first polypeptide and the second polypeptide comprise a dimerization domain.

[0152] In some embodiments, a reconstituted Cas12b effector protein of a truncated Cas12b effector system described herein (including inducer-controlled and autoinducible systems) has an editing efficiency that is at least about 70% (e.g., at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or more, or 100%) of the editing efficiency of the parent Cas12b effector protein.

[0153] In some embodiments, a reconstituted Cas12b effector protein of an inducer-controlled truncated Cas12b effector system described herein has about 50% or less of the editing efficiency of the parent Cas12b effector protein in the absence of an inducer agent (i.e., by auto-induction) (e.g., any of about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less efficiency, or 0% efficiency).

[0154] Fusion Cas12b effector protein The present application also provides engineered Cas12b effector proteins that contain additional protein domains and / or components, such as linkers, nuclear localization / export sequences, functional domains and / or reporter proteins.

[0155] In some embodiments, an engineered Cas12b effector protein is a protein complex that includes one or more heterologous protein domains (e.g., about any of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains) in addition to the nucleic acid targeting domain of an engineered Cas12b nuclease, variant or functional derivative thereof. In some embodiments, an engineered Cas12b effector protein is a fusion protein that includes one or more heterologous protein domains (e.g., about any of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains) fused to an engineered Cas12b nuclease, variant or functional derivative thereof.

[0156] In some embodiments, the engineered Cas12b effector protein of the present application may comprise one or more functional domains (e.g., via a fusion protein, such as one or more peptide linkers, e.g., GS peptide linkers) or may be associated with one or more functional domains (e.g., via co-expression of multiple proteins). In some embodiments, the one or more functional domains are enzymatic domains. The functional domains may have various activities, such as, for example, DNA and / or RNA methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and switching activity (e.g., light induction). In some embodiments, the one or more functional domains are transactivation domains (i.e., transactivation domains) or repressor domains. In some embodiments, the transactivation domain or repressor domain can recruit chromatin modifying agents. In some embodiments, the one or more functional domains are histone modification domains. In some embodiments, the one or more functional domains are a transposase domain, a HR (homologous recombination) mechanical domain, a recombinase domain, and / or an integrase domain. In some embodiments, the functional domain is Kruppel-associated box (KRAB), VP64, VP16, Fok1, P65, HSF1, MyoD1, biotin-APEX, APOBEC1, AID, PmCDA1, Tad1, and M-MLV reverse transcriptase. In some embodiments, the functional domain is selected from the group consisting of a translation promoter domain, a transcription repressor domain, a transactivation domain, an epigenetic modification domain, a nucleobase editing domain (e.g., a CBE or ABE domain), a reverse transcriptase domain, a reporter domain (e.g., a fluorescent domain), and a nuclease domain. In some embodiments, the functional domain is a KRAB domain, such as the KRAB domain of ZIM3. In some embodiments, the KRAB domain comprises the amino acid sequence of SEQ ID NO:72.

[0157] In some embodiments, the localization of one or more functional domains in an engineered Cas12b effector protein allows for the correct spatial orientation of the functional domain to affect a target with the ascribed functional effect. For example, if the functional domain is a transcriptional activator (e.g., VP16, VP64, or p65), the transcriptional activator is placed in a spatial orientation that allows it to affect target transcription. Similarly, a transcriptional repressor is localized to affect target transcription, and a nuclease (e.g., Fok1) is localized to cleave or partially cleave a target. In some embodiments, a functional domain (e.g., a KRAB domain, such as including SEQ ID NO:72) is located at the N-terminus of an engineered Cas12b effector protein (e.g., any of SEQ ID NOs:79-81, such as SEQ ID NO:81). In some embodiments, a functional domain (e.g., a KRAB domain, such as SEQ ID NO:72) is located at the C-terminus of an engineered Cas12b effector protein (e.g., any of SEQ ID NOs:79-81, such as SEQ ID NO:81). In some embodiments, an engineered Cas12b effector protein comprises a first functional domain at the N-terminus and a second functional domain at the C-terminus. In some embodiments, an engineered Cas12b effector protein comprises a catalytically inactive mutant of any of the engineered Cas12b nucleases described herein (e.g., any of SEQ ID NOs:79-81) fused to one or more functional domains (e.g., a KRAB domain).

[0158] In some embodiments, the engineered Cas12b effector protein is a transcriptional activator. In some embodiments, the engineered Cas12b effector protein comprises any of the enzymatically inactivated variants of the engineered Cas12b nucleases described herein (e.g., any of SEQ ID NOs: 79-81) fused to a transactivation domain. In some embodiments, the transactivation domain is selected from the group consisting of VP64, p65, HSF1, VP16, MyoD1, HSF1, RTA, SET7 / 9, and combinations thereof. In some embodiments, the transactivation domain comprises VP64, p65, and HSF1. In some embodiments, the engineered Cas12b effector protein comprises two truncated Cas12b effector polypeptides, each fused to a transactivation domain. In some embodiments, the engineered Cas12b effector protein further comprises one or more nuclear localization sequences (e.g., any of SEQ ID NOs: 61, 62, and 82).

[0159] In some embodiments, the engineered Cas12b effector protein is a transcriptional repressor. In some embodiments, the engineered Cas12b effector protein comprises any of the enzymatically inactivated variants of engineered Cas12b nucleases described herein (e.g., any of SEQ ID NOs: 79-81) fused to a transcriptional repressor domain (e.g., KRAB). In some embodiments, the transcriptional repressor domain is selected from the group consisting of Kruppel-associated box (KRAB), EnR, NuE, NcoR, SID, SID4X, and combinations thereof. In some embodiments, the engineered Cas12b effector protein comprises two truncated Cas12b effector polypeptides, each fused to a transcriptional repressor domain. In some embodiments, the engineered Cas12b effector protein further comprises one or more nuclear localization sequences (e.g., any of SEQ ID NOs: 61, 62, and 82).

[0160] In some embodiments, the engineered Cas12b effector protein is a base editor, such as a cytosine editor or an adenosine editor. In some embodiments, the engineered Cas12b effector protein comprises any of the enzymatically inactivated variants of the engineered Cas12b nuclease described herein (e.g., any of SEQ ID NOs:79-81) fused to a nucleobase editing domain (e.g., a cytosine base editing (CBE) domain or an adenosine base editing (ABE) domain). In some embodiments, the nucleobase editing domain is a domain for editing DNA. In some embodiments, the nucleobase editing domain has deaminase activity. In some embodiments, the nucleobase editing domain is a cytosine deaminase domain. In some embodiments, the nucleobase editing domain is an adenosine deaminase domain. An example of a Cas nuclease-based base editor is described, for example, in WO2018 / 165629A1 and WO2019 / 226953A1, the contents of which are incorporated herein by reference in their entirety. Exemplary CBE domains include, but are not limited to, activation-induced cytosine deaminase or AID (e.g., hAID), apolipoprotein B mRNA editing complex or APOBEC (e.g., rat APOBEC1, hAPOBEC3 A / B / C / D / E / F / G), and PmCDA1. Exemplary ABE domains include, but are not limited to, TadA, ABE8, and variants thereof (e.g., Gaudelli et al., 2017, Nature 551: 464-471; and Richter et al., 2020, Nature Biotechnology 38: 883-891; the contents of each of which are incorporated herein by reference in their entirety). In some embodiments, the functional domain is an APOBEC1 domain, such as a rat APOBEC1 domain. In some embodiments, the functional domain is a TadA domain.In some embodiments, the engineered Cas12b effector protein further comprises one or more nuclear localization sequences (e.g., any of SEQ ID NOs: 61, 62, and 82).

[0161] In some embodiments, the engineered Cas12b effector protein is a prime editor. Prime editors based on Cas9 are described, for example, in A. Anzalone et al., Nature, 2019, 576 (7785): 149-157, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the engineered Cas12b effector protein comprises any of the engineered Cas12b nuclease nickase variants described herein fused to a reverse transcriptase domain. In some embodiments, the functional domain is a reverse transcriptase domain. In some embodiments, the reverse transcriptase domain is M-MLV reverse transcriptase, or a variant thereof, such as M-MLV reverse transcriptase having one or more of the following mutations: D200N, T306K, W313F, T330P, and L603W. In some embodiments, an engineered CRISPR / Cas12b system is provided that comprises a prime editor. In some embodiments, the engineered CRISPR / Cas12b system further comprises a second Cas12b nickase, e.g., based on the same engineered Cas12b nuclease as the prime editor. In some embodiments, the engineered CRISPR / Cas12b system comprises a prime editor guide RNA (pegRNA) that includes a primer binding site and a reverse transcriptase (RT) template sequence.

[0162] In some embodiments, the present application provides a truncated Cas12b effector system having one or more functional domains (e.g., 1, 2, 3, 4, 5, 6 or more) associated (i.e., bound or fused) to one or two truncated Cas12b effector moieties. The functional domains may be provided as fusions within the construct as part of the first and / or second truncated Cas12b effector protein. The functional domains are typically fused to other parts of the truncated Cas12b effector protein (e.g., truncated Cas12b effector moieties) via a peptide linker (e.g., GS linker). This functional domain can be used to retune the function of the truncated Cas12b effector system based on a catalytically inactivated Cas12b effector.

[0163] In some embodiments, the engineered Cas12b effector protein comprises one or more nuclear localization sequences (NLS) and / or one or more nuclear export sequences (NES). Exemplary NLS sequences include, for example, PKKKRKV (SEQ ID NO:82), PKKKRKVPG (SEQ ID NO:61), and ASPKKKRKV (SEQ ID NO:62). The NLS and / or NES are operably linked to the N-terminus and / or C-terminus of the engineered Cas12b effector protein or the polypeptide chain in the engineered Cas12b effector protein.

[0164] In some embodiments, the engineered Cas12b effector protein can encode additional components such as reporter proteins. In some embodiments, the engineered Cas12b effector protein includes a fluorescent protein such as GFP. Such a system can allow imaging of genomic sites (see, for example, "Dynamic Imaging of Genomic Loci in Living Human Cells by an Optimized CRISPR / Cas System" Chen B et al. Cell 2013). In some embodiments, the engineered Cas12b effector protein is an inducible cleavage Cas effector system that can be used for imaging of genomic sites.

[0165] Engineered CRISPR-Cas12b system Also provided is an engineered CRISPR-Cas12b system comprising: (a) any of the engineered Cas12b nucleases described herein, variants or derivatives thereof (e.g., any of SEQ ID NOs: 2-22 and 79-81), engineered Cas12b effector proteins (e.g., engineered Cas12b nucleases, nickases, truncated Cas12b proteins, transcriptional repressors, transcriptional activators, base editors, or prime editors), or nucleic acids encoding same; and (b) a guide RNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or one or more nucleic acids encoding the guide RNA, wherein the engineered Cas12b nuclease or engineered Cas12b effector protein and the guide RNA are capable of specifically binding to a target nucleic acid comprising the target sequence and forming a CRISPR complex that induces modification of the target nucleic acid.In some embodiments, the present invention provides an engineered Cas12b nuclease or its effector protein that comprises one, two or three types of mutations relative to a reference Cas12b nuclease, the mutations being (1) a substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K), and / or (2) a substitution of one or more amino acid residues that are involved in DNA duplex cleavage in the Cas12b nuclease (e.g., one or more of positions 118 and 119). (3) substitution of one or more amino acid residues that interact with ssDNA substrates in the RuvC domain of a reference Cas12b nuclease (e.g., 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 860, 862, 864, 866, 868, 870, 872, 874, 876, 878, 879, 880, 881, 882, 883, 884, 885, 886, 887, 888, 889, 890, 900, 901, 902, 903, 904, 905, 906, 907, 908, 909, 910, 911, 912, 913, 914, 915, 916, 917, 918, 919, 920, 921, 922, 923, 924, 925, 926, 927, 928, 929, 930, 931, 932, 933, 934, 935, 936, 937, 938, 939, 940, 941, 942, 943, 944, 945, 946, 947, 950, 951, 952, 953, 954, 955, 956, 957, 9 56, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease has the structure of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, Provided is an engineered CRISPR-Cas12b system comprising: (a) an engineered Cas12b nuclease or its effector protein, the engineered Cas12b nuclease or its effector protein comprising the amino acid sequence of ID NO:1 or a nucleic acid encoding the engineered Cas12b nuclease or its effector protein; and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA, wherein the engineered Cas12b nuclease or its effector protein and the gRNA are capable of specifically binding to a target nucleic acid comprising the target sequence and forming a CRISPR complex that induces modification of the target nucleic acid.In some embodiments, the engineered CRISPR-Cas12b system comprises one or more nucleic acids encoding an engineered Cas12b nuclease, a variant or derivative thereof, or an engineered Cas12b effector protein and / or a guide RNA. In some embodiments, the gRNA comprises a crRNA and a tracrRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor guide RNA array that can be processed into a plurality of crRNAs, for example, by an engineered Cas12b nuclease, a variant or derivative thereof, or an engineered Cas12b effector protein. In some embodiments, the gRNA is an sgRNA. In some embodiments, the sgRNA comprises a scaffold sequence of any of SEQ ID NOs: 23-53. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding an engineered Cas12b nuclease, a variant or derivative thereof, or an engineered Cas12b effector protein and / or a guide RNA. In some embodiments, the engineered Cas12b nuclease, variants or derivatives thereof, or engineered Cas12b effector proteins and / or guide RNAs are encoded by one or more vectors (e.g., adeno-associated virus (AAV) vectors). In some embodiments, the engineered CRISPR-Cas12b system comprises a ribonucleoprotein (RNP) complex comprising the engineered Cas12b nuclease, variants or derivatives thereof, or engineered Cas12b effector proteins that bind to the guide RNA.

[0166] In some embodiments, the method comprises the steps of: (a) a Cas12b nuclease or an effector protein thereof (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor) comprising the amino acid sequence of SEQ ID NO:1, any of the engineered Cas12b nucleases, variants or derivatives thereof described herein (e.g., any of SEQ ID NOs:2-22 and SEQ ID NOs:79-81), engineered Cas12b effector proteins (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor), or nucleic acids encoding same; and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding a gRNA, wherein the gRNA comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, an engineered CRISPR-Cas12b system is provided, comprising an engineered scaffold comprising any of the sequences of SEQ ID NO:25-53, wherein i) the Cas12b nuclease or its effector protein, an engineered Cas12b nuclease, variant or derivative thereof, or an engineered Cas12b effector protein, and ii) the gRNA form a CRISPR complex that specifically binds to a target nucleic acid and can induce modification of the target nucleic acid.A Cas12b nuclease or effector protein thereof (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor) comprising the amino acid sequence of NO:1, or an engineered Cas12b nuclease or effector protein thereof, comprising one, two, or three types of mutations relative to a reference Cas12b nuclease, the mutations including (1) substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K), and / or (2) substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K), and / or (3) substitution of one or more amino acid residues that interact with DNA duplexes in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K). and / or (3) substitution of one or more amino acid residues involved in cleavage of the RuvC domain of a reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W) (e.g., one or more of positions 118 and 119). and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA, wherein the gRNA comprises one or more of the amino acid sequences of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8,An engineered CRISPR-Cas12b system is provided, which comprises an engineered scaffold having any of the sequences of NO:25 to 53, and wherein i) the Cas12b nuclease or its effector protein, or an engineered Cas12b nuclease or its effector protein, and ii) the gRNA form a CRISPR complex that specifically binds to a target nucleic acid and can induce modification of the target nucleic acid. In some embodiments, an engineered CRISPR-Cas12b system is provided, comprising: (a) a Cas12b nuclease or a Cas12b effector protein, or a nucleic acid encoding the same, comprising any of the amino acid sequences of SEQ ID NOs: 1-22 and SEQ ID NOs: 79-81; and (b) a gRNA, or a nucleic acid encoding the gRNA, comprising a guide sequence complementary to a target sequence of a target nucleic acid, wherein the gRNA comprises an engineered scaffold comprising any of the sequences of SEQ ID NOs: 25-53, wherein the Cas12b nuclease or the Cas12b effector protein and the gRNA form a CRISPR complex that specifically binds to the target nucleic acid and can induce modification of the target nucleic acid. In some embodiments, the gRNA comprises a crRNA and a tracrRNA, wherein the tracrRNA comprises an engineered scaffold or a portion thereof. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor gRNA array encoding a plurality of crRNAs. In some embodiments, the gRNA is an sgRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises an engineered Cas12b nuclease, an engineered Cas12b effector protein, or one or more vectors encoding a Cas12b nuclease or a Cas12b effector protein. In some embodiments, the one or more vectors are AAV vectors. In some embodiments, the one or more vectors further encode a gRNA.

[0167] PAM In some embodiments, the engineered Cas12b nuclease, variant or derivative thereof, engineered Cas12b effector protein, Cas12b nuclease or Cas12b effector protein Cas12b recognizes a PAM that comprises (or consists of) a 5'-TTN-3' sequence, where N is A, T, G, or C. In some embodiments, the PAM comprises or consists of 5'-TTC-3', 5'-TTA-3', 5'-TTT-3', or 5'-TTG-3'.

[0168] Guide RNA The engineered CRISPR-Cas12b system of the present application may comprise any suitable guide RNA. A guide RNA (gRNA) may comprise a guide sequence (or spacer) capable of hybridizing to a target sequence in a target nucleic acid of interest (e.g., a genomic site of interest in a cell). In some embodiments, the gRNA comprises a CRISPR RNA (crRNA) sequence comprising a guide sequence. In some embodiments, the crRNA described herein comprises a direct repeat (DR) sequence and a spacer sequence. In some embodiments, the crRNA comprises (consists essentially of, or consists of) a direct repeat sequence linked to a guide sequence or a spacer sequence. In some embodiments, the direct repeat sequence may be located upstream (i.e., 5') of the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3') of the guide sequence or spacer sequence. In some embodiments, the crRNA comprises a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-spacer-DR), which is a typical precursor crRNA (pre-crRNA) configuration. In some embodiments, the crRNA comprises a truncated direct repeat and a spacer sequence, which is a typical processed or mature crRNA. In some embodiments, the crRNA comprises a mutated DR sequence and a spacer sequence. In some embodiments, the gRNA comprises a trans-activating CRISPR RNA (tracrRNA) sequence. In some embodiments, the tracrRNA is fused to the crRNA at the 5' end of the DR sequence. In some embodiments, the guideRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA or sgRNA comprises a tracrRNA and a crRNA. In some embodiments, the sgRNA comprises any of the sequences of SEQ ID NOs: 23-53. In some embodiments, the tracrRNA comprises any of the sequences of SEQ ID NOs: 23-53 or a portion thereof.

[0169] In some embodiments, the gRNA comprises a non-homologous crRNA and / or tracrRNA sequence that does not naturally occur at the CRISPR locus of the reference Cas12b protein. For example, the homologous tracrRNA and crRNA sequences of AaCas12b, AkCas12b, AmCas12b, BhCas12b, BsCas12b, Bs3Cas12b, LsCas12b, and SbCas12b, as well as exemplary sgRNA sequences, are described in Figures S4 and S8 of Teng F. et al., Cell Discovery (2019) 5: 23, the contents of which are incorporated herein by reference in their entirety.

[0170] In some embodiments, the CRISPR-Cas12b system described herein comprises one or more gRNAs (e.g., crRNA, tracrRNA, or sgRNA) (e.g., 1, 2, 3, 4, 5, 10, 15, or more), or nucleic acids encoding them. In some embodiments, the two or more gRNAs target different target sites, e.g., two target sites in the same target DNA or gene, or two target sites in two different target DNAs or genes.

[0171] The sequence and length of the gRNA described herein may be optimized. In some embodiments, the optimal length of the gRNA can be determined by knowing the processed form of the crRNA or by empirical length survey of the crRNA. In some embodiments, the gRNA contains base modifications, such as the gRNA scaffold region.

[0172] The spacer does not need to be perfectly complementary, as long as it has sufficient complementarity for the gRNA (e.g., crRNA or sgRNA) to function (i.e., direct the Cas12b nuclease (e.g., engineered) or its effector protein to the target site). The gRNA-mediated editing or cleavage efficiency of the Cas12b nuclease (e.g., engineered) or its effector protein can be adjusted by introducing one or more mismatches (e.g., one or two mismatches between the spacer sequence and the target sequence, including positions along the spacer / target sequence mismatches). If the mismatch (e.g., double mismatch) is located in a more central position of the spacer (i.e., not at the 3' or 5' end of the spacer), it will have a large impact on the cleavage efficiency. Thus, by selecting the position of the mismatch on the spacer sequence, the editing or cleavage efficiency of the Cas12b nuclease (e.g., engineered) or its effector protein can be adjusted. For example, if it is desired to edit or cleave the target sequence in less than 100% (e.g., in a population of cells), one or two mismatches between the spacer sequence and the target sequence can be introduced into the spacer sequence.

[0173] In some embodiments, the guide sequence or spacer is designed to have at least one mismatch with the target sequence, and the heteroduplex formed between the guide sequence and the target sequence includes an unpaired C in the guide sequence opposite a target A, or an unpaired A in the guide sequence opposite a target C, to effect deamination (e.g., base editing) on ​​the target sequence. In some embodiments, in addition to this AC or CA mismatch, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using a suitable alignment algorithm.

[0174] The guide sequence may have any suitable length. In some embodiments, the length of the guide sequence or spacer sequence is about 10 nt to about 50 nt. In some embodiments, the length of the guide sequence or spacer sequence is at least about 16 nucleotides, preferably about 16 to about 100 nucleotides, more preferably about 16 to about 50 nucleotides (e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides). In some embodiments, the spacer is about 16 to about 27 nucleotides, for example, about 17 to about 24 nucleotides, about 18 to about 24 nucleotides, or about 18 to about 22 nucleotides. In some embodiments, the guide sequence is about 18 to about 35 nucleotides, including, for example, any of 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides.

[0175] In some embodiments, the guide sequence or spacer sequence is at least about 60% (e.g., at least about any of 70%, 75%, 80%, 85%, 90%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) complementary to the target sequence. In some embodiments, there are at least about 15 (e.g., at least about any of 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or more) base pairs between the spacer sequence and the target sequence of the target nucleic acid (e.g., DNA).

[0176] Optimal alignment can be determined using any suitable sequence alignment algorithm, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows-Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of the guide sequence (within the nucleic acid-targeting guide RNA) to guide sequence-specific binding of a nucleic acid-targeting complex to a target nucleic acid sequence can be assessed by any suitable assay. For example, sufficient nucleic acid-targeting CRISPR system components (including a test guide sequence) to form a nucleic acid-targeting complex can be provided to a host cell having a corresponding target nucleic acid sequence by transfecting with a vector encoding the nucleic acid-targeting complex components, and then evaluating preferential targeting (e.g., cleavage) within the target nucleic acid sequence by the Surveyor assay described herein. Similarly, cleavage of a target nucleic acid sequence can be evaluated in a test tube by providing the target nucleic acid sequence and components of the nucleic acid-targeting complex (including a test guide sequence and a control guide sequence different from the test guide sequence) and comparing the binding or cleavage rate at the target sequence between the test guide sequence and the control guide sequence reactions. Other measurements are possible and can be performed by one of skill in the art.

[0177] As used herein, a target nucleic acid is used interchangeably with a target sequence or a target nucleic acid sequence and refers to a specific nucleic acid that includes a nucleic acid sequence that is complementary to all or a portion of a spacer in a crRNA or a gRNA. In some examples, a target nucleic acid includes a gene or a sequence within a gene. In some examples, a target nucleic acid includes a non-coding region (e.g., a promoter). In some examples, a target nucleic acid is single-stranded. In some examples, a target nucleic acid is double-stranded. A target nucleic acid can be selected to target any target nucleic acid sequence, such as a DNA or RNA sequence (e.g., mRNA).

[0178] The target nucleic acid should be associated with PAM, i.e., the short sequence recognized by CRISPR complex.Depending on the nature of CRISPR-Cas protein, the target sequence should be selected so that the complementary sequence (the complementary sequence of the target sequence) in the DNA double strand is located upstream or downstream of the PAM.In one embodiment of the present application, the complementary sequence of the target sequence is located downstream or 3' of the PAM.The exact sequence and length requirement of the PAM depends on the Cas12b protein used.

[0179] The term "tracrRNA" or similar term includes any polynucleotide sequence that has sufficient complementarity with the crRNA sequence to hybridize. In some embodiments, when optimally aligned, the degree of complementarity between the tracrRNA sequence and the crRNA sequence is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more along the length of the shorter sequence. In some embodiments, the length of the tracr sequence is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more nucleotides. In some embodiments, the tracr sequence and the crRNA sequence are contained within a single transcript, and hybridization between them produces a transcript with secondary structure (e.g., one or more hairpin structures). Generally, the degree of complementarity refers to the optimal alignment of the guide sequence and the tracr sequence along the length of the shorter of the two sequences, which can be determined by any suitable alignment algorithm and can further take into account secondary structure.

[0180] Any gRNA scaffold, tracrRNA or DR sequence capable of mediating binding of the Cas12b protein described herein to a corresponding gRNA (e.g., crRNA) can be used herein. In some embodiments, the gRNA scaffold, tracrRNA or DR sequence comprises a stem-loop structure near the 5' or 3' end (immediately adjacent to the spacer sequence). By "stem-loop structure" is meant a nucleic acid having a secondary structure that includes a region of nucleotides known or predicted to form a double-stranded (stem) portion, one end of which is substantially connected by a linking region (loop) of single-stranded nucleotides. The term "hairpin" structure is also used herein to refer to a stem-loop structure. Such structures are well known in the art, and the terms are used according to their generally known meaning in the art. A stem-loop structure does not require exact base pairing. Thus, the stem may contain one or more base mismatches. Alternatively, the base pairing may be exact, i.e., without mismatches.

[0181] In some embodiments, the gRNA scaffold, tracrRNA or DR is a "functional variant" of the wild-type scaffold, tracrRNA or DR, e.g., a "functionally truncated version," a "functionally extended version," or a "functionally replaced version." A "functional variant" of a gRNA scaffold, tracrRNA or DR is a 5' and / or 3' extended (functionally extended version) or truncated (functionally truncated version) variant of a reference scaffold, tracrRNA or DR (e.g., parent DR), or comprises one or more insertions, deletions, and / or substitutions of one or more nucleotides relative to a reference scaffold, tracrRNA or DR (e.g., parent DR) (functional replacement version), while maintaining at least about 20% (e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more) functionality of said reference scaffold, tracrRNA or DR (e.g., parent DR), i.e., the function of mediating binding of a Cas12b nuclease (e.g., engineered) or its effector protein to the corresponding sgRNA or crRNA. The gRNA scaffold, tracrRNA or DR functional variant typically retains a stem-loop-like secondary structure or portion thereof useful for binding to a Cas12b nuclease (e.g., engineered) or its effector protein. In some embodiments, the gRNA scaffold, tracrRNA or DR or functional variant thereof comprises at least two (e.g., 2, 3, 4, 5 or more) stem-loop-like secondary structures or portions thereof useful for binding to a Cas12b nuclease (e.g., engineered) or its effector protein.

[0182] In some embodiments, the DR or functional variant thereof comprises at least about 16 nucleotides (nt), e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 or more nucleotides. In some embodiments, the DR comprises about 20 nt to about 40 nt, e.g., about 20 nt to about 30 nt, about 22 nt to about 40 nt, about 23 nt to about 38 nt, about 23 nt to about 36 nt, or about 30 nt to about 40 nt. In some embodiments, the DR comprises 22 nt, 23 nt, or 24 nt. In some embodiments, the DR comprises 35 nt, 36 nt, or 37 nt. In some embodiments, the sgRNA scaffold or functional variant thereof comprises between about 50 nt and about 180 nt, e.g., between about 70 nt and about 140 nt, or between about 90 nt and about 120 nt.

[0183] In some embodiments, the sgRNA comprises a scaffold sequence comprising a stem-loop structure (e.g., 1, 2, 3, 4 or more stem-loops) located near the 5' end of the spacer sequence. In some embodiments, the stem comprises at least about 4 bp comprising complementary X and Y sequences, although stems of more (e.g., 5, 6, 7, 8, 9, 10, 11 or 12) or fewer (e.g., 3, 2) base pairs are also contemplated. Thus, for example, X2-10 and Y2-10 (X and Y represent any complementary group of nucleotides) are contemplated. In some embodiments, the stem composed of X and Y nucleotides, together with the loop, forms a perfect hairpin in the overall secondary structure, and this is advantageous, and the number of base pairs may be any number to form a perfect hairpin. In some embodiments, any complementary X:Y base pairing sequence (e.g., in terms of length) is permissible as long as the secondary structure of the entire guide molecule is maintained. In some embodiments, the loop connecting the stems composed of X:Y base pairs may be of any sequence of the same length (e.g., 4 or 5 nucleotides) or more, so long as it does not interrupt the overall secondary structure of the guide molecule. In some embodiments, the stems comprise about 5-7 bp containing complementary X and Y sequences, although stems of more or less base pairs are also contemplated. In some embodiments, non-Watson-Crick base pairing is contemplated, and such pairing typically preserves the stem-loop structure at that position. In some embodiments, the stems contained in the scaffold sequence comprise (e.g., are composed of) 5 pairs of complementary bases hybridized to each other, and the loop length is 6, 7, 8, or 9 nucleotides. In some embodiments, the stems may comprise at least 2, at least 3, at least 4, or at least 5 base pairs.In some embodiments, the stem loop structure comprises a first stem nucleotide strand having a length of 5 nucleotides, a second stem nucleotide strand having a length of 5 nucleotides, and a cyclic nucleotide strand disposed between the first stem nucleotide strand and the second stem nucleotide strand, wherein the first stem nucleotide strand and the second stem nucleotide strand may be hybridized to each other, and the cyclic nucleotide strand comprises 6, 7 or 8 nucleotides.

[0184] In some embodiments, the native hairpin or stem-loop structure of the guide molecule is extended or replaced by an extended stem-loop. In some cases, stem extension has been shown to enhance assembly of the guide molecule with CRISPR-Cas proteins (Chen et al. Cell. (2013); 155(7): 1479-1491). In some embodiments, the stem of the stem-loop is extended by at least 1, 2, 3, 4, 5 or more complementary base pairs (i.e., corresponding to the addition of 2, 4, 6, 8, 10 or more nucleotides to the guide molecule). In some embodiments, they are located at the end of the stem and adjacent to the loop of the stem-loop.

[0185] As used herein, the secondary structures of two or more sgRNAs or tracrRNAs are substantially identical or substantially not different means that the sgRNAs or tracrRNAs contain stems and / or loops that differ in length by no more than 1, 2, or 3 nucleotides. The nucleotide sequences of the sgRNAs or tracrRNAs differ by no more than 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides in terms of nucleotide type (A, U, G, or C) when compared by sequence alignment. In some embodiments, the secondary structures of two or more sgRNAs or tracrRNAs are substantially identical or substantially not different means that the sgRNAs or tracrRNAs contain up to one pair of stems that differ in complementary bases, and / or up to one loop that differs in nucleotide length, and / or contain stems that are identical in length but have base mismatches.

[0186] In some embodiments, a gRNA scaffold sequence capable of directing any engineered Cas12b effector protein of the present application to a target site comprises one or more nucleotide changes selected from the group consisting of nucleotide additions, insertions, deletions, and substitutions that do not result in a substantial difference in secondary structure compared to the scaffold sequence set forth in any of SEQ ID NOs: 23-53 or a functional truncated version thereof. In some embodiments, the gRNA scaffold comprises any of the sequences of SEQ ID NOs: 25-53, or a variant thereof comprising no more than about 10 nt (e.g., 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nt) differences.

[0187] In some embodiments, the guide RNA comprises a crRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises an array of precursor guide RNAs encoding multiple crRNAs. In some embodiments, the Cas12b effector protein cleaves the precursor guide RNA array to generate multiple crRNAs. In some embodiments, the engineered CRISPR-Cas12b system comprises an array of precursor guide RNAs encoding multiple crRNAs, each crRNA comprising a different guide sequence. In some embodiments, the crRNA encoded by the precursor guide RNA array is associated with a tracrRNA.

[0188] Constructs and Vectors The present application also provides constructs, vectors, and expression systems encoding any of the engineered Cas12b effector proteins (including engineered Cas12b nucleases) described herein. In some embodiments, the constructs, vectors, or expression systems further comprise one or more gRNAs (e.g., sgRNAs) or crRNA arrays.

[0189] A "vector" is a composition of matter that contains an isolated nucleic acid and is useful for delivering the isolated nucleic acid into a cell. A variety of vectors are known in the art, including, but not limited to, linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. In general, a suitable vector contains an origin of replication that functions in at least one living organism, a promoter sequence, convenient restriction enzyme sites, and one or more selectable markers. The term "vector" is also meant to include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as, for example, polylysine compounds, liposomes, etc.

[0190] In some embodiments, the vector is a viral vector. Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, lentivirus vectors, retrovirus vectors, vaccinia vectors, herpes simplex virus vectors, and derivatives thereof. In some embodiments, the vector is a phage vector. Viral vector technology is well known in the art and described, for example, in Sambrook et al. (2001, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York) and other virology and molecular biology manuals.

[0191] Many virus-based systems have been developed for gene transfer into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. Heterologous nucleic acids can be inserted into vectors and packaged into retroviral particles using techniques known in the art. Recombinant viruses can then be isolated and delivered to engineered mammalian cells in vitro or ex vivo. Several retroviral systems are known in the art. In some embodiments, adenoviral vectors are used. Various adenoviral vectors are known in the art. In some embodiments, lentiviral vectors are used. In some embodiments, self-inactivating lentiviral vectors are used.

[0192] In certain embodiments, the vector is an adeno-associated virus (AAV) vector, such as AAV2, AAV8, or AAV9, which may be administered in a single dose, said single dose comprising at least 1×10 5 particles (also called particle units, pu) of adenovirus or adeno-associated virus. In some embodiments, the dose is at least about 1 x 10 6 particles, at least about 1 x 10 7 particles, at least about 1 x 108 particles, or at least about 1×10 9 The adeno-associated virus is a single particle. Delivery methods and dosages are described, for example, in WO 2016205764 and U.S. Patent No. 8,454,972, the contents of which are incorporated herein by reference in their entirety.

[0193] In some embodiments, the vector is a recombinant adeno-associated virus (rAAV) vector.For example, in some embodiments, modified AAV vectors can be used for delivery.Modified AAV vectors can be based on one or more of several capsid types, including AAV1, AV2, AAV5, AAV6, AAV8, AAV8.2.AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotyped AAV (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6). Exemplary AAV vectors and techniques that can be used to generate rAAV particles are known in the art (see, e.g., Aponte-Ubillus et al. (2018) Appl. Microbiol. Biotechnol. 102(3): 1045-54; Zhong et al. (2012) J. Genet. Syndr. Gene Ther. S1: 008; West et al. (1987) Virology 160: 38-47 (1987); Tratschin et al. (1985) Mol. Cell. Biol. 5: 3251-60); U.S. Pat. Nos. 4,797,368 and 5,173,414, and International Publication Nos. WO 2015 / 054653 and WO 93 / 24641, each of which is incorporated herein by reference).

[0194] Any of the known AAV vectors for delivering Cas9 and other Cas12b proteins can be used to deliver the engineered Cas12b nucleases, effector proteins or systems of the present application.

[0195] Methods for introducing vectors into mammalian cells are known in the art. The vector can be transferred to the host cell by physical, chemical or biological methods.

[0196] Physical methods of introducing a vector into a host cell include calcium phosphate precipitation, lipofection, particle bombardment, microinjection, electroporation, etc. Methods for producing cells containing vectors and / or exogenous nucleic acids are well known in the art. See, for example, Sambrook et al. (2001) Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York. In some embodiments, the vector is introduced into a cell by electroporation.

[0197] Biological methods for introducing heterologous nucleic acid into host cells include the use of DNA and RNA vectors. Viral vectors have become the most widely used method for inserting genes into mammalian (e.g., human) cells.

[0198] Chemical methods for introducing vectors into host cells include colloidal dispersion systems such as macromolecular complexes, nanocapsules, microspheres, beads, lipid systems such as oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system used as an in vitro delivery vector is liposome (e.g., artificial membrane vesicles). In some embodiments, engineered CRISPR-Cas12b system is delivered as RNP in nanoparticles.

[0199] In some embodiments, vectors or expression systems encoding the CRISPR-Cas12b system or components thereof include one or more selectable or detectable markers that provide a method for isolating or efficiently selecting cells that contain and / or have been modified with the CRISPR-Cas12b system, e.g., at early stages and on a large scale.

[0200] Reporter genes can be used to identify potentially transfected cells and to evaluate the function of regulatory sequences. In general, a reporter gene is a gene that is not present or expressed in a recipient organism or tissue, and whose expression of the encoded polypeptide is reflected by a readily detectable property (e.g., enzymatic activity). Expression of the reporter gene is measured at a suitable time after the DNA is introduced into the recipient cells. Suitable reporter genes can include genes encoding luciferase, β-galactosidase, chloramphenicol acetyltransferase, secreted alkaline phosphatase, or the green fluorescent protein gene (e.g., Ui-Tei et al. FEBS Letters 479: 79-82 (2000)).

[0201] Other methods for determining the presence of heterologous nucleic acid in a host cell include, for example, molecular biology assays well known to those of skill in the art, such as Southern blotting, Northern blotting, RT-PCR, PCR; biochemical assays, such as detecting the presence or absence of specific peptides, such as immunological methods (ELISA and Western blot).

[0202] In some embodiments, the nucleic acid sequence encoding the engineered Cas12b nuclease or effector protein and / or guide RNA is operably linked to a promoter. In some embodiments, the promoter is an endogenous promoter to the cell engineered with the engineered CRISPR-Cas12b system. For example, the nucleic acid encoding the engineered Cas12b effector protein can be knocked into the genome of the engineered mammalian cell downstream of an endogenous promoter using any method known in the art. In some embodiments, the endogenous promoter is a promoter of an abundant protein (e.g., β-actin). In some embodiments, the endogenous promoter is an inducible promoter, e.g., induced by an endogenous activation signal of the engineered mammalian cell. In some embodiments, the engineered mammalian cell is a T cell and the promoter is a T cell activation-dependent promoter (e.g., an IL-2 promoter, an NFAT promoter, or an NFκB promoter).

[0203] In some embodiments, the promoter is a heterologous promoter for the engineered cell using the engineered CRISPR-Cas12b system.Various promoters have been studied for gene expression in mammalian cells, and any promoter known in the art can be used in the present application.Promoters can be broadly classified into constitutive promoters, or regulated promoters such as inducible promoters.

[0204] In some embodiments, the nucleic acid sequence encoding the engineered Cas12b effector protein and / or guide RNA is operably linked to a constitutive promoter. A constitutive promoter allows for constitutive expression of a heterologous gene (also called transgenic) in a host cell. Exemplary constitutive promoters contemplated herein include, but are not limited to, the cytomegalovirus (CMV) promoter, human elongation factor-1α (hEF1α), ubiquitin C promoter (UbiC), phosphoglycerol kinase promoter (PGK), simian virus 40 early promoter (SV40), and chicken β-actin promoter linked to the CMV early enhancer (CAG). In some embodiments, the promoter is a CAG promoter that includes the cytomegalovirus (CMV) early enhancer element, the promoter, the first exon and first intron of the chicken β-actin gene, and the splicing receptor of the rabbit β-globin gene.

[0205] In some embodiments, the nucleic acid sequence encoding the engineered CRISPR-Cas12b protein and / or guide RNA is operably linked to an inducible promoter. Inducible promoters belong to the category of regulated promoters. Inducible promoters can be induced by one or more conditions, such as physical conditions, microenvironment or physiological state of the host cell, inducers (i.e., induction drugs), or combinations thereof. In some embodiments, the induction conditions are selected from the group consisting of inducers, radiation (e.g., ionizing radiation, light), temperature (e.g., heat), redox conditions, tumor environment, and activation states of the cells engineered by the engineered CRISPR-Cas12b system. In some embodiments, the promoter can be induced by a small molecule inducer (e.g., a compound). In some embodiments, the small molecule is selected from the group consisting of doxycycline, tetracycline, alcohol, metals, or steroids. Chemically induced promoters are the most widely studied. Such promoters include promoters whose transcriptional activity is regulated by the presence or absence of small chemical molecules (e.g., doxycycline, tetracycline, alcohol, steroids, metals and other compounds). The doxycycline-inducible system with trans-tetracycline-controlled transactivator (rtTA) and tetracycline-responsive element promoter (TRE) is currently the most mature system. WO9429442 describes stringent control of gene expression in eukaryotic cells by tetracycline-responsive promoters. WO9601313 discloses transcriptional regulators that are regulated by tetracycline. Furthermore, Tet technology (e.g., the Tet-on system) is described on websites such as, for example, TetSystems.com. Known chemically regulated promoters can be used to drive expression of the engineered CRISPR-Cas12b proteins and / or guide RNAs of the present application.

[0206] In some embodiments, the nucleic acid sequence encoding the engineered Cas12b nuclease or effector protein is codon-optimized. In some embodiments, the expression construct encodes a tag (e.g., a 10xHis tag) operably linked to the C-terminus of the engineered Cas12b nuclease or effector protein. In some embodiments, each engineered truncated Cas12b construct encodes a fluorescent protein, such as GFP or RFP. A reporter protein can be used to assess colocalization and / or dimerization of the engineered truncated Cas12b proteins, for example, using microscopy. The nucleic acid sequence encoding the engineered Cas12b effector protein may be fused to a nucleic acid sequence encoding another component using a sequence encoding an autolytic peptide, such as a T2A, P2A, E2A, or F2A peptide.

[0207] In some embodiments, an expression construct for mammalian cells (e.g., human cells) is provided that comprises a nucleic acid sequence encoding an engineered Cas12b nuclease or effector protein. In some embodiments, the expression construct comprises a codon-optimized sequence encoding an engineered Cas12b nuclease or effector protein inserted into a pCAG-2A-eGFP vector, whereby the Cas12b protein is operably linked to eGFP. In some embodiments, a second vector is provided for expressing a guide RNA (e.g., a sgRNA, crRNA, or pre-crRNA array) in mammalian cells (e.g., human cells). In some embodiments, a sequence encoding a guide RNA is expressed within a pUC19-U6-Aa-sgRNA vector backbone.

[0208] In some embodiments, the nucleic acid encoding the Cas12b protein and the nucleic acid encoding the gRNA are on different vectors. In some embodiments, the nucleic acid encoding the Cas12b protein and the nucleic acid encoding the gRNA are on the same vector. In some embodiments, the nucleic acid encoding the Cas12b protein and the nucleic acid encoding the gRNA are under the control of different promoters, such as a CMV promoter and a U6 promoter. In some embodiments, the nucleic acid encoding the Cas12b protein is located upstream of the nucleic acid encoding the gRNA. In some embodiments, the nucleic acid encoding the Cas12b protein is located downstream of the nucleic acid encoding the gRNA. In some embodiments, the nucleic acid encoding the Cas12b protein and the nucleic acid encoding the gRNA are contacted with the target nucleic acid or are introduced into the cell simultaneously. In some embodiments, the nucleic acid encoding the Cas12b protein and the nucleic acid encoding the gRNA are contacted with the target nucleic acid or are introduced into the cell sequentially, such as introducing the nucleic acid encoding the Cas12b protein before the nucleic acid encoding the gRNA or introducing the nucleic acid encoding the Cas12b protein after the nucleic acid encoding the gRNA. In some embodiments, the cell already expresses the Cas12b protein. In some embodiments, only the nucleic acid encoding the gRNA is introduced into the cell. In some embodiments, the cell already expresses the gRNA, hi some embodiments, only the nucleic acid encoding the Cas12b protein is introduced into the cell.

[0209] III.How to use One aspect of the present application provides methods for detecting target or modified nucleic acids in vitro, ex vivo or in vivo using any of the engineered Cas12b nucleases, effector proteins or CRISPR-Cas12b systems described herein, as well as therapeutic or diagnostic methods using the engineered Cas12b nucleases, effector proteins or CRISPR-Cas12b systems. Also provided are the use of the engineered Cas12b nucleases, effector proteins or CRISPR-Cas12b systems described herein in detecting or modifying nucleic acids in cells and treating or diagnosing diseases or conditions in subjects, and the use of compositions comprising one or more components of the engineered Cas12b nucleases, effector proteins or engineered CRISPR-Cas12b systems in the manufacture of medicaments for detecting or modifying nucleic acids (e.g., in cells) and treating or diagnosing diseases or conditions in subjects.

[0210] Modification method In some embodiments, the present application provides a method for modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with any of the engineered CRISPR-Cas12b systems or components thereof described herein. For example, if a Cas12b protein or a nucleic acid encoding the same is already present, a gRNA or a nucleic acid encoding the same may be provided. If a gRNA or a nucleic acid encoding the same is already present, a Cas12b protein or a nucleic acid encoding the same may be provided. In some embodiments, the method includes contacting (e.g., in vitro, ex vivo, or in vivo) a target nucleic acid with a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring), the CRISPR-Cas12b system comprising (a) an engineered Cas12b nuclease or its effector protein (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor) that comprises one, two, or three types of mutations relative to a reference Cas12b nuclease, the mutations being at (1) one or more amino acid residues that interact with a PAM in the reference Cas12b nuclease (e.g., at one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475). and / or (2) substitution of one or more amino acid residues involved in DNA duplex cleavage in the reference Cas12b nuclease (e.g., one or more of positions 118 and 119) with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (3) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (4) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (5) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (6) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (7) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (8) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (9) substitution of one or more amino acid residues involved in DNA duplex cleavage in the RuvC domain of the reference Cas12b nuclease with amino acid residues having an aromatic ring (e.g., F, Y, W), A positively charged amino acid residue (e.g., R, H, or β-amino acid residue) at one or more of the acting amino acid residues (e.g., one or more of positions 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097)and (b) an engineered Cas12b nuclease or its effector protein, the reference Cas12b nuclease comprising an amino acid sequence of SEQ ID NO: 1 or a nucleic acid encoding the engineered Cas12b nuclease or its effector protein, and (b) a gRNA comprising a guide sequence complementary to a target sequence of the target nucleic acid, or a nucleic acid encoding the gRNA, wherein the guide sequence hybridizes to the target sequence of the target nucleic acid, thereby mediating contact of the engineered Cas12b nuclease or its effector protein with the target sequence of the target nucleic acid, thereby leading to modification of the target nucleic acid by the engineered Cas12b nuclease or its effector protein. In some embodiments, the gRNA comprises a scaffold comprising any of the sequences of SEQ ID NO: 23 and SEQ ID NOs: 25-53. In some embodiments, the engineered Cas12b nuclease or its effector protein comprises any of the amino acid sequences of SEQ ID NOs:2-22 and SEQ ID NOs:79-81. In some embodiments, the method includes contacting (e.g., in vitro, ex vivo, or in vivo) a target nucleic acid with a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring), comprising: (a) a Cas12b nuclease or effector protein thereof (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor) comprising the amino acid sequence of SEQ ID NO:1; or an engineered Cas12b nuclease or effector protein thereof, comprising one, two, or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations include: (1) one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease (e.g., 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 89and 475) with a positively charged amino acid residue (e.g., R, H, K), and / or (2) replacement of one or more amino acid residues involved in DNA duplex cleavage in the reference Cas12b nuclease (e.g., one or more of positions 118 and 119) with an amino acid residue having an aromatic ring (e.g., F, Y, W), and / or (3) replacement of one or more amino acid residues that interact with ssDNA substrates in the RuvC domain of the reference Cas12b nuclease (e.g., For example, the reference Cas12b nuclease may comprise a substitution of one or more of the following amino acid residues (e.g., 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), Methods are provided for modifying a target nucleic acid comprising a target sequence, the method comprising: (a) an engineered Cas12b nuclease or its effector protein, the engineered Cas12b nuclease or its effector protein comprising the amino acid sequence of ID NO:1, or a nucleic acid encoding the Cas12b nuclease (e.g., engineered) or its effector protein; and (b) a gRNA comprising a guide sequence complementary to a target sequence of the target nucleic acid, or a nucleic acid encoding the gRNA, the gRNA comprising an engineered scaffold comprising any of the sequences of SEQ ID NOs:25-53, the guide sequence hybridizing to the target sequence of the target nucleic acid mediates contact of the Cas12b nuclease (e.g., engineered) or its effector protein with the target sequence of the target nucleic acid, thereby leading to modification of the target nucleic acid by the Cas12b nuclease (e.g., engineered) or its effector protein. In some embodiments, the engineered Cas12b nuclease or its effector protein comprises any of the amino acid sequences of SEQ ID NOs:2-22 and SEQ ID NOs:79-81. In some embodiments, the method comprises:The method further includes providing a repair / donor template comprising a repair / donor nucleic acid, which can be incorporated (e.g., via homologous recombination) into the modified target nucleic acid at the target sequence. In some embodiments, the modification of the target nucleic acid repairs a mutation (e.g., a loss-of-function mutation) in the target nucleic acid to a wild-type (or harmless version) sequence. In some embodiments, the modification of the target nucleic acid introduces an exogenous sequence. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is present in a cell. In some embodiments, the cell is a bacterial cell, a yeast cell, a plant cell, or an animal cell (e.g., a mammalian cell, such as a human or mouse cell). In some embodiments, the method is performed ex vivo. In some embodiments, the method is performed in vivo. In some embodiments, the target nucleic acid is cleaved by the engineered CRISPR-Cas12b system, or a target sequence in the target nucleic acid is altered (e.g., base-edited) thereby. In some embodiments, the expression of the target nucleic acid is altered by the engineered CRISPR-Cas12b system. In some embodiments, the target nucleic acid is genomic DNA, such as in a cell. In some embodiments, the target sequence is associated with a disease or condition. In some embodiments, the method of modifying the target sequence treats a disease or condition associated with the target sequence. In some embodiments, the engineered CRISPR-Cas12b system comprises an array of precursor guide RNAs encoding multiple crRNAs, each crRNA comprising a different guide sequence.

[0211] In some embodiments, the present application provides a method of treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising modifying the target nucleic acid in the cell of the individual with any of the methods of modifying a target nucleic acid described herein to treat the disease or condition, in some embodiments, the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.

[0212] The engineered CRISPR-Cas12b system described herein can modify a target nucleic acid in a cell in a variety of ways, depending on the type of engineered Cas12b effector protein in the CRISPR-Cas12b system. In some embodiments, the method induces site-specific cleavage of the target nucleic acid. In some embodiments, the method cleaves genomic DNA in a cell, such as a bacterial cell, a plant cell, or an animal cell (e.g., a mammalian cell). In some embodiments, the method kills a cell by cleaving genomic DNA in the cell. In some embodiments, the method cleaves viral nucleic acid in the cell. In some embodiments, the method base edits the target nucleic acid, such as repairing a deleterious or disease-associated mutation to a non-disease-associated sequence. In some embodiments, the method enhances expression of the target nucleic acid (e.g., repairs a deleterious mutation that reduces expression) (e.g., at least about any of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 2-fold, 5-fold, 10-fold, 20-fold, or more). In some embodiments, the method reduces expression of a target nucleic acid (e.g., repairs a deleterious mutation that upregulates expression) (e.g., by at least about any of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 2-fold, 5-fold, 10-fold, 20-fold or more).

[0213] In some embodiments, the method alters (e.g., increases or decreases) the expression level of a target nucleic acid in a cell. In some embodiments, the method enhances the expression level of a target nucleic acid in a cell, for example, using an engineered Cas12b effector protein based on an enzymatically inactivated Cas12b protein (e.g., any of SEQ ID NOs: 79-81) fused to a transactivation domain. In some embodiments, the method reduces the expression level of a target nucleic acid in a cell, for example, using an engineered Cas12b effector protein based on an enzymatically inactivated Cas12b protein (e.g., any of SEQ ID NOs: 79-81) fused to a transcriptional repressor domain (e.g., a KRAB domain). In some embodiments, the method introduces an epigenetic modification to a target nucleic acid in a cell, for example, using an engineered Cas12b effector protein based on an enzymatically inactivated Cas12b protein (e.g., any of SEQ ID NOs: 79-81) fused to an epigenetic modification domain. In some embodiments, the methods introduce base edits into target nucleic acids in cells using engineered Cas12b effector proteins based on, for example, an enzymatically inactivated Cas12b protein (e.g., any of SEQ ID NOs:79-81) fused to a cytosine deaminase domain or an adenosine deaminase domain (e.g., TadA) or a functional fragment thereof. Depending on the functional domains contained in the engineered Cas12b effector protein, the engineered Cas12b systems described herein can be used to introduce other modifications into target nucleic acids.

[0214] In some embodiments, the method alters the target sequence of the target nucleic acid in the cell. In some embodiments, the method introduces a mutation into the target nucleic acid in the cell. In some embodiments, the method uses one or more endogenous DNA repair pathways in the cell (e.g., non-homologous end joining (NHEJ) or homology-directed recombination (HDR)) to repair double-strand breaks in the target DNA caused by sequence-specific cleavage of the CRISPR complex. Exemplary mutations include, but are not limited to, insertions, deletions, substitutions, and frameshifts. In some embodiments, the method inserts donor DNA into the target locus. In some embodiments, the insertion of the donor DNA introduces a selectable marker or reporter protein into the cell. In some embodiments, the insertion of the donor DNA leads to knock-in of a gene. In some embodiments, the insertion of the donor DNA causes a knock-out mutation. In some embodiments, the insertion of the donor DNA causes a substitution mutation, such as a single nucleotide substitution. In some embodiments, the method induces a phenotypic change in the cell.

[0215] In some embodiments, the engineered CRISPR-Cas12b system is used as part of a genetic circuit or to insert a genetic circuit into the genomic DNA of a cell. The engineered cleaved Cas12b effector protein controlled by an inducing agent described herein may be particularly useful as a component of a genetic circuit. The genetic circuit may be useful for gene therapy. Methods and techniques for designing and using genetic circuits are known in the art. See, for example, Brophy, Jennifer AN, and Christopher A. Voigt. "Principles of genetic circuit design." Nature methods 11.5 (2014): 508.

[0216] The engineered CRISPR-Cas12b system described herein can be used to modify a wide range of target nucleic acids. In some embodiments, the target nucleic acid is present in a cell. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is extrachromosomal DNA. In some embodiments, the target nucleic acid is exogenous to the cell. In some embodiments, the target nucleic acid is a viral nucleic acid, such as viral DNA. In some embodiments, the target nucleic acid is a plasmid in a cell. In some embodiments, the target nucleic acid is a horizontally transferred plasmid. In some embodiments, the target nucleic acid is RNA, such as mRNA.

[0217] In some embodiments, the target nucleic acid is an isolated nucleic acid, such as an isolated DNA. In some embodiments, the target nucleic acid is in a cell-free environment. In some embodiments, the target nucleic acid is an isolated vector, such as a plasmid. In some embodiments, the target nucleic acid is an isolated linear DNA fragment.

[0218] The methods described herein are applicable to any suitable cell type. In some embodiments, the cell is a bacterium, a yeast cell, a fungal cell, an algae cell, a plant cell, or an animal cell (e.g., a mammalian cell, such as a human cell). In some embodiments, the cell is a cell isolated from a natural source, such as a tissue biopsy. In some embodiments, the cell is a cell isolated from an in vitro cultured cell line. In some embodiments, the cell is from a primary cell line. In some embodiments, the cell is from an immortalized cell line. In some embodiments, the cell is a genetically engineered cell.

[0219] In some embodiments, the cells are animal cells derived from living organisms, including, but not limited to, cat, dog, mouse, rat, hamster, cow, sheep, goat, horse, donkey, pig, deer, chicken, duck, goose, rabbit, and fish.

[0220] In some embodiments, the cell is a plant cell derived from an organism selected from the group consisting of corn, wheat, barley, oat, rice, soybean, oil palm, safflower, sesame, tobacco, flax, cotton, sunflower, pearl barley, millet, sorghum, oilseed rape, cannabis, vegetable crops, forage crops, industrial crops, woody crops, and biomass crops.

[0221] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a mouse cell, such as a Neuro 2A (N2a) cell. In some embodiments, the cell is a human cell. In some embodiments, the human cell is a human embryonic kidney 293T (HEK293T or 293T) cell or a HeLa cell. In some embodiments, the mammalian cell is selected from the group consisting of an immune cell, a liver cell, a tumor cell, a stem cell, a neuronal cell, a fertilized egg, a muscle cell, and a skin cell.

[0222] In some embodiments, the cell is an immune cell selected from the group consisting of a cytotoxic T cell, a helper T cell, a natural killer (NK) T cell, an iNK-T cell, an NK-T-like cell, a gamma delta T cell, a tumor-infiltrating T cell, and a dendritic cell (DC)-activated T cell. In some embodiments, the method generates a modified immune cell, such as a CAR-T cell, a CAR-NK cell, or a TCR-T cell.

[0223] In some embodiments, the cell is an embryonic stem (ES) cell, an induced pluripotent stem (iPS) cell, a precursor cell of a gamete, a gamete, a fertilized egg, or a cell in an embryo.

[0224] The methods described herein can be used to modify target cells in vivo, ex vivo or in vitro, and can be done in a manner that alters the cells such that the progeny or cell lines of the modified cells retain the modified phenotype. The modified cells and their progeny can be part of a multicellular organism, such as a plant or animal, and can have ex vivo or in vivo applications, such as genome editing or gene therapy.

[0225] In some embodiments, the modification method is carried out ex vivo. In some embodiments, after the engineered CRISPR-Cas12b system is introduced into the cells, the modified cells (e.g., mammalian cells) are grown ex vivo. In some embodiments, the modified cells are cultured to grow for at least about 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 12 days, or 14 days. In some embodiments, the modified cells are cultured for no more than about 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 12 days, or 14 days. In some embodiments, the modified cells are further evaluated or screened to select cells having one or more desired phenotypes or characteristics, or evaluated or screened by PCR or sequencing.

[0226] In some embodiments, the target sequence is a sequence associated with a disease or condition. Exemplary diseases or conditions include, but are not limited to, cancer, blood disease, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, nervous system disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection. In some examples, the disease or condition is a graft-versus-host disease (GvHD) or host-versus-graft disease (HvG) disease. In some embodiments, the disease or condition is a genetic disease. In some embodiments, the disease or condition is a monogenic disease or condition. In some embodiments, the disease or condition is a polygenic disease or condition.

[0227] In some embodiments, the target sequence has a mutation compared to a wild-type sequence, hi some embodiments, the target sequence has a single nucleotide polymorphism (SNP) associated with a disease or condition.

[0228] In some embodiments, the donor DNA inserted into the target nucleic acid encodes a biological product selected from the group consisting of a reporter protein, an antigen-specific receptor, a therapeutic protein, an antibiotic resistance protein, an RNAi molecule, a cytokine, a kinase, an antigen, an antigen-specific receptor, a chimeric receptor, a cytokine receptor, and a suicide polypeptide. In some embodiments, the donor DNA encodes a therapeutic protein, such as a cytokine. In some embodiments, the donor DNA encodes a therapeutic protein useful for gene therapy. In some embodiments, the donor DNA encodes a therapeutic antibody. In some embodiments, the donor DNA encodes an engineered receptor, such as a chimeric antigen receptor (CAR) or an engineered TCR. In some embodiments, the donor DNA encodes a therapeutic RNA, such as a small RNA (e.g., siRNA, shRNA, or miRNA) or a long non-coding RNA (lincRNA).

[0229] The methods described herein can be used for multiplex gene editing or regulation of two or more (e.g., 2, 3, 4, 5, 6, 8, 10 or more) different target loci. In some embodiments, the methods detect or modify multiple target nucleic acids or target nucleic acid sequences. In some embodiments, the methods include contacting a target nucleic acid with a guide RNA comprising multiple (e.g., 2, 3, 4, 5, 6, 8, 10 or more) crRNA sequences, each of which comprises a different target sequence.

[0230] Also provided are engineered cells comprising modified target nucleic acids produced using any of the modification methods described herein. The engineered cells are useful for cell therapy. Autologous or allogeneic cells can be used to generate engineered cells using the methods described herein for cell therapy.

[0231] The methods described herein can also be used to generate isogenic cell lines (e.g., mammalian cells) to study genetic variants.

[0232] Also provided is an engineered plant or non-human animal comprising the engineered cell described herein. In some embodiments, the engineered plant or non-human animal is a genome-edited plant or non-human animal. The engineered non-human animal can be used as a disease model.

[0233] Techniques for generating non-human genome editing or transgenic animals are well known in the art, including but not limited to pronuclear microinjection, viral infection, transformation of embryonic stem cells and induced pluripotent stem (iPS) cells. Detailed methods that can be used include, but are not limited to, those described in Sundberg and Ichiki (2006, Genetically Engineered Mice Handbook, CRC Press) and Gibson (2004, A Primer Of Genome Science 2nd ed. Sunderland, Mass.: Sinauer).

[0234] The engineered animals may be of any suitable species, including, but not limited to, bovine, equine, ovine, canine, cervid, feline, caprine, porcine, primate, and less common mammals such as elephants, deer, zebras, and camels.

[0235] Treatment method Also provided are therapeutic methods using any of the methods described herein for modifying a target nucleic acid in a cell, and diagnostic methods using any of the methods described herein for detecting a target nucleic acid.

[0236] In some embodiments, the present application provides a method of treating a disease or condition associated with a target nucleic acid in an individual's cell, comprising contacting the target nucleic acid with any of the engineered CRISPR-Cas12b systems described herein, wherein a guide sequence of the guide RNA is complementary to a target sequence of the target nucleic acid, and wherein a Cas12b nuclease (e.g., engineered) or its effector protein (e.g., including any of SEQ ID NOs: 1-22 and 79-81) and the guide RNA bind to the target nucleic acid in association with each other to modify the target nucleic acid, thereby treating the disease or condition. In some embodiments, a mutation (e.g., a knock-out or knock-in mutation) is introduced into the target nucleic acid. In some embodiments, expression of the target nucleic acid is enhanced. In some embodiments, expression of the target nucleic acid is inhibited.

[0237] In some embodiments, the present application provides a method of treating a disease or condition in an individual, comprising administering to the individual an effective amount of any of the engineered CRISPR-Cas12b systems described herein and donor DNA encoding a therapeutic agent, wherein a guide sequence of the guide RNA is complementary to a target sequence of a target nucleic acid of the individual, and wherein the Cas12b nuclease (e.g., engineered) or its effector protein (e.g., including any of SEQ ID NOs: 1-22 and 79-81) and the guide RNA bind to the target nucleic acid in association with each other to insert the donor DNA into the target sequence, thereby treating the disease or condition.

[0238] In some embodiments, the present application provides a method of treating a disease or condition in an individual, comprising administering to the individual an effective amount of an engineered cell comprising a modified target nucleic acid, the engineered cell being produced by contacting the cell with any of the engineered CRISPR-Cas12b systems described herein, wherein a guide sequence of the guide RNA is complementary to a target sequence of a target nucleic acid, and the Cas12b nuclease (e.g., engineered) or its effector protein (e.g., including any of SEQ ID NOs: 1-22 and 79-81) and the guide RNA associate with each other to bind to the target nucleic acid and modify the target nucleic acid. In some embodiments, the engineered cell is an immune cell.

[0239] In some embodiments, the method includes contacting a target nucleic acid with an individual (e.g., ex vivo or in vivo) or administering to the individual an effective amount of a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring), the CRISPR-Cas12b system comprising (a) an engineered Cas12b nuclease or its effector protein (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor) that comprises one, two, or three types of mutations relative to a reference Cas12b nuclease, the mutations comprising (1) a positively charged amino acid residue (e.g., R , H, K), and / or (2) substitution of one or more amino acid residues involved in DNA duplex cleavage in a reference Cas12b nuclease (e.g., one or more of positions 118 and 119) with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (3) substitution of one or more amino acid residues that interact with ssDNA substrates in the RuvC domain of a reference Cas12b nuclease (e.g., 300, 301, 304, 329, 6 36, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease has the sequence set forth in SEQ ID NO:and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA, wherein the guide sequence hybridizes to the target sequence of the target nucleic acid, thereby mediating contact of the engineered Cas12b nuclease or its effector protein with the target sequence of the target nucleic acid, thereby leading to modification of the target nucleic acid by the engineered Cas12b nuclease or its effector protein, thereby treating the disease or condition. In some embodiments, the method includes contacting a target nucleic acid with an individual (e.g., ex vivo or in vivo) or administering to an individual an effective amount of a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring), the CRISPR-Cas12b system comprising (a) a sequence encoding a target nucleic acid sequence of the ...A Cas12b nuclease or its effector protein (e.g., a nickase, a truncated Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a prime editor) comprising the amino acid sequence of NO:1, or an engineered Cas12b nuclease or its effector protein, comprising one, two, or three types of mutations relative to a reference Cas12b nuclease, the mutations including (1) substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K), and / or (2) substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K), and / or (3) substitution of one or more amino acid residues that interact with DNA duplexes in the reference Cas12b nuclease (e.g., one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K). (3) substitution of one or more amino acid residues involved in strand cleavage (e.g., one or more of positions 118 and 119) with amino acid residues having an aromatic ring (e.g., F, Y, W), and / or (4) substitution of one or more amino acid residues that interact with the ssDNA substrate in the RuvC domain of the reference Cas12b nuclease (e.g., 300, 301, 304, 329, 636, 639, 647, 682, 757, 75 and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA, wherein the gRNA comprises one or more of the amino acid sequences of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8, 7, 8,The present invention provides a method for treating a disease or condition associated with a target nucleic acid in an individual (e.g., human) cell, comprising hybridizing a guide sequence to a target sequence of the target nucleic acid, thereby mediating contact of a Cas12b nuclease (e.g., engineered) or its effector protein with the target sequence of the target nucleic acid, thereby leading to modification of the target nucleic acid by the Cas12b nuclease (e.g., engineered) or its effector protein, thereby treating the disease or condition. In some embodiments, the engineered Cas12b nuclease or its effector protein comprises any of the amino acid sequences of SEQ ID NOs:2-22 and SEQ ID NOs:79-81. In some embodiments, the method further comprises contacting the target nucleic acid (e.g., ex vivo or in vivo) with an effective amount of a repair / donor nucleic acid or administering to the individual an effective amount of a repair / donor nucleic acid, wherein the repair / donor nucleic acid can incorporate into the modified target nucleic acid (e.g., via homologous recombination) to a target sequence. In some embodiments, modification of the target nucleic acid repairs a mutation (e.g., a loss-of-function mutation) in the target nucleic acid to a wild-type (or harmless version) sequence. In some embodiments, modification of the target nucleic acid introduces a foreign sequence.

[0240] In some embodiments, the individual is a human. In some embodiments, the individual is a model animal, such as, for example, a rodent (e.g., mouse, rat, hamster), a pet (e.g., cat, dog, rabbit), or a farm animal (e.g., horse, cow, sheep, goat, donkey, pig). In some embodiments, the individual is a mammal.

[0241] In some embodiments, the disease or condition is associated with an abnormality (e.g., a pathogenic point mutation) of a target nucleic acid in an individual (e.g., a human). In some embodiments, the disease or condition is treated by modifying (e.g., cleavage, base editing, or repair) the target nucleic acid (e.g., repair abnormality) by the CRISPR-Cas12b system or complex. In some embodiments, the disease is caused by overexpression or misexpression (e.g., missense mutation, frameshift mutation, nonsense mutation) of one or more target genes, and the CRISPR-Cas12b system or complex can target one or more target genes and perform targeted modification such as cleavage, base editing, or sequence repair (e.g., further introducing a repair / donor template to repair the cleaved target gene via the CRISPR-Cas12b system or complex by homologous recombination).

[0242] In some embodiments, the disease or condition is selected from the group consisting of cancer, a cardiovascular disease, a genetic disease, an autoimmune disease, a metabolic disease, a neurodegenerative disease, an eye disease, a bacterial infection, and a viral infection.

[0243] In some embodiments, the disease or condition is transthyretin amyloidosis (ATTR) (e.g., transthyretin-associated wild-type amyloidosis (ATTRwt), transthyretin-associated hereditary amyloidosis (ATTRm), familial amyloidotic polyneuropathy (FAP, ATTR-PN) or familial amyloidotic cardiomyopathy (FAC, ATTR-CM), cystic fibrosis, hereditary angioedema (HAE), diabetes, progressive Duchenne muscular dystrophy, Baker muscular dystrophy (BMD), alpha-1 antitrypsin deficiency (AAT deficiency), Pompe disease, myotonic dystrophy, Habeas disease, or any of a number of other conditions. In some embodiments, the CRISPR-Cas12b system or complex is packaged and delivered via lipid nanoparticles. In some embodiments, the lipid nanoparticles are administered to the individual via intravenous injection or infusion.

[0244] In some embodiments, the target nucleic acid is PCSK9. In some embodiments, the disease or condition is cardiovascular disease. In some embodiments, the disease or condition is coronary artery disease. In some embodiments, the method reduces cholesterol levels in an individual. In some embodiments, the method treats diabetes in an individual. In some embodiments, the disease or condition is hypercholesterolemia, such as familial hypercholesterolemia.

[0245] In some embodiments, the target nucleic acid is HBG1 and / or HBG2. In some embodiments, the disease or condition is sickle cell disease or beta thalassemia. In some embodiments, the disease or condition is hereditary perseveration of fetal hemoglobin (HPFH), HbS-gene deletion HPFH, or HbS-HPFH due to point mutations.

[0246] In some embodiments, the target nucleic acid is CC chemokine receptor (CCR) 5 (CCR5), which encodes the major HIV-1 coreceptor. In some embodiments, the disease or condition is an infectious disease, e.g., AIDS. In some embodiments, the disease or condition is a non-infectious disease, such as cancer (e.g., breast or prostate cancer), atherosclerosis, stroke, or inflammatory bowel disease (IBD).

[0247] In some embodiments, the target nucleic acid is CD34. In some embodiments, the disease or condition is cancer.

[0248] In some embodiments, the target nucleic acid is Ring Finger Protein 2 (RNF2).In some embodiments, the disease or condition is a neurological disease, such as Luo-Schoch-Yamamoto syndrome or non-specific syndromic intellectual disability.

[0249] Detection Method The present application also provides methods of detecting target nucleic acids using engineered Cas12b nucleases or their effector proteins (e.g., including any of SEQ ID NOs: 2-22 and 79-81) with improved activity, or the CRISPR-Cas12b system. The use of Cas12b effector proteins as detection agents takes advantage of the discovery that V-type CRISPR / Cas12 proteins (e.g., Cas12a, Cas12b, Cas12c, Cas12d, Cas12e (CasX) and Cas12i) can promiscuously cleave non-target single-stranded DNA (ssDNA) when activated by detection of target DNA. Methods of using Cas12b proteins as detection agents are described, for example, in US10253365 and WO2020 / 056924, the contents of which are incorporated herein by reference in their entirety. In some embodiments, detection of a target nucleic acid in a sample diagnoses a disease or condition.

[0250] In some embodiments, when the Cas12b effector protein is activated by the guide RNA, if the sample contains a target DNA to which the guide RNA is hybridized (i.e., the sample contains target DNA), the Cas12b nuclease or its effector protein becomes a nuclease that randomly cleaves single-stranded nucleic acids (e.g., non-target ssDNA or RNA, i.e., single-stranded nucleic acids to which the guide sequence of the guide RNA does not hybridize). Thus, if target DNA (double-stranded or single-stranded) is present in the sample (e.g., in some cases in amounts higher than a threshold value), the single-stranded nucleic acids in the sample are cleaved, which can be detected by any convenient detection method (e.g., a labeled single-stranded detection nucleic acid such as DNA or RNA). Cas12b can cleave ssDNA and ssRNA.

[0251] In some embodiments, methods are provided for detecting target DNA (e.g., double-stranded or single-stranded) in a sample, comprising: (a) contacting the sample with (i) any of the engineered Cas12b nucleases or effector proteins thereof described herein (e.g., including any of SEQ ID NOs: 2-22 and 79-81); (ii) a guide RNA comprising a guide sequence that hybridizes to the target DNA; and (iii) a single-stranded detection nucleic acid (i.e., a "single-stranded detection nucleic acid") that does not hybridize to the guide sequence of the guide RNA; and (b) cleaving the single-stranded detection nucleic acid with the engineered Cas12b effector protein to measure a detectable signal. In some embodiments, a method is provided for detecting a target DNA (e.g., double-stranded or single-stranded) in a sample, comprising: (a) contacting the sample with (i) any of the Cas12b nucleases (e.g., engineered or wild-type) described herein or their effector proteins (e.g., including any of SEQ ID NOs: 1-22 and 79-81); (ii) a guide RNA comprising a guide sequence that hybridizes to the target DNA, and an engineered scaffold comprising any of the sequences of SEQ ID NOs: 25-53; and (iii) a single-stranded detection nucleic acid (i.e., a "single-stranded detection nucleic acid") that does not hybridize to the guide sequence of the guide RNA; and (b) cleaving the single-stranded detection nucleic acid with the engineered Cas12b effector protein to measure a detectable signal. In some embodiments, a method is provided for detecting a target nucleic acid in a sample, comprising: (a) contacting the sample with an engineered CRISPR-Cas12b system described herein and a labeled detection nucleic acid, wherein the gRNA comprises a guide sequence complementary to a target sequence of a target nucleic acid, and the labeled detection nucleic acid is single stranded and does not hybridize to the guide sequence of the gRNA; and (b) detecting the target nucleic acid by cleaving the labeled detection nucleic acid with the engineered CRISPR-Cas12b system and measuring a detectable signal resulting from the cleavage.In some cases, the single-stranded detection nucleic acid comprises a fluorescent dye pair (e.g., the fluorescent dye pair is a fluorescence energy resonance transfer (FRET) pair, a quencher / fluorescent pair). In some cases, the target DNA is viral DNA (e.g., papillomavirus, hepadnavirus, herpesvirus, adenovirus, poxvirus, parvovirus, etc.). In some embodiments, the single-stranded detection nucleic acid is DNA. In some embodiments, the single-stranded detection nucleic acid is RNA. In some embodiments, the engineered Cas12b effector protein is an engineered Cas12b nuclease. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is present in a cell, such as a bacterial cell, a yeast cell, a plant cell, or an animal cell. In some embodiments, the method is performed ex vivo. In some embodiments, the method is performed in vivo. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target sequence is associated with a disease or condition.

[0252] The disclosed method for detecting target DNA (single-stranded or double-stranded) in a sample can detect target DNA with high sensitivity. In some cases, the disclosed method can be used to detect target DNA present in a sample containing multiple DNAs (including target DNA and multiple non-target DNAs), where the target DNA is more than 10 6 It is present as one or more copies per 10 non-target DNA (e.g., 6 1 or more copies of target DNA per 10 non-target DNA 5 1 or more copies of target DNA per 10 non-target DNA 4 1 or more copies of target DNA per 10 non-target DNA 3 1 or more copies of target DNA per 10 non-target DNA 2one or more copies of target DNA per non-target DNA, one or more copies of target DNA per 50 non-target DNA, one or more copies of target DNA per 20 non-target DNA, one or more copies of target DNA per 10 non-target DNA, or one or more copies of target DNA per 5 non-target DNA).

[0253] In some embodiments, an engineered Cas12b nuclease or its effector protein described herein (e.g., including any of SEQ ID NOs:2-22) can detect target DNA with greater sensitivity than a reference Cas12b nuclease (e.g., SEQ ID NO:1). In some embodiments, an engineered Cas12b effector protein can detect target DNA with 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more sensitivity than a reference Cas12b nuclease.

[0254] Delivery method In some embodiments, the engineered CRISPR-Cas12b system or components thereof, nucleic acid molecules thereof, or nucleic acid molecules encoding or providing the components thereof described herein can be delivered to a host cell by multiple delivery systems, such as a plasmid or viral vector (e.g., any of the vectors described in the "Constructs and Vectors" subsection above). In some embodiments or methods, the engineered CRISPR-Cas12b system can be delivered by other methods, such as nucleofection or electroporation of a ribonucleoprotein complex composed of an engineered Cas12b nuclease or its effector protein and its homolog RNA guide.

[0255] In some embodiments, the delivery is via nanoparticles or exosomes.

[0256] In some embodiments, the paired Cas12b-nickase complex may be delivered directly by nanoparticles or other direct protein delivery methods, resulting in co-delivery of a complex containing two paired crRNA elements. Additionally, proteins may be delivered to cells by viral vectors or directly to cells to directly deliver a CRISPR array containing two paired spacer regions to perform double nicking. In some cases, to directly deliver RNA, the RNA may be conjugated to at least one sugar moiety, such as N-acetylgalactosamine (GalNAc), particularly triantennary N-acetylgalactosamine. In some embodiments, the CRISPR-Cas12b system or components thereof are packaged and delivered via lipid nanoparticles. In some embodiments, the lipid nanoparticles are administered to an individual by intravenous injection or infusion.

[0257] IV. KITS AND PRODUCTS Also provided are compositions, kits, unit doses, and articles of manufacture that include one or more components of an engineered Cas12b nuclease or its effector protein, an engineered scaffold-containing sgRNA (e.g., any of SEQ ID NOs:25-53), or an engineered CRISPR-Cas12b system described herein.

[0258] In some embodiments, a kit is provided that includes one or more AAV vectors encoding any of the engineered Cas12b nucleases or their effector proteins, or engineered CRISPR-Cas12b systems described herein. In some embodiments, the kit further includes one or more guide RNAs, such as sgRNAs (e.g., any of SEQ ID NOs:25-53) that include an engineered scaffold. In some embodiments, the kit further includes donor DNA. In some embodiments, the kit further includes a cell, such as a human cell.

[0259] The kit may include one or more additional components to allow for the growth of the engineered cells, such as containers, reagents, media, cytokines, buffers, antibodies, etc. The kit may further include a means for administration of the composition.

[0260] The kit may further include instructions for using the engineered CRISPR-Cas12b system described herein, such as a method for detecting or modifying a target nucleic acid. In some embodiments, the kit includes instructions for treating or diagnosing a disease or condition. Instructions for use of the components of the kit typically include information regarding the expected dose of treatment, administration schedule, and route of administration. The containers may be unit doses, bulk packages (e.g., multi-dose packages), or sub-unit doses. For example, the kit may be provided containing a sufficient dose of the compositions disclosed herein to provide a long-term effective treatment to an individual. The kit may further include a plurality of unit doses of the composition and instructions for use packaged in a number sufficient for storage and use in a pharmacy (e.g., a hospital pharmacy or a compounding pharmacy).

[0261] The kits of the present application employ suitable packaging. Suitable packaging includes, but is not limited to, vials, bottles, cans, flexible packaging (e.g., sealed polyester film or plastic bags), and the like. The kits may optionally provide additional components, such as buffers and interpretive information. Thus, the present application also provides articles of manufacture that include vials (e.g., sealed vials), bottles, cans, flexible packaging, and the like.

[0262] The article of manufacture may include a container and a label or package insert on or associated with the container. Suitable containers include, for example, bottles, vials, syringes, and the like. The container may be formed from a number of materials, such as glass or plastic. Generally, the container holds a composition effective for treating a disease or condition described herein and may have a sterile access port (e.g., the container may be an intravenous solution bag or a vial having a stopper pierceable by a hypodermic injection needle). The label or package insert indicates that the composition is used to treat a particular condition in an individual. The label or package insert further includes instructions for administering the composition to an individual.

[0263] Package inserts are instructions typically included in the commercial packaging of therapeutic products that contain information regarding the indications, usage, dosage, administration, contraindications and / or warnings concerning the use of those therapeutic products.

[0264] Moreover, the article of manufacture may further comprise a second container comprising a pharma- ceutically acceptable buffer, such as bacteriostatic water for injection (BWFI), phosphate-buffered saline, Ringer's solution, dextrose solution, etc. It may also include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, and syringes.

[0265] Exemplary embodiments Embodiment 1. An engineered Cas12b nuclease, comprising one, two or three types of mutations relative to a reference Cas12b nuclease, the mutations being: (1) a substitution of one or more amino acid residues that interact with a protospacer adjacent motif (PAM) in the reference Cas12b nuclease with a positively charged amino acid residue; and / or (2) Substitution of one or more amino acid residues involved in DNA double-strand cleavage in the reference Cas12b nuclease with amino acid residues having an aromatic ring, and / or (3) An engineered Cas12b nuclease, comprising a substitution of one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with a positively charged amino acid residue or a hydrophobic amino acid residue.

[0266] Embodiment 2. The engineered Cas12b nuclease of embodiment 1, wherein said reference Cas12b nuclease is a wild-type Cas12b nuclease.

[0267] Embodiment 3. The engineered Cas12b nuclease of embodiment 1 or 2, wherein the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO:1.

[0268] Embodiment 4. The engineered Cas12b nuclease of any one of embodiments 1 to 3, comprising a substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease with a positively charged amino acid residue.

[0269] Embodiment 5. The engineered Cas12b nuclease of embodiment 4, wherein one or more amino acid residues that interact with the PAM are within 9 Å of the PAM in the three-dimensional structure.

[0270] Embodiment 6. The engineered Cas12b nuclease of embodiment 4 or 5, wherein the one or more amino acid residues that interact with the PAM are at one or more of positions 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and / or 475, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0271] Embodiment 7. The engineered Cas12b nuclease of embodiment 6, wherein the one or more amino acid residues that interact with the PAM comprise one or more of the following amino acid residues: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and / or E475, wherein said amino acid residues are numbered according to SEQ ID NO:1.

[0272] Embodiment 8. The engineered Cas12b nuclease of embodiment 7, wherein the one or more amino acid residues that interact with the PAM include one or more of the amino acid residues D116 and E475, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0273] Embodiment 9. The engineered Cas12b nuclease of any one of embodiments 4 to 8, wherein the positively charged amino acid residue is R or K.

[0274] Embodiment 10. The engineered Cas12b nuclease of embodiment 9, wherein the substitution of one or more amino acid residues that interact with the PAM in said reference Cas12b nuclease with a positively charged amino acid residue is one or more of the following substitutions: D116R and / or E475R, said amino acid residues being numbered according to SEQ ID NO:1.

[0275] Embodiment 11. The engineered Cas12b nuclease of any one of embodiments 1 to 10, comprising a substitution of one or more amino acid residues involved in cleaving the DNA double strand in the reference Cas12b nuclease with an amino acid residue having an aromatic ring.

[0276] Embodiment 12. The engineered Cas12b nuclease of embodiment 11, wherein the one or more amino acid residues involved in the cleavage of the DNA duplex interact with the last base pair in the PAM relative to the 3' end of the target strand.

[0277] Embodiment 13. The engineered Cas12b nuclease of embodiment 11 or 12, wherein the one or more amino acid residues involved in the cleavage of the DNA double strand are at one or more of positions 118 and / or 119, said amino acid residues being numbered according to SEQ ID NO:1.

[0278] Embodiment 14. The engineered Cas12b nuclease of any one of embodiments 11 to 13, wherein the amino acid residue having an aromatic ring is Y, F or W.

[0279] Embodiment 15. The engineered Cas12b nuclease of embodiment 14, wherein the substitution of one or more amino acid residues involved in DNA double-strand cleavage in said reference Cas12b nuclease with an amino acid residue having an aromatic ring is Q119Y, Q119F or Q119W, said amino acid residues being numbered according to SEQ ID NO:1.

[0280] Embodiment 16. The engineered Cas12b nuclease of any one of embodiments 1 to 16, comprising a substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease that interact with the single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue.

[0281] Embodiment 17. The engineered Cas12b nuclease of embodiment 16, wherein one or more amino acid residues in the RuvC domain that interact with the single-stranded DNA substrate are within 9 Å of the single-stranded DNA substrate in a three-dimensional structure.

[0282] Embodiment 18. The engineered Cas12b nuclease of embodiment 17, wherein the one or more amino acid residues in the RuvC domain that interact with the single-stranded DNA substrate are at one or more of positions 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and / or 1097, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0283] Embodiment 19. The engineered Cas12b nuclease of embodiment 18, wherein the one or more amino acid residues located within the RuvC domain and interacting with a single-stranded DNA substrate comprise one or more of the following amino acid residues: D300, K301, E304, N329, E636, Q639, T647, Q682, I757, E758, E761, E764, K768, E852, Q854, N856, N857, D858, P860, S862, E863, N865, Q866, L867, Q869, E938, E956, G957, E958, I994, Q1093, and / or W1097, wherein said amino acid residues are numbered according to SEQ ID NO:1.

[0284] Embodiment 20. The engineered Cas12b nuclease of embodiment 19, comprising a substitution of one or more of the following amino acid residues with a positively charged amino acid residue: E636, I757, E758, E761, Q854, N857, N865, Q866, Q869, and / or Q1093, wherein said amino acid residues are numbered according to SEQ ID NO:1.

[0285] Embodiment 21. The engineered Cas12b nuclease of embodiment 20, wherein the positively charged amino acid residue is R or K.

[0286] Embodiment 22. The engineered Cas12b nuclease of embodiment 21, wherein the substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease and interacting with the single-stranded DNA substrate is one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R and / or N857K, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0287] Embodiment 23. The engineered Cas12b nuclease of embodiment 19, comprising a substitution of one or more of the following amino acid residues with a hydrophobic amino acid residue: E758, E761, E863, N865, Q866, Q869, Q956 and / or Q1093, wherein said amino acid residues are numbered according to SEQ ID NO:1.

[0288] Embodiment 24 The engineered Cas12b nuclease of embodiment 23, wherein the hydrophobic amino acid residue is W, Y, F or M.

[0289] Embodiment 25. The engineered Cas12b nuclease of embodiment 24, wherein the substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease and interacting with the single-stranded DNA substrate is one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W and / or Q1093Y, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0290]

[0036] Embodiment 26. The engineered Cas12b nuclease is selected from the group consisting of: (1) D116R, (2) E475R, (3) Q119F and E475R, (4) Q119F, E475R, and E758R, (5) Q119Y, (6) Q119F, (7) Q119W, (8) I757R, (9) E758R, (10) E761R, (11) K768R, (12) I757R and E758R, (13) I757R and E761R, (14) I757R and K768R, (15) E758R and E761R, (16) E758 4. The engineered Cas12b nuclease of any one of embodiments 1 to 3, comprising any one or combination of the following substitutions: R and K768R, (17) E761R and K768R, (18) I757R, E758R, and E761R, (19) I757R, E758R, and K768R, (20) I757R, E761R and K768R, (21) E758R, E761R, and K768R, (22) I757R, E758R, E761R, and K768R, (23) Q866M, (24) Q869M, and (25) Q866M and Q869M, wherein the amino acid residues are numbered according to SEQ ID NO:1.

[0291] Embodiment 27. The engineered Cas12b nuclease of any one of embodiments 1 to 26, comprising an amino acid sequence having at least 85% sequence identity to any of SEQ ID NOs: 20 to 22.

[0292] Embodiment 28. The engineered Cas12b nuclease of any one of embodiments 1 to 27, further comprising one or more mutations that increase flexibility of a flexible region comprising amino acid residues 855 to 859, said amino acid positions being numbered according to SEQ ID NO:1.

[0293] Embodiment 29. The engineered Cas12b nuclease of embodiment 28, wherein the one or more mutations that improve flexibility include N856G.

[0294] Embodiment 30. An engineered Cas12b nuclease comprising: (1) D116R; (2) E475R; (3) Q119F and E475R; (4) Q119F, E475R, and E758R; (5) Q119Y; (6) Q119F; (7) Q119W; (8) Q119F and E475R; (9) Q119F, E475R, and E758R. 58R(10)E636R, (11)I757R, (12)E758R, (13)E761R, (14)Q854R, (15)N857K, (16)Q119F, E475R, and E758R, (17)K768R, (18)I757R and E758R, (19)I757R and E761R, (20)I757R and K768R , (21) E758R and E761R, (22) E758R and K768R, (23) E761R and K768R, (24) I757R, E758R, and E761R, (25) I757R, E758R, and K768R, (26) I757R, E761R, and K768R, (27) E758R, E761R, and K768R , (28) I757R, E758R, E761R, and K768R (29) N865W, (30) N865Y, (31) Q866M, (32) Q869M, (33) Q1093W, (34) Q1093Y, and / or (35) any one or more of the following mutations, wherein the amino acid positions are numbered according to SEQ ID NO:1.

[0295] Embodiment 31. An engineered Cas12b nuclease comprising an amino acid sequence set forth in any one of SEQ ID NOs: 2-22.

[0296] Embodiment 32. An engineered Cas12b effector protein comprising an engineered Cas12b nuclease or a functional derivative thereof according to any one of embodiments 1 to 31.

[0297] Embodiment 33 The engineered Cas12b effector protein of embodiment 32, wherein the engineered Cas12b nuclease or functional derivative thereof has enzymatic activity.

[0298] Embodiment 34. The engineered Cas12b effector protein of embodiment 32 or 33, wherein the engineered Cas12b effector protein is capable of inducing a double-stranded break in a DNA molecule.

[0299] Embodiment 35. The engineered Cas12b effector protein of embodiment 32 or 33, wherein the engineered Cas12b effector protein is capable of inducing a single-strand break in a DNA molecule.

[0300] Embodiment 36 The engineered Cas12b effector protein of embodiment 32, wherein the engineered Cas12b effector protein comprises an enzymatically inactivating mutation of the engineered Cas12b nuclease.

[0301] Embodiment 37. The engineered Cas12b effector protein of embodiment 36, wherein the enzymatically inactivating mutations comprise D570A, R785A, E848A, R911A and / or D977A.

[0302] Embodiment 38. The engineered Cas12b effector protein of any one of embodiments 32 to 37, further comprising a functional domain fused to said engineered Cas12b nuclease or functional derivative thereof.

[0303] Embodiment 39. The engineered Cas12b effector protein of embodiment 38, wherein the functional domain is selected from the group consisting of a translation initiator domain, a transcriptional repressor domain, a transactivation domain, an epigenetic modification domain, a nucleobase editing domain, a reverse transcriptase domain, a reporter domain, and a nuclease domain.

[0304] Embodiment 40. The engineered Cas12b effector protein of any one of embodiments 32 to 37, comprising a first polypeptide comprising an N-terminal portion of an engineered Cas nuclease or a functional derivative thereof, and a second polypeptide comprising a C-terminal portion of an engineered Cas nuclease or a functional derivative thereof, wherein the first and second polypeptides are capable of associating with each other in the presence of a guide RNA comprising a guide sequence to form a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence.

[0305] Embodiment 41. The engineered Cas12b effector protein of embodiment 40, comprising a first polypeptide comprising N-terminal amino acid residues 1 to X of an engineered Cas12b nuclease or functional derivative thereof, and a second polypeptide comprising a residue at position X+1 to the C-terminus of said engineered Cas12b nuclease or functional derivative thereof, wherein said first and second polypeptides are capable of associating with each other in the presence of a guide RNA comprising a guide sequence to form a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence.

[0306] Embodiment 42 The engineered Cas12b effector protein of embodiment 40 or 41, wherein the first polypeptide and the second polypeptide each comprise a dimerization domain.

[0307] Embodiment 43 The engineered Cas12b effector protein of embodiment 42, wherein the first dimerization domain and the second dimerization domain associate with each other in the presence of an inducer.

[0308] Embodiment 44 The engineered Cas12b effector protein of embodiment 40 or 41, wherein the first and second polypeptides do not comprise a dimerization domain.

[0309] Embodiment 45. An engineered CRISPR-Cas12b system comprising: (a) an engineered Cas12b effector protein according to any one of embodiments 32 to 44, or a nucleic acid encoding the engineered Cas12b effector protein; and (b) a guide RNA comprising a guide sequence complementary to a target sequence, or a nucleic acid encoding the guide RNA, wherein the engineered Cas12b effector protein and the guide RNA form a CRISPR complex that specifically binds to a target nucleic acid comprising the target sequence and is capable of inducing modification of the target nucleic acid.

[0310] Embodiment 46. The engineered CRISPR-Cas12b system of embodiment 45, wherein the gRNA comprises a crRNA and a tracrRNA.

[0311] Embodiment 47. The engineered CRISPR-Cas12b system of embodiment 45 or 46, comprising a precursor guide RNA array encoding multiple crRNAs.

[0312] Embodiment 48. The engineered CRISPR-Cas12b system of any one of embodiments 45 to 47, wherein the guide RNA is a single guide RNA (sgRNA).

[0313] Embodiment 49. The engineered CRISPR-Cas12b system of any one of embodiments 45 to 48, comprising one or more vectors encoding the engineered Cas12b effector proteins.

[0314] Embodiment 50. The engineered CRISPR-Cas12b system of embodiment 49, wherein the one or more vectors are adeno-associated virus (AAV) vectors.

[0315] Embodiment 51. The engineered CRISPR-Cas12b system of embodiment 50, wherein the AAV vector further encodes the guide RNA.

[0316] Embodiment 52. A method for detecting a target nucleic acid in a sample, comprising: (a) contacting the sample with an engineered CRISPR-Cas12b system according to any one of embodiments 45 to 51 and a labeled detection nucleic acid, wherein the labeled detection nucleic acid is single-stranded and does not hybridize to a guide sequence of the guide RNA; and (b) detecting the target nucleic acid by measuring a detectable signal produced by cleavage of the labeled detection nucleic acid by the engineered Cas12b effector protein.

[0317] Embodiment 53. A method for modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with an engineered CRISPR-Cas12b system described in any one of embodiments 45 to 51.

[0318] Embodiment 54 The method of embodiment 53, which is carried out in vitro.

[0319] Embodiment 55 The method of embodiment 53, wherein the target nucleic acid is present in a cell.

[0320] Embodiment 56 The method of embodiment 55, wherein the cell is a bacterial cell, a yeast cell, a mammalian cell, a plant cell, or an animal cell.

[0321] Embodiment 57 The method of embodiment 53, which is carried out ex vivo.

[0322] Embodiment 58 The method of embodiment 53, which is carried out in vivo.

[0323] Embodiment 59. The method of any one of embodiments 53 to 58, wherein the engineered CRISPR-Cas12b system cleaves the target nucleic acid or alters the target sequence in the target nucleic acid.

[0324] Embodiment 60. The method of any one of embodiments 53 to 58, wherein the engineered CRISPR-Cas12b system alters expression of the target nucleic acid.

[0325] Embodiment 61. The method of any one of embodiments 53 to 60, wherein the target nucleic acid is genomic DNA.

[0326] Embodiment 62. The method of any one of embodiments 53 to 61, wherein the target sequence is associated with a disease or condition.

[0327] Embodiment 63. The method of any one of embodiments 53 to 62, wherein the engineered CRISPR-Cas12b system comprises a precursor guide RNA array encoding multiple crRNAs, each crRNA comprising a different guide sequence.

[0328] Embodiment 64. A method for treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising the step of modifying the target nucleic acid in a cell of the individual using an engineered CRISPR-Cas12b system described in any one of embodiments 45 to 51, thereby treating the disease or condition.

[0329] Embodiment 65. The method of embodiment 64, wherein the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.

[0330] Embodiment 66. An engineered cell comprising a modified target nucleic acid, wherein the target nucleic acid has been modified by the method of any one of embodiments 53 to 63.

[0331] Embodiment 67. An engineered non-human animal comprising one or more engineered cells of embodiment 66. EXAMPLES

[0332] The following examples are used only to illustrate the present application and therefore should not be considered as limiting the present application in any way. The following examples and detailed description are intended to be illustrative and not limiting.

[0333] method Plasmid construction The coding sequence of AaCas12b was codon-optimized for expression and synthesis in human cells. A nucleic acid sequence encoding an engineered AaCas12b protein variant was generated by PCR-based site-directed mutagenesis. Specifically, the DNA sequence encoding the reference AaCas12b protein was split into two parts around the mutation site. Two pairs of primers were designed to amplify the two parts of the DNA sequence, respectively, and assembled into one DNA fragment by Gibson cloning and then incorporated into the pCAG-2A-eGFP vector. The DNA encoding the reference AaCas12b protein was cut into multiple segments, followed by amplification and assembly using PCR and Gibson cloning to construct the mutant combination. The DNA encoding the engineered AaCas12b protein was inserted between the XmaI and NhheI sites of the pCAG-2A-eGFP vector. The mutation sites of the AaCas12b protein variants were designed based on the analysis of the crystal structure of AaCas12b using protein structure visualization software (e.g., PyMol or Chimera) commonly used in the field. The crystal structure of AaCas12b is available in the RCSB PDB database under the accession numbers 6LTU, 6LTR, 6LU0, and 6LTP. The AaCas12b variants were expressed in human 293T cells using the pCAG-2A-eGFP vector. The DNA sequence encoding the sgRNA scaffold was de novo synthesized and assembled into a pUC19-U6 backbone by Gibson cloning. The nucleic acid encoding the spacer sequence was also ligated into the same pUC19-U6 backbone.

[0334] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded into 24-well plates (Corning) and cultured for 16 h until cell confluence reached 70%. 600 ng of plasmid encoding AaCas12b protein and different amounts of plasmid encoding sgRNA were transfected into cells in each well of the 24-well plate using Lipofectamine 3000 (Invitrogen). 68 h after transfection, the HEK293T cells were digested with trypsin-EDTA (0.05%) (Gibco) and FACS sorted using MoFlo XDP (Beckman Coulter) based on GFP signal (indicating successful transfection).

[0335] Targeted deep sequencing analysis for genome modifications FACS-sorted GFP-positive HEK293T cells were lysed with buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. The corresponding primer pairs were used to PCR-amplify dsDNA fragments containing target sites at different loci. For targeted deep sequencing, the cell lysate was directly used as template DNA to perform barcode PCR amplification. The PCR products were purified and then collected into several libraries for high-throughput sequencing. The insertion / deletion frequency (%) was analyzed by calculating the percentage of reads containing insertions or deletions using CRISPResso2 software. In this application, the insertion / deletion frequency (%) index was used to compare and analyze gene editing efficiency in the presence of different engineered Cas12b proteins and / or different sgRNA scaffolds. Reads that were less than 0.05% of the total reads were discarded.

[0336] Example 1: Substitution of one or more amino acid residues that interact with the PAM in a reference AaCas12b nuclease with a positively charged amino acid residue An engineered AaCas12b enzyme with single point mutations in amino acid residues interacting with the PAM was designed and expressed according to the above method. Briefly, 10 amino acids within 9 Å of the PAM in AaCas12b were selected: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and E475, and each amino acid residue was replaced with arginine (R). Nucleic acids encoding sgRNAs for target sites CCR5-11 (SEQ ID NO:63), CD34-7 (SEQ ID NO:64), and RNF2-1 (SEQ ID NO:65) were designed and cloned into a pUC19-U6 backbone, which includes, from 5' to 3', DNA encoding the Aa-sg sgRNA scaffold sequence (SEQ ID NO:23)-DNA encoding the spacer sequence. Using Lipofectamine 3000 (Invitrogen), 600 ng of the plasmid encoding the AaCas12b protein and 300 ng of the plasmid encoding the sgRNA were transfected into HEK293T cells in each well of a 24-well plate as described above. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control. The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiencies are shown in Figure 1 and Table 1. AaCas12b variants with amino acid substitutions D116R (SEQ ID NO: 2) or E475R (SEQ ID NO: 3) showed improved gene editing efficiency compared to wild-type AaCas12b. As shown in Figure 1, the average gene editing efficiency of AaCas12b-D116R (SEQ ID NO:2) and AaCas12b-E475R (SEQ ID NO:3) at the three genomic sites was greater than about 20%, while the average gene editing efficiency of the reference wild-type AaCas12b nuclease was about 6%. The insertion / deletion frequencies of the AaCas12b-D116R (SEQ ID NO:2) and AaCas12b-E475R (SEQ ID NO:3) mutants were significantly higher than the insertion / deletion frequencies of other AaCas12b mutants using such tests.AaCas12b-D395R achieved higher gene editing efficiency at the CD34-7 locus, but not at other loci tested, compared with wild-type AaCas12b.

[0337] [Table 1]

[0338] Example 2: Substitution of one or more amino acid residues involved in DNA duplex cleavage in the reference AaCas12b nuclease with amino acid residues having an aromatic ring An engineered AaCas12b nuclease with a single substitution in an amino acid residue involved in DNA double strand cleavage was designed and expressed according to the method described above. Briefly, amino acid residue Q118 or Q119 is replaced by an aromatic amino acid residue (e.g., Y, F, or W). Here, the same sgRNA-encoding plasmid as in Example 1 was used. 600 ng of the plasmid encoding the AaCas12b protein and 300 ng of the plasmid encoding the sgRNA were transfected into HEK293T cells in each well of a 24-well plate using Lipofectamine 3000 (Invitrogen) as described above. Wild-type AaCas12b (SEQ ID NO:1) was used as a control. The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiencies are shown in Figure 2 and Table 2. AaCas12b with amino acid substitutions Q119Y, Q119F, or Q119W showed improved gene editing efficiency at all loci tested compared to wild-type AaCas12b. The insertion / deletion frequencies of the AaCas12b-Q119Y, AaCas12b-Q119F and AaCas12b-Q119W mutants were significantly higher than the insertion / deletion frequencies of other AaCas12b mutants (Q118Y, Q118F, Q118W) using such tests.

[0339] [Table 2]

[0340] Example 3: Substitution of one or more amino acid residues that interact with single-stranded DNA substrates in the RuvC domain of the reference AaCas12b nuclease with positively charged or hydrophobic amino acid residues An engineered AaCas12b nuclease with single amino acid substitutions in amino acid residues in the RuvC domain that interact with single-stranded DNA substrates was designed and expressed according to the methods described above. Nucleic acids encoding sgRNAs for target sites CCR5-3 (SEQ ID NO:66) and RNF2-5 (SEQ ID NO:67) were designed and cloned into a pUC19-U6 backbone, which contains, from 5' to 3', DNA encoding the Aa-sg sgRNA scaffold sequence (SEQ ID NO:23)-DNA encoding the spacer sequence. 600 ng of the plasmid encoding the AaCas12b protein and 300 ng of the plasmid encoding the sgRNA were transfected into HEK293T cells in each well of a 24-well plate using Lipofectamine 3000 (Invitrogen). Wild-type AaCas12b (SEQ ID NO:1) was used as a control.

[0341] In the first group of AaCas12b mutants, each amino acid residue in Table 3 was replaced by arginine (R), a positively charged amino acid residue. The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiencies are shown in Figures 3 to 4B and Table 3.

[0342] [Table 3]

[0343] In the second group of AaCas12b mutants, each amino acid residue in Table 4 was replaced by a positively charged amino acid residue, lysine (K). The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiencies are shown in Table 4 and Figures 4A-4B.

[0344] As shown in Tables 3-4 and Figures 3-4B, AaCas12b variants with amino acid substitutions D300R, K301R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, K768R, Q854R, N857R, D858R, N865R, Q866R, I994R, Q1093R, W1097R, E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, I994K, Q1093K, or W1097K improved gene editing efficiency at both loci tested compared to wild-type AaCas12b. The insertion / deletion frequencies of the AaCas12b E636R, I757R, E758R, E761R, Q854R, D858R, E758K, N857K, I994R, and D858K mutants were significantly higher than the insertion / deletion frequencies with other AaCas12b mutants tested (which were substituted with positively charged amino acid residues).

[0345] [Table 4]

[0346] In the third group of AaCas12b mutants, the following amino acid residues were replaced by hydrophobic amino acid residues (e.g., Y, F, M, or W), respectively: E758, E761, E863, N865, Q866, Q869, E956, and Q1093. The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiencies are shown in Figure 5 and Table 5. AaCas12b mutants with amino acid substitutions E758W, E758Y, E758M, E761Y, N865W, N865Y, N865F, Q866M, Q869M, Q1093W, Q1093Y, Q1093F, or Q1093M showed improved gene editing efficiency at the two loci tested compared to wild-type AaCas12b. The insertion / deletion frequencies of the AaCas12b N865W, N865Y, Q866M, Q869M, Q1093W and Q1093Y mutants were significantly higher than the insertion / deletion frequencies of other such AaCas12b mutants (substituted with hydrophobic amino acid residues) tested.

[0347] [Table 5]

[0348] Example 4: Combinations of mutations from Examples 1-3 and characterization of their gene editing efficiency Using the amino acid substitutions with desired gene editing efficiency screened in Examples 1, 2 and 3, i.e., Q866M, Q869M, I757R, E758R, E761R, K768R and I757R, multiple mutations, i.e., Q866M+Q869M, I757R+E758R, I757R+E761R, I757R+K768R, E7 AaCas12b proteins were produced having the following structures: 58R+E761R, E758R+K768R, E761R+K768R, I757R+E758R+E761R, I757R+E758R+K768R, I757R+E761R+K768R, E758R+E761R+K768R, and I757R+E758R+E761R+K768R. Nucleic acids encoding sgRNAs for target sites CCR5-3 (SEQ ID NO:66), CCR5-11 (SEQ ID NO:63), CD34-1 (SEQ ID NO:68), and RNF2-5 (SEQ ID NO:67) were prepared and cloned into a pUC19-U6 backbone, which contains, from 5' to 3', a DNA encoding an Aa-sg sgRNA scaffold sequence (SEQ ID NO:23)-a DNA encoding a spacer sequence. Wild-type AaCas12b (SEQ ID NO:1) was used as a control. 600ng of the plasmid encoding the above AaCas12b protein and 300ng of the plasmid encoding the sgRNA were transfected into HEK293T cells in each well of a 24-well plate using Lipofectamine 3000 (Invitrogen). Their gene editing efficiencies are shown in Figure 6 and Table 6. AaCas12b mutants with combinations of amino acid substitutions showed significantly improved gene editing efficiency at all loci tested compared to wild-type AaCas12b. At certain loci, AaCas12b combination mutants, such as Q866M+Q869M, E758R+E761R, E758R+E768R, I757R+E758R+K768R, and E758R+E761R+K768R, have improved gene editing efficiency compared to the corresponding single mutants.

[0349] [Table 6]

[0350] AaCas12b-Q119F+E475R and AaCas12b-Q119F+E475R+E758R were generated as described above. Here, the same sgRNA-encoding plasmid as in Example 1 was used. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control. 600 ng of the plasmid encoding the AaCas12b protein and 300 ng of the plasmid encoding the sgRNA were transfected into HEK293T cells in each well of a 24-well plate using Lipofectamine 3000 (Invitrogen). Their gene editing efficiencies are shown in Figure 7 and Table 7. As a result, it was revealed that compared with the wild-type AaCas12b, AaCas12b-Q119F+E475R and AaCas12b-Q119F+E475R+E758R significantly improved gene editing efficiency at all loci tested. AaCas12b-Q119F+E475R+E758R showed the most significant improvement in gene editing efficiency at all loci tested (CCR5-11, CD34-7, and RNF2-1) compared to wild-type AaCas12b or corresponding AaCas12b variants with single substitutions.

[0351] [Table 7]

[0352] Example 5: Enhancement of gene editing activity of engineered AaCas12b using sgRNAs with engineered scaffolds In this example, the gene editing activity of multiple sgRNAs with engineered scaffolds was tested using the AaCas12b mutant (Q119F+E475R+E758R) of Example 4. A nucleic acid encoding an sgRNA for the target site CCR5-11 (SEQ ID NO:63) was designed and cloned into a pUC19-U6 backbone, which contains, from 5' to 3', a DNA encoding a DNA-spacer sequence encoding the sgRNA scaffold sequence. Using Lipofectamine 3000 (Invitrogen), 600 ng of plasmid encoding AaCas12b variant proteins, 300 ng of plasmid encoding sgRNA with engineered scaffolds (SEQ ID NO: 25-53; modified based on AacCas12b sgRNA scaffold V0), AacCas12b sgRNA scaffold (SEQ ID NO: 24; V0, control; H. Yang et al., Cell. 2016;167(7):1814-1828.e12), or AaCas12b Aa-sg scaffold (SEQ ID NO: 23; control) were transfected into HEK293T cells in each well of a 24-well plate. Their gene editing efficiencies are shown in Figure 9. The data in Figure 9 showed that all sgRNA engineered scaffolds significantly improved the gene editing efficiency of the AaCas12b (Q119F + E475R + E758R) variant compared to the AacCas12b sgRNA scaffold (V0). All sgRNA engineered scaffolds (except V1 and V8) significantly improved the gene editing efficiency of the AaCas12b (Q119F + E475R + E758R) variant compared to the Aa-sg scaffold.

[0353] Example 6: Engineered AaCas12b with inactivated nuclease activity To generate an inactivated AaCas12b protein, the AaCas12b (Q119F+E475R+E758R) variant (SEQ ID NO:22) of Example 4 was further modified to include an additional single point mutation (D570A) in the nucleolytic domain (Figure 10A). Plasmids encoding i) AaCas12b(Q119F+E475R+E758R) or AaCas12b(Q119F+E475R+E758R+D570A) (SEQ ID NO:79) under the control of the CMV promoter, and ii) control sgRNA (not targeting any sequence in hemoglobin subunit gamma 1 / 2 (HBG1 / 2)), sgRNA1 (target sequence SEQ ID NO:70 for HBG1 / 2), or sgRNA2 (target sequence SEQ ID NO:71 for HBG1 / 2) under the control of the U6 promoter were transfected into HEK293 cells in a similar manner as described above (Table 8; see FIG. 10A for plasmid construction). The sgRNAs were constructed using the sgRNA scaffold (V9) (SEQ ID NO:53). Three days after transfection, genomic DNA was extracted from the transfected cells. To determine the cleavage efficiency, a T7 endonuclease I (T7EI) mismatch detection assay was performed (M. Crispo et al., PLoS One. 2015;10(8):e0136690). The primer sequences used for the T7EI assay are shown in Table 9.

[0354] As shown in Figure 10B, the catalytic activity of AaCas12b(Q119F+E475R+E758R+D570A) was significantly reduced compared with AaCas12b(Q119F+E475R+E758R) when it cleaved two different target sites in HBG1 / 2 under the guidance of sgRNA1 or sgRNA2.

[0355] [Table 8]

[0356] [Table 9]

[0357] To further reduce the nuclease activity of engineered AaCas12b, additional point mutations were introduced into AaCas12b(Q119F+E475R+E758R+D570A) to generate AaCas12b(Q119F+E475R+E758R+D570A+E848A) (SEQ ID NO:80) or AaCas12b(Q119F+E475R+E758R+D570A+D977A) (SEQ ID NO:81) (Figure 11A). Using a method similar to that described above, HEK293 cells were transfected with plasmids co-encoding i) AaCas12b(Q119F+E475R+E758R), AaCas12b(Q119F+E475R+E758R+D570A+E848A), or AaCas12b(Q119F+E475R+E758R+D570A+D977A) under the control of a CMV promoter, and ii) sgRNA1 (target sequence SEQ ID NO:70 for HBG1 / 2) or sgRNA2 (target sequence SEQ ID NO:71 for HBG1 / 2) under the control of a U6 promoter (see FIG. 11A for plasmid construction). As a negative control, plasmids encoding AaCas12b(Q119F+E475R+E758R), AaCas12b(Q119F+E475R+E758R+D570A+E848A), or AaCas12b(Q119F+E475R+E758R+D570A+D977A) were transfected into HEK293 cells lacking any sgRNA coding sequence, as well as a control sgRNA (not targeting any sequence within hemoglobin subunit gamma 1 / 2 (HBG1 / 2)). As shown in Figure 11B, both AaCas12b (Q119F+E475R+E758R+D570A+E848A) and AaCas12b (Q119F+E475R+E758R+D570A+D977A) completely eliminated the nuclease activity of AaCas12b (Q119F+E475R+E758R).

[0358] Example 7: Transcriptional repression using engineered AaCas12b fusion proteins AaCas12b (Q119F+E475R+E758R+D570A+D977A) (SEQ ID NO:81) from Example 6 was further engineered to generate a fusion protein that silences the transcription of a target gene. AaCas12b (Q119F+E475R+E758R+D570A+D977A) (flanked by two copies of the nuclear localization sequence NLS) was fused to the Kruppel-associated box (KRAB) domain (SEQ ID NO:72) of the transcriptional repression module ZIM3 to recruit repressive chromatin modifiers. KRAB was fused to the C-terminus or N-terminus of AaCas12b (Q119F+E475R+E758R+D570A+D977A), and the fusion proteins were named Cd12bk and Nd12bk, respectively. The same plasmid also expresses the SCN9A gene (voltage-gated sodium channel 1.7Na) under the control of the U6 promoter. v 1.7) (Figure 12A; Table 10). To construct those sgRNAs, the sgRNA scaffold (V9) (SEQ ID NO:53) was used.

[0359] To investigate whether Cd12bk and Nd12bk fusion proteins can recruit chromatin-modifying complexes to silence SCN9A transcription, we transfected plasmids encoding the fusion proteins and sgRNAs into Neuro 2A (N2a; a mouse neural crest-derived cell line) cells. As a control, a plasmid encoding Cd12bk was transfected into N2a cells as well as a control sgRNA (not targeting any sequence within SCN9A). Three days after transfection, transfected cells were harvested and RNA was extracted using an RNA extraction kit (Vazyme, product no. RC112-01). The mRNA levels of Nav1.7 in each sample were determined by qPCR. Data were normalized to Cd12bk using a control sgRNA ("Cd12bk-non-target"). As shown in Figure 12B, Cd12bk or Nd12bk, together with sgRNA-msg6, sgRNA-msg8, sgRNA-msg13 or sgRNA-msg18, significantly suppressed the transcription of SCN9A, among which sgRNA-msg8 and sgRNA-msg13 showed the strongest suppression. Nd12bk can also significantly suppress the transcription of SCN9A together with sgRNA-msg11. These results suggest that dAaCas12b fused to KRAB (e.g., AaCas12b(Q119F+E475R+E758R+D570A+D977A)) is useful as a target transcription control tool in eukaryotic cells.

[0360] [Table 10]

[0361] Although the embodiments of the present application have been described above with reference to the drawings, the present application is not limited to the above specific embodiments or application fields. The above specific embodiments are illustrative and instructive, and are not limiting. With the teachings of this specification and without departing from the scope of the claims of the present application, a person skilled in the art can come up with many more forms, all of which fall within the scope of protection of the present application.

[0362] JPEG2025501473000012.jpg245158

[0363] JPEG2025501473000013.jpg241151

[0364] JPEG2025501473000014.jpg241151

[0365] JPEG2025501473000015.jpg241151

[0366] JPEG2025501473000016.jpg241151

[0367] JPEG2025501473000017.jpg241151

[0368] JPEG2025501473000018.jpg241151

[0369] JPEG2025501473000019.jpg241151

[0370] JPEG2025501473000020.jpg241151

[0371] JPEG2025501473000021.jpg241151

[0372] JPEG2025501473000022.jpg241151

[0373] JPEG2025501473000023.jpg241151

[0374] JPEG2025501473000024.jpg241151

[0375] JPEG2025501473000025.jpg241151

[0376] JPEG2025501473000026.jpg241151

[0377] JPEG2025501473000027.jpg241151

[0378] JPEG2025501473000028.jpg241151

[0379] JPEG2025501473000029.jpg241151

[0380] JPEG2025501473000030.jpg241151

[0381] JPEG2025501473000031.jpg241151

[0382] JPEG2025501473000032.jpg241151

[0383] JPEG2025501473000033.jpg241151

[0384] JPEG2025501473000034.jpg241151

[0385] JPEG2025501473000035.jpg241151

[0386] JPEG2025501473000036.jpg241151

[0387] JPEG2025501473000037.jpg241151

[0388] JPEG2025501473000038.jpg241151

[0389] JPEG2025501473000039.jpg241151

[0390] JPEG2025501473000040.jpg241151

[0391] JPEG2025501473000041.jpg241151

[0392] JPEG2025501473000042.jpg241151

[0393] JPEG2025501473000043.jpg241151

[0394] JPEG2025501473000044.jpg241151

[0395] JPEG2025501473000045.jpg211151

Claims

1. It comprises one, two or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations are: (1) Substitution of one or more amino acid residues that interact with a protospacer adjacent motif (PAM) in the reference Cas12b nuclease with positively charged amino acid residues, and / or (2) Substitution of one or more amino acid residues involved in DNA double-strand cleavage in the reference Cas12b nuclease with amino acid residues having an aromatic ring, and / or (3) The reference Cas12b nuclease includes a substitution of one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain with a positively charged amino acid residue or a hydrophobic amino acid residue; Preferably, the reference Cas12b nuclease is an engineered Cas12b nuclease comprising the amino acid sequence of SEQ ID NO:

1. Claim 2: The reference Cas12b nuclease comprises a substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease with positively charged amino acid residues, wherein the one or more amino acid residues that interact with PAM are at one or more of positions 475, 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, and 400, and wherein the amino acid residues are numbered according to SEQ ID NO: 1; Preferably, the one or more amino acid residues that interact with PAM include one or more of the amino acid residues D116 and E475; More preferably, the positively charged amino acid residue is R or K; More preferably, the substitution of one or more amino acid residues that interact with PAM in the reference Cas12b nuclease with positively charged amino acid residues is one or more of the following substitutions: E475R, and D116R; More preferably, the engineered Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 2 or 3.

2. The engineered Cas12b nuclease of claim 1.

3. the reference Cas12b nuclease comprises a substitution of one or more amino acid residues involved in the cleavage of the DNA double strand with an amino acid residue having an aromatic ring, wherein the one or more amino acid residues involved in the cleavage of the DNA double strand are at one or more of positions 118 and 119, and the amino acid residues are numbered according to SEQ ID NO: 1; Preferably, the amino acid residue having an aromatic ring is Y, F, or W; More preferably, the substitution of one or more amino acid residues involved in DNA double-strand cleavage in the reference Cas12b nuclease with an amino acid residue having an aromatic ring is Q119F, Q119Y, or Q119W; More preferably, the engineered Cas12b nuclease comprises an amino acid sequence set forth in any one of SEQ ID NOs: 4 to 6.

2. The engineered Cas12b nuclease of claim 1.

4. The engineered Cas12b nuclease comprises a substitution of one or more amino acid residues located within the RuvC domain in the reference Cas12b nuclease and interacting with the single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue; The one or more amino acid residues located within the RuvC domain and interacting with the single-stranded DNA substrate are at one or more of positions 758, 866, 300, 301, 304, 329, 636, 639, 647, 682, 757, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097, wherein the amino acid residues are numbered according to SEQ ID NO: 1; Preferably, it comprises a substitution of one or more of the amino acid residues E636, Q639, T647, Q682, I757, E758, E761, K768, Q854, N857, D858, N865, Q866, I994, Q1093, and W1097 with a positively charged amino acid residue; More preferably, the positively charged amino acid residue is R or K; More preferably, the substitution of one or more amino acid residues located in the RuvC domain in the reference Cas12b nuclease and interacting with the single-stranded DNA substrate is one or more of the following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R, and W1097R; More preferably, it comprises any of the amino acid sequences of SEQ ID NOs: 7 to 13, or a substitution of one or more of the amino acid residues E758, E761, E863, N865, Q866, Q869, Q956, and Q1093 with a hydrophobic amino acid residue; More preferably, the hydrophobic amino acid residue is W, Y, F, or M; Or, The substitution of one or more amino acid residues located within the RuvC domain in the reference Cas12b nuclease and interacting with the single-stranded DNA substrate is one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W, and Q1093Y; Preferably, it comprises any of the amino acid sequences of SEQ ID NO: 14 to 19.

2. The engineered Cas12b nuclease of claim 1.

5. (1) E475R, (2) D116R, (3) Q119F+E475R, (4) Q119F+E475R+E758R, (5) Q11 9Y, (6) Q119F, (7) Q119W, (8) I757R, (9) E758R, (10) E761R, (11) K768R, (1 2) I757R+E758R, (13) I757R+E761R, (14) I757R+K768R, (15) E758R+E761 R, (16) E758R+K768R, (17) E761R+K768R, (18) I757R+E758R+E761R, (19) I (20) I757R+E758R+K768R, (21) E758R+E761R+K768R, (22) I757R+E758R+E761R+K768R, (23) Q866M, (24) Q869M, (25) Q866M+Q869M, (26) E636R, (27) Q854R, (28) N857K, (29) N865W, (30) N865Y, (31) Q1093W, (32) Q1093Y, and (33) D858R substitutions or combinations thereof, wherein the amino acid residues are numbered according to SEQ ID NO: 1; Preferably, it comprises any of the amino acid sequences of SEQ ID NO: 20 to 22.

2. The engineered Cas12b nuclease of claim 1.

6. one or more mutations of the reference Cas12b nuclease, wherein the mutations increase the flexibility of a flexible region comprising amino acid residues 855 to 859, wherein the amino acid residues are numbered according to SEQ ID NO: 1; the one or more mutations that increase flexibility include N856G; 2. The engineered Cas12b nuclease of claim 1.

7. The engineered Cas12b nuclease of any one of claims 1 to 6 or a functional derivative thereof, Preferably, the engineered Cas12b nuclease comprises an enzymatically inactivated mutant thereof, More preferably, the engineered enzymatically inactive mutant of Cas12b nuclease comprises a substitution of one or more amino acid residues selected from the group consisting of D570A, E848A, R785A, E848A, R911A, and D977A, wherein said amino acid residues are numbered according to SEQ ID NO:

1. More preferably, the engineered enzymatically inactivated mutant of Cas12b nuclease comprises any of the amino acid sequences of SEQ ID NO: 79-81. Engineered Cas12b effector proteins.

8. 8. The engineered Cas12b effector protein of claim 7, further comprising a functional domain fused to said engineered Cas12b nuclease or functional derivative thereof.

9. A single guide RNA (sgRNA) comprising any of the sequences of SEQ ID NOs: 25-53.

10. (I) (a) an engineered Cas12b nuclease according to any one of claims 1 to 6, an engineered Cas12b effector protein according to any one of claims 7 to 8, or nucleic acids encoding them; (b) a guide RNA (gRNA) comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA; The engineered Cas12b nuclease or the engineered Cas12b effector protein and the gRNA form a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid; (II) (a) a Cas12b nuclease or a Cas12b effector protein, or a nucleic acid encoding the same, comprising any of the amino acid sequences of SEQ ID NOs: 1-22 and 79-81; (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising any of the sequences of SEQ ID NOs: 25 to 53; The Cas12b nuclease or the Cas12b effector protein and the gRNA form a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid. Engineered CRISPR-Cas12b system.

11. 1. A method for detecting a target nucleic acid in a sample, comprising:

11. A method comprising: (a) contacting the sample with the engineered CRISPR-Cas12b system of claim 10 and a labeled detection nucleic acid, wherein the gRNA comprises a guide sequence complementary to a target sequence of the target nucleic acid, and the labeled detection nucleic acid is single-stranded and does not hybridize to the guide sequence of the gRNA; and (b) detecting the target nucleic acid by cleaving the labeled detection nucleic acid with the engineered Cas12b nuclease or its effector protein to produce a detectable signal, or by Cas12b nuclease or its effector protein.

12. 11. A method of modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with the engineered CRISPR-Casl2b system of claim 10.

13. 1. An agent for treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising: A drug comprising the engineered CRISPR-Cas12b system of claim 10.

14. 1. An engineered cell comprising a modified target nucleic acid, 11. An engineered cell, wherein the target nucleic acid is modified in the engineered CRISPR-Cas12b system of claim 10.

15. 15. An engineered non-human animal comprising one or more engineered cells of claim 14.