Engineered Cas12i nucleases, effector proteins, and their applications

JP7913729B2Active Publication Date: 2026-09-01INST OF ZOOLOGY CHINESE ACAD OF SCI +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023573318
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-15
Filing Date
2022-05-25
Publication Date
2026-09-01
Estimated Expiration
2042-05-25

Smart Images

  • Figure 0007913729000027
    Figure 0007913729000027
  • Figure 0007913729000028
    Figure 0007913729000028
  • Figure 0007913729000029
    Figure 0007913729000029
Patent Text Reader

Abstract

The present application provides an engineered Cas12i nuclease that includes one or more mutations based on a reference Cas12i nuclease: (1) replacing one or more amino acids that interact with the PAM in the reference Cas12i nuclease with positively charged amino acids, and / or (2) replacing one or more amino acids involved in opening the DNA duplex in the reference Cas12i nuclease with amino acids with aromatic rings, and / or (3) replacing one or more amino acids located in the RuvC domain in the reference Cas12i nuclease that interact with a single-stranded DNA substrate with positively charged amino acids, and / or (4) replacing one or more amino acids that interact with the DNA-RNA duplex in the reference Cas12i nuclease with positively charged amino acids, and / or (5) replacing one or more polar or positively charged amino acids that interact with the DNA-RNA duplex in the reference Cas12i nuclease with hydrophobic amino acids. The engineered Cas12i nuclease has significantly improved gene editing efficiency and / or significantly reduced off-target phenomena compared to the reference Cas12i nuclease.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority Information Claims of this application: This application claims priority under Chinese patent application CN202111347952.7 filed on 15 November 2021, Chinese patent application CN202110581290.3 filed on 27 May 2021, and international application PCT / CN2021 / 096477 filed on 27 May 2021, the contents of which are incorporated herein by reference in their entirety.

[0002] Sequence listing file submitted at the same time All contents of the following ASCII code text file are incorporated herein by reference: Sequence listing in computer-readable format (CRF) (FF00614PCT-NSequence6thVersion-20220525.txt, date: May 25, 2022, size: 218kb)

[0003] This application belongs to the field of biotechnology. More specifically, this application relates to Cas12i nucleases, effector proteins, and applications thereof that have improved catalytic activity (e.g., gene editing activity). [Background technology]

[0004] Genome editing is an important and useful technology in genome research. Several systems are available for genome editing, including the CRISPR-Cas system, which uses clustered, regularly spaced short palindromic repeat sequences; the TALEN system, which uses transcription factor-like effector nucleases; and the Zinc finger nuclease (ZFN) system.

[0005] The CRISPR-Cas system is an efficient and cost-effective genome editing technology that can be widely applied to a range of eukaryotes, from yeast and plants to zebrafish and humans (see review articles: Van der Oost 2013, Science 339: 768-770, and Charpentier and Doudna, 2013, Nature 495: 50-51). The CRISPR-Cas system provides adaptive immunity in archaea and bacteria by binding Cas12i effect proteins to CRISPR RNA (crRNA). To date, the CRISPR-Cas system, comprising six types (types I-VI) and two subtypes (class I and class II), has been characterized based on the system's remarkable functional and evolutionary modularity. In the Category II CRISPR-Cas system, the Type II Cas9 system and the Type V A / B / E / J Cas12a / Cas12b / Cas12e / Cas12j systems are used for genome editing, offering broad potential for biomedical research.

[0006] However, current CRISPR-Cas systems have various limitations, including limited gene editing efficiency. Therefore, improved methods and systems are needed for effective genome editing across multiple gene loci. [Overview of the project]

[0007] This application provides the following technical solutions: 1. Includes one or more mutations based on the reference Cas12i nuclease, wherein the mutations are (1) Substitution of one or more amino acids that interact with PAM in the reference Cas12i nuclease with positively charged amino acids; (2) Substituting one or more amino acids involved in the release of the DNA double strand in the reference Cas12i nuclease with aromatic ring-containing amino acids; (3) Substitution of one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with a single-stranded DNA substrate with positively charged amino acids; (4) Substituting one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with positively charged amino acids; and (5) Selected from substituting one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with hydrophobic amino acids; Preferably, the reference Cas12i nuclease is an engineered Cas12i nuclease, which is a wild-type Cas12i nuclease with amino acid sequence SEQ ID NO:1.

[0008] 2. One or more amino acids that interact with the PAM are amino acids located within 9 Å of the PAM in three-dimensional structure; preferably, one or more amino acids that interact with the PAM are amino acids located at one or more of the positions 176, 178, 226, 227, 229, 237, 238, 264, 447, and 563. More preferably, the one or more amino acids that interact with the PAM are one or more amino acids from among E176, E178, Y226, A227, N229, E237, K238, K264, T447, and E563. More preferably, the one or more amino acids that interact with the PAM are one or more amino acids from among E176, K238, T447, and E563. Among these, the positional number of the amino acid is defined as the corresponding amino acid position shown in SEQ ID NO:1, the engineered Cas12i nuclease as described in item 1.

[0009] 3. The engineered Cas12i nuclease according to item 1 or 2, wherein the positively charged amino acid is R, K, or H.

[0010] 4. Substituting one or more amino acids that interact with PAM in the aforementioned reference Cas12i nuclease with positively charged amino acids refers to one or more substitutions of E176R, K238R, T447R, and E563R; Preferably, the engineered Cas12i nuclease includes any of the following mutations or combinations of mutations: (1) E563R; (2) E176R, T447R, E176R and E563R; (3) K238R and E563R; (4) E176R, K238R and T447R; (5) E176R, K238R and E563R; (6) E176R, T447R and E563R; and (7) any mutation or combination of mutations among E176R, K238R, T447R and E563R; Among these, the position of the amino acid is defined as the corresponding amino acid position shown in SEQ ID NO:1, an engineered Cas12i nuclease according to any one of items 1 to 3.

[0011] 5. One or more amino acids involved in the release of the DNA double helix are amino acids that interact with the last base pair at the 3' end of the target strand in the PAM; Preferably, one or more amino acids involved in the unwinding of the DNA double helix are amino acids located at one or more positions among positions 163 and 164; More preferably, the one or more amino acids involved in the release of the DNA double helix are one or more amino acids Q163 and N164; more preferably, the amino acid involved in the release of the DNA double helix is ​​N164; Among these, the positional number of the amino acid is defined as the corresponding amino acid position shown in SEQ ID NO:1, an engineered Cas12i nuclease according to any one of items 1 to 4.

[0012] 6. An engineered Cas12i nuclease according to any one of claims 1 to 5, wherein one or more amino acids involved in the unwinding of the DNA double helix are replaced with aromatic amino acids, the aromatic amino acids being F, Y, or W, preferably F or Y.

[0013] 7. The substitution of one or more amino acids involved in the unwinding of the DNA double strand in the reference Cas12i nuclease with aromatic ring-attached amino acids includes one or more substitutions of Q163F, Q163Y, Q163W, and N164F; preferably, the Cas12i nuclease includes the N164Y or N164F mutation; more preferably, the engineered Cas12i nuclease includes N164Y, as described in any one of items 1 to 6.

[0014] 8. The one or more amino acids located in the RuvC domain that interact with the single-stranded DNA substrate are amino acids that are within 9 Å of the single-stranded DNA substrate in three-dimensional structure; Preferably, one or more amino acids located in the RuvC domain that interact with the single-stranded DNA substrate are amino acids at one or more of the following positions: 323, 327, 355, 359, 360, 361, 362, 388, 390, 391, 392, 393, 414, 417, 418, 421, 424, 425, 650, 652, 653, 696, 705, 708, 709, 751, 752, 755, 840, 848, 851, 856, 885, 897, 925, 926, 928, 929, 932, and 1022. More preferably, the one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are one or more amino acids from among E323, L327, V355, G359, G360, K361, D362, L388, N390, N391, F392, K393, Q414, L417, L418, K421, Q424, Q425, S650, E652, G653, I696, K705, K708, E709, L751, S752, E755, N840, N848, S851, A856, Q885, M897, N925, I926, T928, G929, Y932, and A1022; More preferably, the one or more amino acids located in the RuvC domain that interact with the single-stranded DNA substrate are one or more amino acids selected from the group consisting of E323, D362, Q425, N925, I926, N391, Q424 and G929; The engineered Cas12i nuclease according to any one of Items 1 to 7, wherein the position number of said amino acid is defined as the corresponding amino acid position shown in SEQ ID NO: 1.

[0015] 9. The engineered Cas12i nuclease according to any one of Items 1 to 8, wherein substituting one or more amino acids located in the RuvC domain of a reference Cas12i nuclease that interact with a single-stranded DNA substrate with positively charged amino acids comprises substitution with R or K, and preferably said positively charged amino acid is R.

[0016] 10. Substituting one or more amino acids located in the RuvC domain of a reference Cas12i nuclease that interact with a single-stranded DNA substrate with positively charged amino acids comprises one or more substitutions selected from the group consisting of E323R, D362R, N391R, Q424R, Q425R, N925R, I926R and G929R; Preferably, said engineered Cas12i nuclease comprises any one mutation or combination of mutations selected from the group consisting of: (1) E323R; (2) D362R; (3) Q425R; (4) N925R; (5) I926R; (6) E323R and D362R; (7) E323R and Q425R; (8) E323R and I926R; (9) Q425R and I926R; (10) D362R and I926R; (11) N925R and I926R; (12) E323R, D362R and Q425R; (13) E323R, D362R and I926R; (14) E323R, Q425R and I926R; (15) D362R, N925R and I926R; and (16) E323R, D362R, Q425R and I926R; The engineered Cas12i nuclease according to any one of Items 1 to 9, wherein said amino acid position is defined as the corresponding amino acid position set forth in SEQ ID NO: 1.

[0017] 11. The one or more amino acids that interact with the DNA-RNA double helix are amino acids located within 9Å of the DNA-RNA double helix in terms of three-dimensional structure; Preferably, the one or more amino acids that interact with the DNA-RNA double helix are amino acids at one or more positions selected from the group consisting of 116, 117, 156, 159, 160, 161, 247, 293, 294, 297, 301, 305, 306, 308, 312, 313, 316, 319, 320, 343, 348, 349, 427, 433, 438, 441, 442, 679, 683, 691, 782, 783, 797, 800, 852, 853, 855, 861, 865, 957, and 958; More preferably, the one or more amino acids that interact with the DNA-RNA double helix are selected from the group consisting of G116, E117, A156, T159, S161, T301, I305, K306, T308, N312, F313, D427, K433, V438, N441, Q442, M852, L855, N861, Q865, E160, Q316, E319, Q320, E247, E343, E348, E349, N679, E683, E691, D782, E783, E797, E800, D853, S957, D958, G293, E294 and N297; one or more amino acids selected from the group consisting of G116, E117, A156, T159, E160, S161, E247, G293, E294, N297, T301, I305, K306, T308, N312, F313, Q316, E319, Q320, E343, E348, E349, D427, K433, V438, N441, Q442, N679, E683, E691, D782, E783, E797, E800, M852, D853, L855, N861, Q865, S957, and D958; More preferably, the one or more amino acids that interact with the DNA-RNA double helix are one or more amino acids from G116, E117, T159, S161, E319, E343 and D958; the position of the amino acid is defined as the corresponding amino acid position shown in SEQ ID NO:1, the engineered Cas12i nuclease according to any one of items 1 to 10.

[0018] 12. An engineered Cas12i nuclease according to any one of items 1 to 11, wherein the substitution of one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with a positively charged amino acid includes substitution with R or K, preferably the positively charged amino acid being R.

[0019] 13. Substitution of one or more amino acids that interact with the DNA-RNA double helix in a reference Cas12i nuclease with positively charged amino acids includes one or more substitutions of G116R, E117R, T159R, S161R, E319R, E343R, and D958R; Preferably, the engineered Cas12i nuclease contains a D958R substitution; Among these, the position of the amino acid is defined as the corresponding amino acid position shown in SEQ ID NO:1, an engineered Cas12i nuclease according to any one of items 1 to 12.

[0020] 14. One or more polar or positively charged amino acids that interact with the DNA-RNA double helix are selected from amino acids at one or more positions among 357, 394, 715, 719, 807, 844, 848, 857, and 861, i.e., one or more amino acids among H357, K394, R715, R719, K807, K844, N848, R857, and R861; Among these, the position of the amino acid is defined as the corresponding amino acid position shown in SEQ ID NO:1, an engineered Cas12i nuclease according to any one of items 1 to 13.

[0021] 15. An engineered Cas12i nuclease according to any one of items 1 to 14, wherein one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease are replaced with a hydrophobic amino acid, which includes replacement with alanine (A).

[0022] 16. Containing one or more mutations selected from the group consisting of H357A, K394A, R715A, R719A, K807A, K844A, N848A, R857A, and R861A; Preferably, the mutant comprises one or more mutations selected from the group consisting of K394A, R719A, K844A, and R857A; Among these, the aforementioned amino acid position is defined as the corresponding amino acid position shown in SEQ ID NO:1, wherein the engineered Cas12i nuclease is as described in any one of items 1 to 15.

[0023] 17. An engineered Cas12i nuclease according to any one of items 1 to 16, comprising the amino acid substitution of R719A and K844A, or comprising the amino acid substitution of R857A and K844A, wherein the amino acid positions are defined as the corresponding amino acid positions shown in SEQ ID NO:1.

[0024] 18. Further comprising one or more flexible region mutations, the mutations increasing the flexibility of the flexible region in the reference Cas12i nuclease, the flexible region being selected from amino acid residues 439-443 or amino acid residues 925-929; Preferably, the flexible region mutation is located at positions 439 and / or 926; More preferably, the flexible region mutation is a mutation at L439 and / or I926; Among these, the aforementioned amino acid position is defined as the corresponding amino acid position shown in SEQ ID NO:1, wherein the engineered Cas12i nuclease is as described in any one of items 1 to 17.

[0025] 19. The one or more flexible region mutations consist of substituting an amino acid in the flexible region with G and / or inserting one or two Gs after it; Preferably, the one or more flexible region mutations include I926G, L439(L+G), or L439(L+GG); More preferably, the one or more flexible region mutations include L439(L+G) or L439(L+GG); Among these, the aforementioned amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1, the engineered Cas12i nuclease described in item 18.

[0026] 20. An engineered Cas12i nuclease; The aforementioned engineered Cas12i nuclease is (1) E563R; (2) E176R and T447R; (3) E176R and E563R; (4) K238R and E563R; (5) E176R, K238R and T447R; (6) E176R, T447R and E563R; (7) E176R, K238R and E563R; (8) E176R, K238R, T447R and E563R; (9) N164Y; (10) N164F; (11) E323R; (12) D362R; (13) Q425R; (14) N925R; (15) I926R; (16) D958R; (17) E323R and D362R; (18) E323R and Q425R; (19) E323R and I926R; (20) Q425R and I926R; (21) D362R and I926R; (22) N925R and I926R; (23) E323R, D362R and Q425R; (24) (25) E323R, D362R and I926R; (26) D362R, N925R and I926R; (27) E323R, D362R, Q425R and I926R; (28) D362R and I926G; (29) N925R and I926G; (30) D362R, N925R and I926G; (31) I926R and L439(L+G); (32) I926R and L439(L+GG); (33) E323R, D362R and I926G; (34) R719A and K844A; and (35) R857A and K844A, including one or more groups of mutations; Preferably, the engineered Cas12i nuclease includes one or more groups of mutations from (1) E176R, K238R, T447R and E563R; (2) N164Y; (3) I926R; (4) E323R and D362R; (4) I926G; (5) I926R and L439(L+G); (6) I926R and L439(L+GG); and (7) D958R; Of these, the aforementioned amino acid position is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0027] 21. Engineered Cas12i nucleases, wherein the engineered Cas12i nucleases are (1) E176R, K238R, T447R, E563R and N164Y; (2) E176R, K238R, T447R, E563R and I926R; (3) N164Y, E323R and D362R; (4) E176R, K238R, T447R, E563R, E323R and D362R; (5) N164Y and I926R; (6) E176R, K238R, T447R, E563R, N164Y and I926R; (7) E176R, K238R, T447R, E563R, N 164Y, E323R and D362R; (8) E176R, K238R, T447R, E563R, N164Y, I926R, E323R and D362R; (9) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and I926G; (10) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G and L439(L+GG); (11) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G and L439(L+G);(12)E176R, K238R, T447R, E563R, N164Y and D958R;(13)E176R, K238R, T447R, E563R, I926R and D958R;(14)E176R, K238R, T447R, E563R, E323R, D362R and D958R;(15)N164Y, I926R and D958R;(16)N164Y, E323R, D362R and D958R;(17)E176R, K238R, T447R, E563R, N164Y, I926R and D958R;( 18) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and D958R; (19) E176R, K238R, T447R, E563R, N164Y, I926R, E323R, D362R and D958R; (20) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G and D958R; (21) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, L439(L+GG) and D958R;(22) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, L439(L+G) and D958R; (23) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and R857A; (24) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and N861A (25) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and K807A; (26) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and N848A; (27) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and R715A; (28) E176R, K238 R, T447R, E563R, N164Y, E323R, D362R and R719A; (29) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and K394A; (30) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and H357A; (31) E176R, K238R, T447R, E563R, N (32)Variations comprising any of the following groups: 164Y, E323R, D362R and K844A; (32)E176R, K238R, T447R, E563R, N164Y, E323R, D362R, R719A and K844A; or (33)Variations comprising any of the following groups: E176R, K238R, T447R, E563R, N164Y, E323R, D362R, R857A and K844A; of which, the amino acid positions are defined as the corresponding amino acid positions shown in SEQ ID NO:1.

[0028] 22. An engineered Cas12i nuclease having an amino acid sequence shown in any of SEQ ID NOs:2-24, or an engineered Cas12i nuclease according to any one of items 1 to 21, comprising an amino acid sequence having at least 80% identity with the amino acid sequence shown in any of SEQ ID NOs:2-24.

[0029] 23. An engineered Cas12i effector protein comprising an engineered Cas12i nuclease or a functional derivative thereof as described in any one of items 1 to 22; An engineered Cas12i effector protein wherein the engineered Cas12i nuclease or its functional derivative is enzymatically active, or the engineered Cas12i nuclease or its functional derivative is an enzymatically inactive variant.

[0030] 24. The engineered Cas12i effector protein according to item 23, wherein the Cas12i effector protein can induce double-strand or single-strand breaks in a DNA molecule.

[0031] 25. The engineered Cas12i nuclease or its functional derivative is an enzyme-inactive variant comprising one or more mutations of D599A, E833A, S883A, H884A, R900A, and D1019A; the amino acid position of which is defined as the corresponding amino acid position shown in SEQ ID NO:1, as described in item 23.

[0032] 26. An engineered Cas12i effector protein according to any one of claims 23 to 25, further comprising a functional domain that fuses with the engineered Cas12i nuclease or a functional derivative thereof.

[0033] 27. The functional domain is selected from one or more of the following: a translation initiation domain, a transcriptional inhibition domain, a transactivation domain, an epigenetic modification domain, a nucleic acid base editing domain, a reverse transcriptase domain, a reporter molecule domain, and a nuclease domain, in the engineered Cas12i effector protein described in item 26.

[0034] 28. The engineered Cas12i effector protein comprises a first polypeptide comprising the N-terminal portion of the engineered Cas12i nuclease or a functional derivative thereof, and a second polypeptide comprising the C-terminal portion of the engineered Cas12i nuclease or a functional derivative thereof, wherein the first polypeptide and the second polypeptide can associate with each other in the presence of a guide RNA containing a guide sequence to form a clustered, regularly spaced, short palindromic repeat sequence (CRISPR) complex that specifically binds to a target nucleic acid, the target nucleic acid comprising a target sequence complementary to the guide sequence; Preferably, the first polypeptide comprises N-terminal partial amino acid residues 1 to X of the engineered Cas12i nuclease described in any of items 1 to 22, and the second polypeptide comprises amino acid residues X+1 to the C-terminus of the engineered Cas12i nuclease described in any of items 1 to 22; Selectively, the first polypeptide and the second polypeptide each contain a dimerizing domain; An engineered Cas12i effector protein according to any one of claims 23 to 27, wherein the first dimer domain and the second dimer domain selectively associate with each other in the presence of an inducer.

[0035] 29. An engineered CRISPR-Cas12i system (a) an engineered Cas12i nuclease as described in any one of items 1 to 28, or an engineered Cas12i effector protein as described in any one of items 23 to 28; and (b) A guide RNA comprising a guide sequence complementary to the target sequence, or one or more nucleic acids encoding the guide RNA, Among these, the engineered Cas12i nuclease or engineered Cas12i effector protein and the guide RNA can form a CRISPR complex, the specific binding of the CRISPR complex includes the target nucleic acid of the target sequence and induces modification of the target nucleic acid; Preferably, the guide RNA is a crRNA containing the guide sequence; More preferably, the engineered CRISPR-Cas12i system includes a precursor guide RNA array encoding multiple crRNAs; Alternatively, preferably, an engineered CRISPR-Cas12i system in which the engineered Cas12i nuclease or engineered Cas12i effector protein is the main editor and the guide RNA is a prime editing guide RNA (pegRNA).

[0036] 30. comprising one or more vectors encoding the engineered Cas12i nuclease or engineered Cas12i effector protein; Preferably, the one or more vectors are selected from the group consisting of: retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated vectors, and herpes simplex vectors; More preferably, the one or more vectors are adeno-associated virus (AAV) vectors; More preferably, the AAV vector also encodes the guide RNA in the engineered CRISPR-Cas12i system described in Section 29.

[0037] 31. A method for detecting a target nucleic acid in a sample, (a) The sample is brought into contact with the engineered CRISPR-Cas12i system and tagged detection nucleic acid described in Section 29, wherein the detection nucleic acid is single-stranded and does not hybridize with the guide sequence of the guide RNA; and (b) A method for detecting a target nucleic acid, comprising measuring a detectable signal generated by the engineered Cas12i nuclease or engineered Cas12i effector protein cleaving the tagged detection nucleic acid.

[0038] 32. A method for modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with an engineered CRISPR-Cas12i system as described in item 29 or 30; Preferably, the method is carried out in vitro, ex vivo, or in vivo; More preferably, the target nucleic acid is present in the cell; More preferably, the cells are bacterial cells, yeast cells, mammalian cells, plant cells, or animal cells; More preferably, the target nucleic acid is genomic DNA; More preferably, the target sequence is associated with a disease or disorder; More preferably, the engineered CRISPR-Cas12i system comprises a precursor guide RNA array encoding multiple crRNAs, of which each crRNA comprises a different guide sequence.

[0039] 33. Use of an engineered CRISPR-Cas12i system as described in Section 29 or 30 in the manufacture of a drug for treating a disease or disorder related to a target nucleic acid in the cells of an individual; preferably, the disease or disorder is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, ophthalmological disease, bacterial infection, and viral infection.

[0040] 34. A method for treating a disease or disorder related to a target nucleic acid in the cells of an individual, the method comprising modifying the target nucleic acid in the cells of the individual by the method described in item 32, to treat the disease or disorder, preferably the disease or disorder being selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, ophthalmological disease, bacterial infection, and viral infection.

[0041] 35. A method for modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with an engineered CRISPR-Cas12i system as described in item 29 or 30.

[0042] 36. A composition or kit comprising an engineered Cas12i nuclease as described in any one of items 1 to 22, or an engineered Cas12i effector protein as described in any one of items 23 to 28.

[0043] 37. Engineered cells containing a target nucleic acid modified by the method described in section 32 or 35.

[0044] 38. An engineered non-human animal comprising one or more engineered cells as described in paragraph 37. Beneficial effects obtained by the technical invention of this application

[0045] The Cas12i nucleases and their effector proteins engineered in this application exhibit higher activity, including catalytic efficiency for cleaving nucleic acid substrates and intracellular gene editing efficiency. The engineered Cas12i nucleases in this application have superior gene editing efficiency in mammalian cells (e.g., human cells) compared to conventional Cas gene editing tools. For example, when several exemplary Cas12i2 nuclease variants in this application were tested for gene editing efficiency at multiple sites (e.g., 62 sites) in human cells, gene editing efficiency exceeded approximately 60% at 57 sites, and the average gene editing efficiency was close to 70%. In some embodiments, the engineered Cas12i nucleases and their effector proteins in this application also have one or more advantages, such as being small proteins (1054aa), having simple crRNA components, having simple PAM sequences, and the protein itself being able to process precursor crRNA. Furthermore, this application provides a Cas12i nuclease with lower off-target rates and higher specificity, further artificially modified based on a highly active engineered Cas12i nuclease (e.g., SEQ ID NO. 8). These advantages make the efficient engineered Cas12i nuclease and its effector protein of this application highly suitable for in vivo gene editing or gene regulation. [Brief explanation of the drawing]

[0046] [Figure 1a] By replacing the amino acids that interact with PAM in the reference Cas12i nuclease (wild-type Cas12i nuclease shown in SEQ ID NO. 1) with positively charged amino acids, gene editing efficiency is enhanced. As shown in the figure, the four mutants E176R, K238R, T447R, and E563R can significantly improve gene editing efficiency in human 293T cells. [Figure 1b]In Figure 1a, when amino acid mutations that significantly enhance gene editing efficiency (E176R, K238R, T447R, E563R) are combined, the resulting mutants exhibit higher gene editing efficiency in human 293T cells. [Figure 2] By substituting amino acids involved in the release of the DNA double strand in the reference Cas12i nuclease with aromatic ring-containing amino acids, the efficiency of gene editing can be enhanced. As shown in the figure, mutants such as Q163F, Q163Y, Q163W, N164F, and N164Y can significantly improve gene editing efficiency in human 293T cells. [Figure 3a-3c] By replacing the amino acid located in the RuvC domain of the reference Cas12i nuclease, which interacts with the single-stranded DNA substrate, with a positively charged amino acid, the efficiency of gene editing is enhanced. As shown in Figures 3a, 3b, and 3c, mutants such as E323R, L327R, V355R, G359R, G360R, D362R, N391R, Q424R, Q425R, N925R, I926R, and G929R can significantly improve gene editing efficiency in human 293T cells. [Figure 3d] When the point mutations with improved efficiency shown in Figures 3a and 3b were combined, the resulting mutants exhibited higher gene editing efficiency in human 293T cells. [Figure 3e] When point mutations that enhance efficiency, as shown in Figures 3a and 3b, were combined with modified mutations (L439(L+GG), I926G) based on the molecular flexibility principle, the combined mutants showed higher gene editing efficiency in human 293T cells. [Figure 4] By substituting positively charged amino acids in the reference Cas12i enzyme that interact with the DNA-RNA double helix, gene editing efficiency can be enhanced. As shown in Figure 4, mutants with G116R, E117R, T159R, S161R, E319R, E343R, and D958R can significantly improve gene editing efficiency in human 293T cells. Of these, the D958R mutation is optimal. [Figure 5]When high-efficiency mutants obtained from the three modification strategies shown in Figures 1a to 3e were combined with mutants modified based on the molecular flexibility principle (L439(L+GG), I926G), the combined mutants showed higher gene editing efficiency in human 293T cells. By combining them, gene editing efficiency can be significantly improved. The mutant with the best gene editing effect was selected, named CasXX, and used in subsequent experiments. CasXX (SEQ ID NO:8) is a Cas12i engineering enzyme with a combination of E176R+K238R+T447R+E563R+N164Y+E323R+D362R mutations (reference Cas12i2 based on amino acid sequence SEQ ID NO:1). [Figure 6a] This is a summary of the gene editing efficiency of CasXX at 62 locations in the human genome. PAM=NTTN. [Figure 6b] This is a comparison of the gene editing efficiency of CasXX, AsCas12a, and BhCas12b v4. [Figure 6c] This is a comparison of the gene editing efficiencies of CasXX, SpCas9, SaCas9, and SaCas9-KKH. [Figure 6d] This is a statistical analysis of the gene editing efficiency of CasXX in the mouse Hepa1-6 cell line. CasXX showed strong gene editing ability at 65 sites within mouse Hepa1-6 cells, and the average gene editing efficiency exceeded 60%. [Figure 7] Figure 7 shows the gene editing efficiency of CasXX at 64 human genome sites. Of these 64 sites, the PAM sequences included cover all NNNN combinations: NTTN, NTAN, NTCN, NTGN, NATN, NAAN, NACN, NAGN, NCTN, NCAN, NCCN, NCGN, NGTN, NGAN, NGCN, and NGGN. As shown in Figure 7, CasXX exhibits efficient gene editing at the PAM sequences NTTN, NTAN, NTCN, NTGN, NATN, NAAN, NCTN, NCAN, and NGTN, with an average gene editing efficiency exceeding 40% at these sites. [Figure 8]Figure 8 shows the results of in vitro cleavage of double-stranded DNA by wild-type Cas12i2 protein (SEQ ID NO:1) and CasXX protein. The wild-type Cas12i2 protein exhibits partial cleavage efficiency only for double-stranded DNA containing NTTN PAM, but has little cleavage activity for the remaining PAMs. However, the CasXX protein shows efficient cleavage efficiency for all double-stranded DNA containing NTTN, NTAN, NTCN, NTGN, NATN, NAAN, NACN, NCTN, NCAN, NGTN, and NGAN PAMs. It can almost completely cleave double-stranded DNA containing NTTN, NTAN, NTCN, NATN, NAAN, NACN, NCTN, NCAN, and NGTN PAMs. [Figure 9] This is a homology alignment of the amino acid sequences of Cas12i2 (SEQ ID NO:1) and Cas12i1 (SEQ ID NO:13). The amino acids shown in shaded areas are those that are the same for both Cas12i proteins, while the amino acids shown in white boxes are those that have similar properties for both Cas12i proteins. [Figure 10] Figure 10 shows the detection of off-target effects of CasXX at the EMX1-7 target site using GUIDE-Seq. The numbers following the sequences represent the reads measured for each sequence. The first row of sequences is the reference target sequence (SEQID NO: 77), and the sequences below represent the target sequence, the off-target sequence, and the reads with double-stranded DNA tags enriched in the target and off-target sequences, respectively. This figure shows that CasXX has off-target effects at the EMX1-7 site. [Figure 11] Figure 11 shows the process of adding a new point mutation to the original CasXX sequence or deleting the original point mutation to construct a new HF mutant. Experimental results showed that single-point mutants based on the CasXX sequences R857A, N861a, K807A, N848A, R715A, R719A, K394A, H357A, and K844A could effectively reduce the rate of insertions and deletions in EMX1-7-OT-1, EMX1-7-OT-2, and EMX1-7-OT-3. [Figure 12] Figure 12 shows the process of adding a new point mutation to the original CasXX sequence or deleting the original point mutation to construct a new HF variant. Experimental results showed that single-point mutants based on the CasXX sequences R857A, N861a, K807A, N848A, R715A, R719A, K394A, H357A, and K844A could effectively reduce the rate of insertions and deletions when using crRNA-RNF2-1-Mis-1 / 2, crRNA-RNF2-1-Mis-5 / 6, crRNA-RNF2-1-Mis-17 / 18, and crRNA-RNF2-1-Mis-19 / 20, which mimic off-target mutations. [Figure 13] Figure 13 shows the experimental flow for screening gene editing specificity variants (CasXX-based HF variants) using a fluorescence reporting system. 600 ng of a plasmid encoding the Cas protein, 300 ng of a plasmid encoding crRNA, and 100 ng of a plasmid encoding mCherry were transfected into cells cultured in a 24-well cell culture dish. After 3 days of transfection, the percentage of residual mCherry-positive cells in each sample was calculated using flow cytometry. The calculation formula is shown in the figure. [Figure 14] Figure 14 shows how CasXX and its variants edit the mCherry gene located on the plasmid in 293T cells. After plasmid transfection, the cells were cultured for 3 days and flow cytometry analysis was performed. The higher the gene editing efficiency of the Cas protein, the lower the percentage of mCherry-positive cells. The experimental flow corresponds to the schematic diagram of the experimental flow in Figure 13; spacer sequences for mCherry-FM, mCherry-Mis-1 / 2, mCherry-Mis-5 / 6, and mCherry-Mis-19 / 20 are shown as SEQ ID NO: 37, 38, 39, and 40, respectively. [Figure 15]Figure 15 shows the construction of new HF mutant combinations based on the original CasXX sequence. Experimental results showed that CasXX sequence-based mutants such as R857A, R719A, K394A, K844A, R719A / K844A, and R857A / K844A effectively reduced the rate of off-target insertions and deletions in EMX1-7-OT-1, EMX1-7-OT-2, and EMX1-7-OT-3 without sacrificing efficiency at the target site. [Figure 16] This figure shows the construction of new HF mutant combinations based on the original CasXX sequence. Experimental results showed that CasXX sequence-based mutants such as R857A, R719A, K394A, K844A, R719A / K844A, and R857A / K844A could effectively reduce the rate of insertions and deletions when using crRNA-RNF2-1-Mis-1 / 2, crRNA-RNF2-1-Mis-5 / 6, crRNA-RNF2-1-Mis-17 / 18, and crRNA-RNF2-1-Mis-19 / 20, which mimic off-target sites, without sacrificing efficiency at target sites. [Figure 17] Figure 17 shows the detection of off-target effects of CasXX and CasXX+K394A mutants at the CD34-7 target site using GUIDE-Seq. The number following the sequence represents the number of reads measured for each sequence. The first row of sequences is the reference target sequence (SEQID NO: 78), and the sequences below represent the target sequence, the off-target sequence, and the number of reads enriched with double-stranded DNA tags in the target and off-target sequences, respectively. This figure shows that CasXX exhibits off-target effects at the CD34-7 site, while the CasXX+K394A mutant does not exhibit off-target effects at this site. [Modes for carrying out the invention]

[0047] Note that several terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that different terms may be used for the same component. This specification and claims use functional differences in components as the criterion for classification, rather than using nominal differences as a way to distinguish components. The terms “contains” or “includes” as used throughout the specification and claims are open terms and should be interpreted as “includes, but not limited to.” The following descriptions in the specification are preferred embodiments for carrying out the invention, but these descriptions are intended for the general principles of the specification and not to limit the scope of the invention. The scope of protection of the invention is subject to those defined in the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as that commonly understood by those skilled in the art to which this disclosure belongs.

[0048] I. Terminology As used herein, “effector protein” refers to a protein that has activities such as site-directed binding activity, single-strand DNA cleavage activity, double-strand DNA cleavage activity, single-strand RNA cleavage activity, DNA or RNA modification (e.g., cleavage, base substitution, insertion, deletion) or transcriptional regulatory activity.

[0049] As used herein, “guide RNA” and “gRNA” refer to RNAs that are interchangeable and capable of forming complexes with Cas12i effector proteins and target nucleic acids (e.g., double-stranded DNA). This specification also considers precursor guide RNA arrays that can be processed into multiple crRNAs. A “crRNA” or “CRISPR RNA” includes a guide sequence that is sufficiently complementary to the target sequence of the target nucleic acid (e.g., double-stranded DNA) and directs the CRISPR complex to specifically bind to the target sequence of the target nucleic acid.

[0050] As used herein, the term “CRISPR array” refers to a nucleic acid (e.g., DNA) fragment containing CRISPR repeats and spacers, starting with the first nucleotide of the first CRISPR repeat and ending with the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeats. As used herein, the terms “CRISPR repeat,” “CRISPR direct repeat,” or “direct repeat” refer to multiple short direct repeat sequences in a CRISPR array that have very small or no sequence changes. Appropriately, direct repeats can form a stem-loop structure.

[0051] The terms “nucleic acid,” “polynucleotide,” and “nucleotide sequence” are used interchangeably to refer to polymeric forms of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and their analogues. “Oligonucleotide” and “Oligonucleotide” are used interchangeably to refer to short polynucleotides of less than approximately 50 nucleotides. As used herein, “complementarity” means the ability of a nucleic acid to form hydrogen bonds with another nucleic acid via conventional Watson-Crick base pairs. Complementarity percentages represent the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (i.e., Watson-Crick base pairs) with a second nucleic acid (e.g., 5 / 10, 6 / 10, 7 / 10, 8 / 10, 9 / 10, 9 / 10, 10 / 10, respectively). “Fully complementary” means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “basically complementary” means two nucleic acids that are at least about 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% complementary in a region of about 40, 50, 60, 70, 80, 100, 150, 200, 250 or more nucleotides, or that have been hybridized under stringent conditions.

[0052] As used herein, “stringent conditions” for hybridization refer to conditions under which nucleic acids complementary to the target sequence primarily hybridize with the target sequence and are not fundamentally hybridized with non-target sequences. Stringent conditions are usually sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature required for specific hybridization between the sequence and its target sequence. Non-limiting examples of stringent conditions are described in detail in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology - Hybridization With Nucleic Acid Probes, Part I, Chapter 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assay,” Elsevier, N, Y.

[0053] Hybridization refers to a reaction in which one or more polynucleotides react to form a complex, which is stabilized by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonds can occur through Woznieck base pairs, Hoogstein bonds, or any other sequence-specific method. A sequence that can hybridize with a given sequence is called a "complement" of that sequence.

[0054] For nucleic acid sequences, the "sequence identity percentage (%)" is defined as the percentage of nucleotides in a candidate sequence that are identical to a specific nucleic acid sequence after the sequence has been aligned by allowing gaps (if necessary) to achieve the maximum possible sequence identity percentage. For peptide, polypeptide, or protein sequences, the "sequence identity percentage (%)" is the percentage of substituted amino acid residues in a candidate sequence that are identical to an amino acid residue in a specific peptide or amino acid sequence after the sequence has been aligned by allowing gaps (if necessary) to achieve the maximum possible sequence identity percentage. For the purpose of determining the amino acid sequence identity percentage, for example, BLAST, BLAST-2, ALIGN, or MEGALIGN are used. TM This can be implemented in various ways within the scope of the art using commonly available computer software such as DNASTAR software. Those skilled in the art can determine appropriate parameters for measurement and alignment, including any algorithm necessary to achieve maximum alignment over the entire length of the sequences being compared.

[0055] The terms “polypeptide” and “peptide” are used interchangeably herein and refer to polymers of amino acids of any length. Such polymers may be linear or branched, may contain modified amino acids, or may be interrupted by non-amino acids. A protein may have one or more polypeptides. The term also includes modified amino acid polymers, such as those modified by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other operation (binding with a labeling component).

[0056] As used herein, “Variant” is interpreted as a polynucleotide or polypeptide that is different from a reference polynucleotide or polypeptide but retains the desired properties. A typical variant of a polynucleotide differs from the nucleic acid sequence of another reference polynucleotide. Changes in the variant nucleic acid sequence may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. Nucleotide changes may result in amino acid substitutions, additions, deletions, fusions, and truncations in the polypeptide encoded by the reference sequence, as described below. A typical variant of a polypeptide differs from another reference polypeptide in terms of its amino acid sequence. Generally, the differences are limited, and the sequences of the reference polypeptide and the variant are very similar overall and identical in many regions. The amino acid sequences of the variant and the reference polypeptide can differ by any combination of one or more substitutions, additions, and deletions. The substituted or inserted amino acid residues may or may not be amino acid residues encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring (e.g., allele variants) or unknown naturally occurring variants. Variants of polynucleotides and polypeptides that do not exist in nature can be produced by mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.

[0057] As used herein, the term “wild type” has the meaning commonly understood by those skilled in the art, meaning a typical form of an organism, strain, gene, or characteristic that is distinguishable from variants or offshoots when present in nature; isolated from natural resources and not intentionally modified.

[0058] As used herein, the terms “not naturally occurring” and “engineered” are interchangeable and mean artificial intervention. When these terms are used to describe a nucleic acid molecule or polypeptide, it means that such nucleic acid molecule or polypeptide does not contain, at least fundamentally, its natural binding or at least one other naturally occurring component.

[0059] As used herein, the term “orthologue” has the meaning generally understood by those skilled in the art. Further guidance, as used herein, an “orthologue” of a protein refers to a protein belonging to a different species that performs the same or similar function as the orthologue protein.

[0060] As used herein, the term “identity” is used to describe sequence matching between two polypeptides or two nucleic acids. If one position in two comparison sequences is occupied by the same base or amino acid monomer subunit (for example, one position in each of two DNA molecules is occupied by adenine, or one position in each of two polypeptides is occupied by lysine), then that position is the same for both molecules. The “identity percentage” between two sequences is a function of the number of matching positions common to the two sequences divided by the number of positions being compared, multiplied by 100. For example, if 6 out of 10 positions in two sequences match, the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGTT have 50% identity (3 out of 6 positions match). Typically, this comparison is performed to align the two sequences to achieve maximum identity. This comparison can be performed, for example, by the method described in Needleman et al. (1970) J. Mol. Biol. 48: 443-453, which can be easily carried out by a computer program such as the Align program (DNAstar, Inc.). A weighted residue table of PAM120 may also be used, and the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4: 11-17 (1988)) is integrated into the ALIGN program (version 2.0). To determine the percentage of identity between the two amino acid sequences, a gap length penalty of 12 and a gap penalty of 4 are used. Furthermore, to determine the percentage of identity between two amino acid sequences, the Needleman and Wunsch (J MoI Biol. 48:444-453 (1970)) algorithm can be used within the GAP program integrated into the GCG package (available from www.gcg.com), using a Blossum62 matrix or PAM250 matrix with gap weights of 16, 14, 12, 10, 8, 6, or 4 and length weights of 1, 2, 3, 4, 5, or 6.

[0061] As used herein, “cell” should be understood to refer not only to a specific single cell, but also to its offspring (progeny) or potential offspring. Such offspring may actually differ from the parent cell due to mutations or environmental influences, but they are included within the scope of the terminology herein.

[0062] As used herein, the terms “transduction” and “transfection” include methods of introducing DNA into cells using infectious agents known in the art (e.g., viruses) or other methods in order to express a protein or molecule of interest. In addition to viruses or virus-like reagents, there are chemical transfection methods such as transfection using calcium phosphate, dendritic polymers, liposomes or cationic polymers (e.g., DEAE-dextran or polyethyleneimine); non-chemical methods such as electroporation, cell squeezing, sonoporation, optical transfection, impalefection, plasma fusion, plasmid delivery or transposon; particle-based methods such as gene guns, magnetic transfection or magnet-assisted transfection, particle impact; and hybridization methods (e.g., nucleofection).

[0063] As used herein, the terms “transfected,” “transformed,” or “transduced” refer to the process of moving or introducing an exogenous nucleic acid into a host cell. A “transfected,” “transformed,” or “transduced” cell is a cell that has been transfected, transformed, or transduced with an exogenous nucleic acid.

[0064] The term "in vivo" refers to the inside of the organism from which cells are obtained. "In vitro" or "ex vivo" refers to the outside of this organism from which cells are obtained.

[0065] As used herein, “treatment / treating” is a method for obtaining a beneficial or desired outcome (including clinical outcomes). For the purposes of the present invention, beneficial or desired clinical outcomes include, but are not limited to, one or more of the following: relief of one or more symptoms of a disease; reduction of the severity of a disease; stabilization of a disease (e.g., prevention or delay of disease exacerbation); prevention or delay of disease spread (e.g., metastasis); prevention or delay of disease recurrence; reduction of disease recurrence rate; delay or delay of disease progression; improvement of a disease state; providing (partial or complete) relief of a disease; reducing the dose of one or more other drugs required to treat the disease; slowing disease progression; improving quality of life; and / or extending survival. “Treatment” also includes reducing a disease, a disease state, or the pathological consequences of a disease. The methods of the present invention have considered one or more of these forms of treatment.

[0066] As used herein, the term “effective dose” means an amount of compound or composition sufficient to treat a particular disorder, condition, or disease (e.g., improvement, relief, reduction, and / or delay of one or more symptoms). As understood in the art, “effective dose” may be administered in one or more doses, i.e., one or more doses may be required to achieve the desired therapeutic endpoint.

[0067] "Subject," "individual," or "patient" are interchangeable terms used herein and refer to any animal classified as a mammal for the purpose of achieving therapeutic objectives, including humans, livestock, and farm animals, such as dogs, horses, cats, and zoo, farm, or pet animals. In some embodiments, the individual is a human individual.

[0068] It should be understood that the embodiments of the present invention described herein include embodiments consisting of and / or essentially consisting of. References to "about" values ​​or parameters herein include (and are described) variations on the value or parameter itself. For example, a description of "about X" includes a description of "X".

[0069] Where used herein, a “no” value or parameter generally refers to a value or parameter other than… For example, “the method is not used for the treatment of type X cancer” means that it is used for the treatment of cancers other than type X.

[0070] As used herein, the term "approximately XY" has the same meaning as "approximately X to approximately Y".

[0071] As used herein and in the appended claims, the singular forms “one / one kind (a / an)” and “the said” include multiple objects unless the context specifically indicates otherwise. Also note that claims may be written to exclude any optional element; therefore, this statement is intended, in conjunction with the description of elements in the claims, as a priori basis for using exclusive terms such as “only” or “only,” or as a restriction for using “no.”

[0072] As used herein, the term "and / or" is intended to include both A and B in terms such as "A and / or B," A or B, A (alone), and B (alone). Similarly, as used herein, the term "and / or" is intended to include embodiments of the following in terms such as "A, B, and / or C": A, B, and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone), B (alone), and C (alone).

[0073] II. Cas12i nucleases and effector proteins Engineered Cas12i nuclease In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising one or more (e.g., two, three, four, or five) mutations based on a reference Cas12i nuclease, the mutations being as follows: (1) Substitution of the amino acid that interacts with PAM in the reference Cas12i nuclease with a positively charged amino acid (R, H, K, etc.); (2) Substituting the amino acids involved in the release of the DNA double strand in the reference Cas12i nuclease with aromatic ring amino acids (e.g., F, Y, or W); (3) Substitution of an amino acid located in the RuvC domain of the reference Cas12i nuclease that interacts with the single-stranded DNA substrate with a positively charged amino acid (e.g., R, H, or K); (4) Substituting the amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with positively charged amino acids (e.g., R, H, or K); and (5) Selected from substituting one or more polar or positively charged amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with hydrophobic amino acids (e.g., A, V, I, L, M, F, Y, P, C, or W). In some embodiments, the reference Cas12i nuclease is a natural Cas12i nuclease, such as the wild-type Cas12i2 nuclease having an amino acid sequence as shown in SEQ ID NO:1. In some embodiments, the reference Cas12i nuclease is a variant of the Cas12i nuclease, such as a natural variant. In some embodiments, the reference Cas12i nuclease is an engineered Cas12i (e.g., Cas12i2 or Cas12i1) nuclease that does not contain one or more mutations of (1) to (5) above. In some embodiments, the amino acid positions are defined by the corresponding amino acid positions shown in SEQ ID NO:1. In some embodiments, the amino acid position is defined by the amino acid position corresponding to SEQ ID NO:1 of a wild-type Cas12i nuclease having a different sequence from SEQ ID NO:1.

[0074] As used herein, "an amino acid is located at position X, where the amino acid position is defined by the corresponding amino acid position of the wild-type Cas12i nuclease shown in SEQ ID NO:1," or "the amino acid is located at position X such that it is defined by the amino acid position corresponding to SEQ ID NO:1 of a wild-type Cas12i nuclease having a different sequence from SEQ ID NO:1, and the amino acid residue is located at a position in the reference enzyme Cas12i, corresponding to position X of SEQ ID NO:1, and the amino acid sequences of the reference enzyme Cas12i and the amino acid sequence of SEQ ID NO:1 are aligned with each other based on sequence identity. For example, Figure 7 shows a homology comparison of the amino acid sequences of Cas12i2 (SEQ ID NO:1) and Cas12i1 (SEQ ID NO:13). A person skilled in the art can use software commonly used by a person skilled in the art, such as Clustal Omega, to compare the amino acid sequences of the reference Cas12i nucleases described in either of the above and the SEQ ID By performing sequence identity comparison and alignment with NO:1, it is possible to obtain the amino acid sites in the reference Cas12i nuclease corresponding to the amino acid sites defined based on SEQ ID NO:1 described in this application.

[0075] This application describes how introducing amino acid mutations based on any one or more combinations of the five modification principles described above can increase in vitro and / or in vivo enzyme activity (e.g., DNA single-strand or double-strand break activity) (e.g., gene cleavage efficiency can be increased by approximately 100 times), increase the number of PAMs that can be identified (e.g., as can be seen from Figure 8, the wild-type enzyme can identify PAMs, and the CasXX enzyme based on the above amino acid mutations can identify at least 11 types of PAMs), and / or reduce the off-target rate and increase specificity (e.g., as can be seen from Table 18, the off-target efficiency of HF-Cas against CasXX can be reduced by approximately 99%). The engineered Cas12i nuclease includes one or more specific mutations described in sections 1) to 6) below. In some embodiments, any one or more mutations described herein can be combined with existing Cas12i mutations (e.g., mutations described in section 7) below) to provide an engineered Cas12i nuclease with higher activity.

[0076] In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease includes a mutation that replaces one or more amino acids in a reference Cas12i nuclease that interact with PAM with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease includes a mutation that replaces one or more amino acids in a reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease includes a mutation that replaces one or more amino acids located in the RuvC domain of a reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising a mutation that replaces one or more amino acids that interact with the DNA-RNA double helix in a reference Cas12i nuclease with positively charged amino acids. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising a mutation that replaces one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in a reference Cas12i nuclease with hydrophobic amino acids.

[0077] In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that interact with PAM with a positively charged amino acid; and 2) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that interact with PAM with a positively charged amino acid; and 2) a mutation that replaces one or more amino acids located in the RuvC domain of a reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that are involved in the release of the DNA double helix with aromatic ring-bearing amino acids; and 2) a mutation that replaces one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid. In some embodiments, the engineered Cas12i nuclease also comprises a mutation that replaces one or more polar or positively charged amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with a hydrophobic amino acid. In some embodiments, the engineered Cas12i nuclease further comprises one or more mutations in the flexible region (e.g., substitution with G and / or insertion of one or two Gs thereafter) to increase the flexibility of the flexible region. In some embodiments, the Cas12i-engineered enzyme thus modified increases flexibility by at least about 10%, for example, at least about 20%, 30%, 50%, 100%, 150%, 200%, 500%, 1000%, or higher.

[0078] In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that interact with PAM with a positively charged amino acid; and 2) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that interact with the DNA-RNA double helix with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring-containing amino acid; and 2) a mutation that replaces one or more amino acids in a reference Cas12i nuclease that interact with the DNA-RNA double helix with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation in which one or more amino acids located in the RuvC domain of a reference Cas12i nuclease and interacting with a single-stranded DNA substrate are replaced with positively charged amino acids; and 2) a mutation in which one or more amino acids interacting with the DNA-RNA double helix in the reference Cas12i nuclease are replaced with positively charged amino acids. In some embodiments, the engineered Cas12i nuclease also comprises a mutation in which one or more polar or positively charged amino acids interacting with the DNA-RNA double helix in the reference Cas12i nuclease are replaced with hydrophobic amino acids. In some embodiments, the engineered Cas12i nuclease further comprises one or more mutations in the flexible region to increase the flexibility of the flexible region.

[0079] In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with PAM with a positively charged amino acid; 2) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring-containing amino acid; and 3) a mutation that replaces one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with PAM with a positively charged amino acid; 2) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring-containing amino acid; and 3) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with a positively charged amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease comprising: 1) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with PAM with a positively charged amino acid; 2) a mutation that replaces one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid; and 3) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with a DNA-RNA double helix with a positively charged amino acid.In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease includes: 1) a mutation that replaces one or more amino acids involved in the release of the DNA double helix in the reference Cas12i nuclease with an aromatic ring-containing amino acid; 2) a mutation that replaces one or more amino acids located in the RuvC domain of the reference Cas12i nuclease and interacting with a single-stranded DNA substrate with a positively charged amino acid; and 3) a mutation that replaces one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with a positively charged amino acid. In some embodiments, the engineered Cas12i nuclease also includes a mutation that replaces one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with a hydrophobic amino acid. In some embodiments, the engineered Cas12i nuclease further includes one or more flexible region mutations to enhance the flexibility of the flexible region.

[0080] In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease includes: 1) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with PAM with a positively charged amino acid; 2) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring amino acid; 3) a mutation that replaces one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid; and 4) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with a positively charged amino acid. In some embodiments, the engineered Cas12i nuclease also includes a mutation that replaces one or more polar or positively charged amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with a hydrophobic amino acid. In some embodiments, an engineered Cas12i nuclease is provided, the engineered Cas12i nuclease includes: 1) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with PAM with a positively charged amino acid; 2) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that are involved in the release of the DNA double helix with an aromatic ring amino acid; 3) a mutation that replaces one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with a single-stranded DNA substrate with a positively charged amino acid; 4) a mutation that replaces one or more amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with a positively charged amino acid; and 5) a mutation that replaces one or more polar or positively charged amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with a hydrophobic amino acid. In some embodiments, the engineered Cas12i nuclease further includes one or more flexible region mutations to enhance the flexibility of the flexible region.

[0081] In some embodiments, the engineered Cas12i nuclease contains a mutation at the corresponding amino acid position shown in SEQ ID NO:1: N164Y+E176R+K238R+E323R+D362R+T447R+E563R (hereafter referred to as "CasXX"). In some embodiments, the engineered Cas12i nuclease contains the sequence of SEQ ID NO:8.

[0082] 1) Substitution of the amino acid that interacts with PAM in the reference Cas12i nuclease with a positively charged amino acid. In some embodiments, the engineered Cas12i nuclease includes one or more mutations based on a reference Cas12i nuclease (e.g., Cas12i2) for substituting amino acids that interact with PAM in the reference Cas12i nuclease with positively charged amino acids such as R, H, or K. In some embodiments, the engineered Cas12i nuclease includes one, two, three, four, five, six, seven, eight or more of the aforementioned amino acid residue substitutions.

[0083] In some embodiments, the amino acid interacting with the PAM is an amino acid that is within 9 Å of the PAM in three-dimensional structure, for example, an amino acid within 9 Å of the PAM, an amino acid within 8 Å of the PAM, an amino acid within 7 Å of the PAM, an amino acid within 6 Å of the PAM, an amino acid within 5 Å of the PAM, an amino acid within 4 Å of the PAM, an amino acid within 3 Å of the PAM, an amino acid within 2 Å of the PAM, or an amino acid that is closer.

[0084] The spatial distance between PAM and amino acids is defined by the distance between atoms in the 3D structure (PDB file) of the analyzed Cas protein-RNA-DNA three-dimensional complex, and the interatomic distances can be displayed by PDB file identification software. In some embodiments, the spatial distance between PAM and amino acids is defined by the minimum distance between an amino acid residue and an atom contained in a nucleotide. Programs or PDB file identification software that can be used to measure the spatial distance between PAM and amino acids are widely known in the art and include, but are not limited to, PyMOL, ChimeraX, Swiss-pdbviewer, etc.

[0085] In some embodiments, the mutation that replaces one or more amino acids that interact with PAM in the reference Cas12i nuclease with a positively charged amino acid is a mutation of an amino acid at one or more of the positions of 176, 178, 226, 227, 229, 237, 238, 264, 447, and 563. In some embodiments, the mutation that replaces one or more amino acids that interact with PAM in the reference Cas12i nuclease with a positively charged amino acid is a mutation of one or more of the amino acids E176, E178, Y226, A227, N229, E237, K238, K264, T447, and E563. In some embodiments, the mutation that replaces one or more amino acids that interact with PAM in the reference Cas12i nuclease with a positively charged amino acid is a mutation of one or more of the amino acids E176, K238, T447, and E563. In some embodiments, the mutation that replaces the PAM-interacting amino acid in the reference Cas12i nuclease with a positively charged amino acid is located at amino acid residue 563, such as E563. In some embodiments, the amino acid position is defined by the corresponding amino acid position in the wild-type Cas12i nuclease shown in SEQ ID NO:1.

[0086] In some embodiments, the engineered Cas12i nuclease comprises mutations in one or more amino acids from E176R, E178R, Y226R, A227R, N229R, E237R, K238R, K264R, T447R, and E563R, and is defined by the corresponding amino acid position of the wild-type Cas12i nuclease shown in SEQ ID NO:1.

[0087] In the context of this specification, E176 means the 176th amino acid E (glutamic acid) in the cited amino acid sequence (for example, for SEQ ID NO: 1), of which common amino acids and their three-letter and one-letter abbreviations are described below: alanine Ala A, arginine Arg R, asparagine Asp D, cysteine ​​Cys C, glutamine Gln Q, glutamic acid Glu E, histidine HisH, isoleucine Ile I, glycine Gly G, asparagine Asn N, leucine Leu L, lysine Lys K, methionine Met M, phenylalanine Phe F, proline ProP, serine Ser S, threonine Thr T, leucine Trp W, tyrosine Tyr Y, valine Val V.

[0088] In some embodiments, the mutation that replaces the amino acid interacting with PAM in the reference Cas12i nuclease with a positively charged amino acid is to replace the corresponding amino acid residue in the reference Cas12i nuclease with R, H, or K, for example, R or K. In some embodiments, the mutation that replaces the amino acid interacting with PAM in the reference Cas12i nuclease with a positively charged amino acid is to replace the corresponding amino acid residue in the reference Cas12i nuclease with R.

[0089] In some embodiments, the engineered Cas12i nuclease includes one or more amino acid mutations from 176R, 238R, 447R, and 563R, where the amino acid position is defined by the corresponding amino acid position of the wild-type Cas12i nuclease shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes one or more mutations based on the reference Cas12i nuclease: E176, K238, T447, and E563, where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes one or more mutations based on the reference Cas12i nuclease: E176R, K238R, T447R, and E563R, where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease contains the E563R mutation, of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, for the purpose of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 85% (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with the above-mentioned engineered Cas12i nuclease (in which one or more amino acids that interact with PAM in the reference Cas12i nuclease are replaced with positively charged amino acids) may also be used.

[0090] In the context of this specification, E176R means that in the cited amino acid sequence, the 176th amino acid E (glutamic acid) is replaced with R (arginine).

[0091] In some embodiments, the engineered Cas12i nuclease includes any mutation or combination of mutations at the amino acid residue positions of (i) 176, 238, 264, 447, 563, 176+238, 176+447, 176+563, 238+447, 238+563, 447+563, 176+238+447, 176+238+563, 176+447+563, 238+447+563, 176+238+447+563, where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, substituting one or more amino acids that interact with PAM in the reference Cas12i nuclease with positively charged amino acids includes substituting R, H, or K, for example R or K, preferably R. In some embodiments, the engineered Cas12i nuclease includes any mutation or combination of mutations in any of the following amino acid residues: E176, K238, E264, T447, E563, E176+K238, E176+T447, E176+E563, K238+T447, K238+E563, T447+E563, E176+K238+T447, E176+K238+E563, E176+T447+E563, K238+T447+E563, and E176+K238+T447+E563; of which the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0092] In some embodiments, the engineered Cas12i nuclease includes any of the following mutation / mutation combinations: E176R, K238R, E264R, T447R, E563R, E176R+K238R, E176R+T447R, E176R+E563R, K238R+T447R, K238R+E563R, T447R+E563R, E176R+K238R+T447R, E176R+K238R+E563R, E176R+T447R+E563R, K238R+T447R+E563R, and E176R+K238R+T447R+E563R; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes any of the following mutation / mutation combinations: E563R, E176R+T447R, E176R+E563R, K238R+E563R, E176R+K238R+T447R, E176R+K238R+E563R, E176R+T447R+E563R, and E176R+K238R+T447R+E563R; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. For the purpose of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 85% (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with the aforementioned engineered Cas12i nuclease (in which one or more amino acids that interact with PAM in the reference Cas12i nuclease are replaced with positively charged amino acids) may also be used.

[0093] 2) Replacing the amino acids involved in the release of the DNA double helix in the reference Cas12i nuclease with aromatic ring-containing amino acids. In some embodiments, the engineered Cas12i nuclease includes mutations based on one or more reference Cas12i nucleases (e.g., Cas12i2) that substitute one or more amino acids involved in the unwinding of the DNA double strand in the reference Cas12i nuclease with aromatic ring amino acids (such as F, Y, or W). In some embodiments, the engineered Cas12i nuclease includes substitutions of one, two, three, four, five, six, or more of the aforementioned amino acid residues.

[0094] Of these, one or more amino acids involved in the release of the DNA double helix are amino acids that interact with the last base pair at the 3' end of the PAM relative to the target strand. For example, the PAM sequence identified by Cas12i2 is a 5'-NTTN-3' base pair, and the base pair formed between the 3' end N base in the PAM sequence and the target strand is the "last base pair at the 3' end of the PAM relative to the target strand," and the sequence following this base pair is the target site sequence.

[0095] In some embodiments, one or more amino acids involved in the uncoupling of the DNA double helix are located at positions 163 and / or 164; of which, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, one or more amino acids involved in the uncoupling of the DNA double helix are one or more amino acids at Q163 and N164; of which, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the amino acid involved in the uncoupling of the DNA double helix is ​​N164; of which, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0096] In some embodiments, the amino acid involved in the unwinding of the DNA double helix is ​​substituted with F, Y, or W. In some embodiments, the amino acid involved in the unwinding of the DNA double helix is ​​substituted with F. In some embodiments, the amino acid involved in the unwinding of the DNA double helix is ​​substituted with Y.

[0097] In some embodiments, the engineered Cas12i nuclease comprises mutations in one or more amino acid residues of 163F, 163Y, 163W, 164W, 164F, or 164Y; of which, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0098] In some embodiments, the engineered Cas12i nuclease includes either the Q163 and / or N164 mutation; of which the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes either the Q163F, Q163Y, Q163W, N164W, N164F, or N164Y mutation; of which the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes either the Q163F, Q163Y, Q163W, N164F, or N164Y mutation; of which the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes either the N164Y or N164F mutation; of which the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease contains an N164Y mutation; of which, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, to achieve the objective of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 85% (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with the above-mentioned engineered Cas12i nuclease (in which one or more amino acids involved in the release of the DNA double helix in the reference Cas12i nuclease are replaced with aromatic ring-bearing amino acids) may also be used.

[0099] 3) Substituting a positively charged amino acid located in the RuvC domain of the reference Cas12i nuclease that interacts with the single-stranded DNA substrate. In some embodiments, the engineered Cas12i nuclease comprises one or more mutations based on a reference Cas12i nuclease (e.g., Cas12i2), wherein the mutations replace an amino acid located in the RuvC domain of the reference Cas12i nuclease that interacts with a single-stranded DNA substrate with a positively charged amino acid (such as R, H, or K). In some embodiments, the engineered Cas12i enzyme comprises the substitution of one, two, three, four, five, six, or more amino acid residues.

[0100] Of these, one or more amino acids located in the RuvC structural region and interacting with the single-stranded DNA substrate are amino acids that are within 9 Å of the single-stranded DNA substrate in the three-dimensional structure. For example, they may be amino acids within 8 Å of the single-stranded DNA substrate in the three-dimensional structure, within 7 Å of the single-stranded DNA substrate in the three-dimensional structure, within 6 Å of the single-stranded DNA substrate in the three-dimensional structure, within 5 Å of the single-stranded DNA substrate in the three-dimensional structure, within 4 Å of the single-stranded DNA substrate in the three-dimensional structure, within 3 Å of the single-stranded DNA substrate in the three-dimensional structure, within 2 Å of the single-stranded DNA substrate in the three-dimensional structure, or amino acids that are closer.

[0101] The RuvC domain is the enzymatically active domain in the Cas12i protein responsible for cleaving single-stranded or double-stranded DNA. In the protein's primary sequence, the RuvC domain of Cas12i is divided into three parts: RuvC-1, RuvC-2, and RuvC-3. These three parts are adjacent in the three-dimensional structure and together constitute a catalytic pocket with enzymatic cleavage activity. For a description of the three-dimensional crystal structure, domain composition, and interaction with DNA substrates of Cas12i2, see Huang X. et al., Nature Communications, 11, Article number: 5241 (2020). For a description of the three-dimensional crystal structure, domain composition, and interaction with DNA substrates of Cas12i1, see Zhang H. et al., Nature Structural & Molecular Biology 27, 1069-1076 (2020). Through homology structure comparison and modeling, a three-dimensional structural model of the interaction between a reference Cas12i and a substrate can be obtained through known Cas12i three-dimensional crystal structures. Example 3 describes a modeling method for obtaining amino acids located in the RuvC domain of Cas12i2 and within 9 Å of a single-stranded DNA substrate.

[0102] In some embodiments, the spatial distance between an amino acid in the RuvC domain and a single-stranded DNA substrate can be defined by the interatomic distance in the 3D structure (PDB file) of the analyzed Cas protein-RNA-DNA three-dimensional complex, and the interatomic distance can be displayed by PDB file identification software. In some embodiments, the spatial distance between an amino acid in the RuvC domain and a single-stranded DNA substrate is defined by the minimum distance between an amino acid residue and an atom contained in a nucleotide. Programs or PDB file identification software that can be used to measure the spatial distance between an amino acid in the RuvC domain and a single-stranded DNA substrate are widely known in the art and include, but are not limited to, PyMOL, ChimeraX, and Swiss-pdbviewer.

[0103] In some embodiments, the one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are one or more amino acids at positions 323, 327, 355, 359, 360, 361, 362, 388, 390, 391, 392, 393, 414, 417, 418, 421, 424, 425, 650, 652, 653, 696, 705, 708, 709, 751, 752, 755, 840, 848, 851, 856, 885, 897, 925, 926, 928, 929, 932, and 1022. In some embodiments, one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are one or more of the following amino acids: E323, L327, V355, G359, G360, K361, D362, L388, N390, N391, F392, K393, Q414, L417, L418, K421, Q424, Q425, S650, E652, G653, I696, K705, K708, E709, L751, S752, E755, N840, N848, S851, A856, Q885, M897, N925, I926, T928, G929, Y932, and A1022. In some embodiments, the one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are one or more amino acids from E323, D362, L388, N391, L417, Q424, Q425, N925, I926, and G929. In some embodiments, the amino acid positions are defined by the corresponding amino acid positions of the wild-type Cas12i nuclease shown in SEQ ID NO:1.

[0104] In some embodiments, the engineered Cas12i nuclease includes a mutation in which one or more amino acids located in the RuvC domain of the reference Cas12i nuclease and interacting with a single-stranded DNA substrate are replaced with R, H, or K (e.g., R or K).

[0105] In some embodiments, the engineered Cas12i nucleases are N390R, N391R, F392R, L751R, E755R, N840R, N848R, S851R, A856R, Q885R, M897R, I926R, G929R, Y932R, E323R, L327R, V355R, G359R, G360R, K361R, D362R, Q414R, K421R, Q425R, S650R, E652R, K705R, K708R, E709R, S752R, N925R, The amino acid positions include one or more mutations or combinations of mutations of T928R, E323R+D362R, E323R+Q425R, E323R+I926R, Q425R+I926R, E323R+D362R+Q425R+I926R, E323R+Q425R+I926R, E323R+D362R+Q425R+I926R, D362R+I926R, N925R+I926R, D362R+N925R+I926R, and D362R+N925R; the amino acid positions are defined by the corresponding amino acid positions of the wild-type Cas12i nuclease shown in SEQ ID NO:1.

[0106] In some embodiments, the engineered Cas12i nuclease includes a mutation or combination of mutations at the position of one of the amino acid residues E323, D362, Q425, N925, I926, or G929; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes a mutation or combination of mutations at the position of any of the following amino acid residues: E323, D362, Q425, N925, I926, E323+D362, E323+Q425, E323+I926, D362+Q425, D362+N925, D362+I926, Q425+I926, N925+I926, E323+D362+Q425, E323+D362+I926, E323+Q425+I926, D362+N925+I926, D362+Q425+I926, E323+D362+Q425+I926; of which the amino acid position number is the SEQ ID It is defined by the corresponding amino acid position shown in NO:1. In some embodiments, the mutation is a mutation that replaces the amino acid residue at the said position with R, H, or K (e.g., R). In some embodiments, the engineered Cas12i nuclease comprises any of the amino acids or combinations of amino acids 323R, 362R, 425R, 925R, 926R, 323R+362R, 323R+425R, 323R+926R, 362R+425R, 362R+926R, 425R+926R, 925R+926R, 323R+362R+425R, 323R+362R+926R, 323R+425R+926R, 362R+925R+926R, and 362R+425R+926R; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease comprises any or a combination of mutations of E323R, D362R, Q424R, Q425R, N925R, I926R, and G929R; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1.In some embodiments, the engineered Cas12i nucleases are E323R, D362R, Q425R, N925R, I926R, E323R+D362R, E323R+Q425R, E323R+I926R, D362R+Q425R, Q425R+I926R, D362R+I926R, N925R+I926R, E3 The engineered Cas12i nuclease includes any of the following mutations or combinations of mutations: 23R+D362R+Q425R, E323R+D362R+I926R, E323R+Q425R+I926R, D362R+N925R+I926R, D362R+Q425R+I926R, and E323R+D362R+Q425R+I926R; of which, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes the I926R mutation; of which, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes the E323R+D362R mutation; of which, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, to achieve the objective of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 85% (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with the above-described engineered Cas12i nuclease (in which the RuvC domain of the reference Cas12i nuclease is positioned and one or more amino acids that interact with the single-stranded DNA substrate are replaced with positively charged amino acids) may also be used.

[0107] 4) Replacing one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with positively charged amino acids. In some embodiments, the engineered Cas12i nuclease includes mutations based on one or more reference Cas12i nucleases (e.g., Cas12i2) to replace one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with positively charged amino acids such as R, H, or K. In some embodiments, the engineered Cas12i enzyme includes substitutions of one, two, three, four, five, six, or more of the aforementioned amino acid residues.

[0108] Of these, one or more amino acids that interact with the DNA-RNA double helix are amino acids that are within 9 Å of the DNA-RNA double helix in the three-dimensional structure, for example, amino acids that are within 8 Å of the DNA-RNA double helix in the three-dimensional structure, amino acids that are within 7 Å of the DNA-RNA double helix in the three-dimensional structure, amino acids that are within 6 Å of the DNA-RNA double helix in the three-dimensional structure, amino acids that are within 5 Å of the DNA-RNA double helix distance in the three-dimensional structure, amino acids that are within 4 Å of the DNA-RNA double helix distance in the three-dimensional structure, amino acids that are within 3 Å of the DNA-RNA double helix distance in the three-dimensional structure, amino acids that are within 2 Å of the DNA-RNA double helix distance in the three-dimensional structure, or amino acids that are closer to the DNA-RNA double helix distance in the three-dimensional structure. The operating principle of several Cas nucleases is as follows: Cas forms a complex with a guide RNA such as crRNA, and within this complex, the crRNA and target DNA pair up to form a DNA-RNA double helix. This then interacts with the Cas nuclease, releasing the double-stranded target DNA, forming an R-loop, and completing the cleavage of dsDNA at the Cas enzymatic cleavage active site. For a description of the three-dimensional crystal structure of Cas12i2, its domain composition, and its interaction with the DNA-RNA double helix, see Huang X. et al., Nature Communications, 11, Article number: 5241 (2020).

[0109] In some embodiments, the spatial distance between the DNA-RNA double helix and the Cas amino acid can be defined by the minimum distance between the amino acid residue and the atom contained in the nucleotide in the 3D structure (PDB file) of the analyzed Cas protein-RNA-DNA three-dimensional complex, and the interatomic distance can be displayed by PDB file identification software. In some embodiments, the spatial distance between the Cas amino acid and the DNA-RNA double helix is ​​defined by the distance from the hypothetical atomic position. Programs or PDB file identification software that can be used to measure the spatial distance between the Cas amino acid and the DNA-RNA double helix are widely known in the art and include, but are not limited to, PyMOL, ChimeraX, Swiss-pdbviewer, and others.

[0110] In some embodiments, one or more amino acids that interact with the DNA-RNA double helix are one or more amino acids at positions 116, 117, 156, 159, 160, 161, 247, 293, 294, 297, 301, 305, 306, 308, 312, 313, 316, 319, 320, 343, 348, 349, 427, 433, 438, 441, 442, 679, 683, 691, 782, 783, 797, 800, 852, 853, 855, 861, 865, 957, and 958. In some embodiments, one or more amino acids that interact with the DNA-RNA double helix are one or more of the following amino acids: G116, E117, A156, T159, E160, S161, E247, G293, E294, N297, T301, I305, K306, T308, N312, F313, Q316, E319, Q320, E348, E348, E349, D427, K433, V438, N441, Q442, N679, E683, E691, D782, E783, E797, E800, M852, D853, L855, N861, Q865, S957, D958. In some embodiments, the one or more amino acids that interact with the DNA-RNA double helix are one or more amino acids from G116, E117, T159, S161, E319, E343, or D958. In some embodiments, the amino acid that interacts with the DNA-RNA double helix is ​​D958. In some embodiments, the amino acid position is defined by the corresponding amino acid position of the wild-type Cas12i nuclease shown in SEQ ID NO:1.

[0111] In some embodiments, the engineered Cas12i nuclease includes mutations that substitute one or more amino acids in the reference Cas12i nuclease that interact with the DNA-RNA double helix with R, H, or K (e.g., R or K).

[0112] In some embodiments, the engineered Cas12i nucleases are G116R, E117R, A156R, T159R, S161R, T301R, I305R, K306R, T308R, N312R, F313R, D427R, K433R, V438R, N441R, Q442R, M852R, L855R, N861R, Q865R It includes one or more amino acid mutations of E160R, Q316R, E319R, Q320R, E247R, E343, E348R, E349R, N679R, E683R, E691R, D782R, E783R, E797R, E800R, D853R, S957R, D958R, G293R, E294R, and N297R; the amino acid position is defined by the corresponding amino acid position of the wild-type Cas12i nuclease shown in SEQ ID NO:1.

[0113] In some embodiments, the engineered Cas12i nuclease includes a mutation or combination of mutations at any of the amino acid residue positions G116, E117, T159, S161, E319, E343, or D958; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the mutation is a mutation that replaces the amino acid residue at the said position with R, H, or K (e.g., R). In some embodiments, the engineered Cas12i nuclease includes an amino acid or combination of amino acid mutations at any of the sites 116R, 117R, 159R, 161R, 319R, 343R, or 958R; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes one of the mutations or combinations of mutations G116R, E117R, T159R, S161R, E319R, E343R, or D958R; therein, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the engineered Cas12i nuclease includes the D958R mutation; therein, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, to achieve the objective of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 85% (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with the engineered Cas12i nuclease described above may also be used.

[0114] 5) Replacing one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with hydrophobic amino acids. In some embodiments, the engineered Cas12i nuclease includes mutations based on one or more reference Cas12i nucleases, which replace one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with hydrophobic amino acids (such as A, V, I, L, M, F, Y, P, C, or W). The mutations (of which also refer to “high specificity” or “HF” mutations) can reduce the off-target rate of the Cas12i nuclease (i.e., increase specificity). In some embodiments, the off-target rate of CasXX-HF compared to CasXX can be reduced by at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or more. In some embodiments, the engineered Cas12i enzyme includes one, two, three, four, five, six, or more of the above amino acid residue substitutions.

[0115] Of these, one or more amino acids that interact with the DNA-RNA double helix are amino acids that are within 9 Å of the DNA-RNA double helix in the three-dimensional structure. For example, these may be amino acids that are within 8 Å of the DNA-RNA double helix in the three-dimensional structure, amino acids that are within 7 Å of the DNA-RNA double helix in the three-dimensional structure, amino acids that are within 6 Å of the DNA-RNA double helix in the three-dimensional structure, amino acids with a double helix distance of 5 Å or less from the DNA-RNA in the three-dimensional structure, amino acids with a double helix distance of 4 Å or less from the DNA-RNA in the three-dimensional structure, amino acids with a double helix distance of 3 Å or less from the DNA-RNA in the three-dimensional structure, amino acids with a double helix distance of 2 Å or less from the DNA-RNA in the three-dimensional structure, or amino acids that are closer. For a description of the three-dimensional crystal structure of Cas12i2, its domain composition, and its interaction with the DNA-RNA double helix, see Huang X. et al., Nature Communications, 11, Article number: 5241 (2020).

[0116] In some embodiments, the one or more polar or positively charged amino acids that interact with the DNA-RNA double helix are one or more amino acids at positions 119, 164, 297, 308, 309, 312, 346, 357, 394, 395, 402, 441, 433, 565, 715, 719, 766, 782, 807, 841, 844, 845, 848, 857, 861, and 865. In some embodiments, one or more polar or positively charged amino acids that interact with the DNA-RNA double helix are one or more amino acids from Y119, Y164, N297, T308, R309, N312, S346, H357, K394, E395, R402, N441, K433, S565, R715, R719, S766, D782, K807, N841, K844, K845, N848, R857, N861, and Q865. In some embodiments, one or more polar or positively charged amino acids that interact with the DNA-RNA double helix are one or more amino acids from R857, R861, K807, N848, R715, R719, K394, H357, and K844. In some embodiments, the one or more polar or positively charged amino acids that interact with the DNA-RNA double helix are one or more amino acids R857, R719, K394, and K844. Of these, the amino acid positions are defined as shown above by the corresponding amino acid positions indicated by SEQ ID NO: 1 or 8.

[0117] In some embodiments, the engineered Cas12i nuclease comprises one or more mutations of S565A, N297A, Q865A, T308A, R309A, N312A, N441A, R857A, N861A, Y119F, K433A, K807A, N841A, N848A, K845A, D782A, R715A, R719A, S766A, K394A, H357A, K844A, E395A, S346A, R402A, and YY164N; of which, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8.

[0118] In some embodiments, the engineered Cas12i nuclease includes a mutation that replaces one or more polar or positively charged amino acids located at 357, 394, 715, 719, 807, 844, 848, 857, and 861 in the reference Cas12i nuclease with a hydrophobic amino acid. In some embodiments, the hydrophobic amino acid is selected from A, V, L, I, P, and F, for example, A, V, L, or I. In some embodiments, the engineered Cas12i nuclease includes a mutation that replaces one or more polar or positively charged amino acids located at H357, K394, R715, R719, K807, K844, N848, R857, and / or R861 in the reference Cas12i nuclease with A. Of these, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8. In some embodiments, the engineered Cas12i nuclease includes mutations in an amino acid or combination of amino acids at any of the following positions: H357A, K394A, R715A, R719A, K807A, K844A, N848A, R857A, and R861A; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8.

[0119] In some embodiments, the engineered Cas12i nuclease includes a mutation or combination of mutations at any of the amino acid residue positions 857, 719, 394, or 844, where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, the mutation is a mutation that replaces the aforementioned positional amino acid residue with a hydrophobic amino acid (e.g., A). In some embodiments, the engineered Cas12i nuclease includes a mutation of an amino acid or combination of amino acids at any of the positions R857, R719, K394, or K844; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1 or 8. In some embodiments, the engineered Cas12i nuclease includes any mutation or combination of mutations at the position of any of the following amino acid residues: R857, R719, K394, K844, R719+K394, K394+K844, R857+K394, R719+K844, R857+R719+K394+K844, R857+R719+K394, R857+K394+K844, R857+R719+K394+K844; where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8; the mutation is a mutation that replaces a positional amino acid residue with a hydrophobic amino acid (e.g., A). In some embodiments, the engineered Cas12i nuclease includes any mutation or combination of mutations at the position of any of the following amino acid residues: R857A, R719A, K394A, K844A, R719A+K394A, K394A+K844A, R857A+K394A, R719A+K844A, R857A+R719A, R857A+K844A, R857A+R719A+K394A+K844A, R857A+R719A+K394A+K844A; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8.In some embodiments, the engineered Cas12i nuclease includes one or a combination of the R857A, R719A, K394A, or K844A mutations; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8. In some embodiments, the engineered Cas12i nuclease includes the K844A mutation; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8. In some embodiments, the engineered Cas12i nuclease includes the R719A and K844A mutations; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8. In some embodiments, the engineered Cas12i nuclease includes the R857A and K844A mutations; of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO: 1 or 8. In some embodiments, to achieve the objective of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 85% (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with the above-described engineered Cas12i nuclease (in which one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease are replaced with hydrophobic amino acids) may also be used.

[0120] 6) Other mutations Unless otherwise specified, the mutations described herein may include one or more insertions, deletions, and substitutions, and may be mutations of one or more amino acids.

[0121] Any or more of the mutations described in sections 1) to 5) can be combined with any or more known mutations that increase Cas12i activity, such as target binding, double-strand break activity, nickase activity, and / or gene editing activity. Exemplary row mutations are found, for example, in the literature PCT / CN2020 / 0134249 and CN112195164A, which are incorporated herein by general reference. Any or more of the mutations described in sections 1) to 5) can also be combined with any or more known mutations that decrease Cas12i activity, such as target binding, double-strand break activity, nickase activity, and / or gene editing activity.

[0122] In some embodiments, the engineered Cas12i nuclease (e.g., an engineered Cas12i nuclease comprising one or more mutations in sections (1) to (5) above) further comprises one or more flexible region mutations, the mutations increasing the flexibility of the flexible region in the reference Cas12i nuclease (or an engineered Cas12i nuclease comprising one or more mutations in any one of sections (1) to (5)) (e.g., at least about 10%, 20%, 50%, 60%, 70%, 80%, 90%, 1x, 1.1x, 1.2x, 1.5x, 2x, 3x, 4x, 5x, 10x, 20x, 50x, 100x or more). The flexible region in the reference Cas12i nuclease can be determined using any method known in the art. In some embodiments, multiple flexible regions are determined based solely on the amino acid sequence of the reference Cas12i nuclease. In some embodiments, multiple flexible regions, including secondary structures, crystal structures, and NMR structures, are determined based on the structural information of the reference Cas12i nuclease.

[0123] The method for engineering the flexible region of the Cas12i nuclease here comprises (a) obtaining a plurality of engineered Cas12i nucleases, each engineered Cas12i nuclease containing one or more mutations, the mutations increasing the flexibility of the flexible region in one or more flexible regions of the reference Cas12i nuclease, and (b) selecting one or more engineered Cas12i nucleases from the plurality of engineered Cas12i nucleases, of which the one or more engineered Cas12i nucleases have increased activity (such as target binding, double-strand break activity, nickasing activity, and / or gene editing activity) compared to the reference Cas12i nuclease. In some embodiments, the method further comprises determining one or more flexible regions in the reference Cas12i nuclease. In some embodiments, the method further comprises measuring the activity of the engineered Cas12i nuclease in eukaryotic cells such as mammalian cells (e.g., human cells).

[0124] In some embodiments, multiple flexible regions are determined using a program selected from the group consisting of PredyFlexy, FoldUnfold, PROFbval, Flexserv, FlexPred, DynaMine, and Disomine. In some embodiments, one or more flexible regions are located in a random coil. In some embodiments, one or more flexible regions are located within a domain that interacts with DNA and / or RNA of a reference Cas12i nuclease (or an engineered Cas12i nuclease containing one or more mutations from any of the sections (1) to (5) above). In some embodiments, the flexible region has a length of at least about five (e.g., five) amino acids.

[0125] In some embodiments, the one or more mutations include inserting one or more (e.g., two) glycine (G) residues into the flexible region. In some embodiments, the one or more G residues are inserted into the N-terminus of a flexible amino acid residue in the flexible region, the flexible amino acid residue being selected from the group consisting of G, serine (S), asparagine (N), asparagine (D), histidine (H), methionine (M), threonine (T), glutamic acid (E), glutamine (Q), lysine (K), arginine (R), alanine (A), and proline (P). In some embodiments, the flexible amino acid residue is selected in the order of priority G>S>N>D>H>M>T>E>Q>K>R>A>P. In some embodiments, the one or more mutations include substituting one or more G residues with one or more non-G residues.

[0126] In some embodiments, the one or more mutations include substituting a hydrophobic amino acid residue in the flexible region with a G residue, the hydrophobic amino acid residue being selected from the group consisting of leucine (L), isoleucine (I), valine (V), cysteine ​​(C), tyrosine (Y), phenylalanine (F), and tryptophan (W).

[0127] In some embodiments, the activity is site-directed nuclease activity. In some embodiments, the activity is gene editing activity in eukaryotic cells (e.g., human cells). In some embodiments, the gene editing efficiency is measured using a T7 endonuclease 1 (T7E1) assay, target DNA sequencing, degradation-tracking insertion / deletion (TIDE) assay, or insertion / deletion detection by amplification (IDAA) assay.

[0128] In some embodiments, the engineered Cas12i nuclease (e.g., an engineered Cas12i nuclease comprising one or more mutations from any of sections (1) to (5)) comprises one or more flexible region mutations that increase the flexibility of the flexible region in the reference Cas12i nuclease (e.g., a Cas12i2 nuclease, or an engineered Cas12i nuclease comprising one or more mutations from any of sections (1) to (5)), wherein the flexible region is selected from the group to the regions of amino acid residues 228-232, 439-443, 478-482, 500-504, 775-779, and 925-929, and the amino acid residue numbers are defined by the corresponding amino acid positions shown in SEQ ID NO:1. In some embodiments, the flexible region is selected from amino acid residues 439-443 or 925-929, where the amino acid residue numbers are defined by the corresponding amino acid positions indicated in SEQ ID NO:1. In some embodiments, the reference Cas12i enzyme is Cas12i2 (SEQ ID NO:1). In some embodiments, the one or more flexible region mutations involve inserting one or more (e.g., two) G residues into the flexible region. In some embodiments, the one or more G residues are inserted at the N-terminus of a flexible amino acid residue in the flexible region, where the flexible amino acid residue is selected from the group G, S, N, D, H, M, T, E, Q, K, R, A, P. In some embodiments, the flexible amino acid residue is selected according to the priority order G>S>N>D>H>M>T>E>Q>K>R>A>P. In some embodiments, the one or more flexible region mutations include substituting a hydrophobic amino acid residue in the flexible region with a G residue, wherein the hydrophobic amino acid residue is selected from A, V, I, L, M, F, Y, P, C, or W, preferably L, I, V, C, Y, F, and W.

[0129] In some embodiments, the flexible region mutations are located at 439 and / or 926. In some embodiments, they are one or more amino acids at L439, I926. Of these, the amino acid positions are defined by the corresponding amino acid positions shown in SEQ ID NO:1.

[0130] In some embodiments, the engineered Cas12i nuclease (e.g., an engineered Cas12i nuclease containing one or more mutations from any of sections (1) to (5)) contains the 926G and / or 439(L+G) mutation. In some embodiments, the engineered Cas12i nuclease contains one or more flexible region mutations of I926G, L439(L+G), and L439(L+GG). In some embodiments, the engineered Cas12i nuclease contains the I926G mutation. In some embodiments, the engineered Cas12i nuclease contains the L439(L+G) mutation. In some embodiments, the engineered Cas12i nuclease contains the L439(L+GG) mutation. Of these, the amino acid residue numbers are based on SEQ ID NO:1.

[0131] In the context of this specification, L439(L+G) means that in the cited amino acid sequence (e.g., SEQ ID NO:1), one glycine (G) is inserted after the 439th amino acid, while the original L sequence at position 439 remains unchanged. This may also be expressed as 439G in the context and drawings of this application. L439(L+GG) means that in the cited amino acid sequence, two glycine (GG) molecules are inserted after the 439th amino acid, while the original L sequence at position 439 remains unchanged. This may also be expressed as 439GG in the context and drawings of this application.

[0132] In some embodiments, the engineered Cas12i nuclease (e.g., an engineered Cas12i nuclease comprising one or more mutations from any of sections (1) to (5)) comprises any mutation or combination of mutations at the position of any of the amino acid residues 926, 439, 925+926, 362+925+926, 439+926, 323+362+926 (e.g., I926, L439, N925+I926, D362+N925+I926, L439+I926, E323+D362+I926); of which the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, mutations located at amino acid positions 323, 362, 925, or 926 are mutations that substitute the amino acid residue at the said position with R, H, or K (e.g., R), where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, mutations located at amino acid positions 439 or 926 are mutations that substitute the amino acid residue at the said position with G, or insert G or GG after the said amino acid residue, where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1.

[0133] In some embodiments, the engineered Cas12i nuclease (e.g., an engineered Cas12i nuclease comprising one or more mutations from any of sections (1) to (5)) comprises mutations in amino acid residues or combinations of amino acid residues as described in 926G, 439(L+GG), 925R+926G, 362R+925R+926G, 439(L+GG)+926R, or 323R+362R+926G, where the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1.

[0134] In some embodiments, the engineered Cas12i nuclease includes any of the following mutations or combinations of mutations: I926G, L439(L+GG), L439(L+GG)+I926R, N925R+I926G, D362R+N925R+I926G, E323R+D362R+I926G; of which the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0135] In some embodiments, to achieve the objective of improving gene editing efficiency, the engineered Cas12i nuclease (e.g., an engineered Cas12i nuclease containing one or more mutations in any of sections (1) to (5), and / or the engineered Cas12i nuclease containing the flexible region mutation) may also be an engineered Cas12i nuclease having at least about 85% sequence identity (e.g., at least about 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%).

[0136] 7) Variation of combinations The engineered Cas12i nucleases obtained by the mutations described in sections 1) to 6) of this specification and by combinations of one or more amino acid substitutions / insertions in Tables 1 to 5, 9, 12, 14, and 16 to 18 are all within the scope of the claims of this application.

[0137] In some embodiments, the engineered Cas12i nucleases are 164, 176, 238, 323, 357, 362, 394, 439, 447, 563, 715, 719, 807, 844, 848, 857, 861, 925, 926, 958, 176+238+447+563, 323+362, 176+238+447+563+164, 176+238+447+563+926, 176+238+447+563+323+362, 164+926, 164+323+362, 176+ This includes any mutation or combination of mutations at the positions of any of the following amino acid residues: 238+447+563+164+926, 176+238+447+563+164+323+362, 176+238+447+563+164+926+323+362, 176+238+447+563+164+323+362+926+439, 719+844, 857+844, 362+926, 925+926, 362+925+926, 439+926, and 323+362+926; of which, the amino acid position number is defined by the corresponding amino acid position shown in SEQ ID NO:1. In some embodiments, mutations located at amino acid positions 176, 238, 323, 362, 447, 563, 926, or 958 are mutations that replace the amino acid residue at that position with R, H, or K (e.g., R). In some embodiments, mutations located at amino acid position 164 are mutations that replace the amino acid residue at that position with Y or F (e.g., Y). In some embodiments, mutations located at amino acid positions 439 or 926 are mutations that replace the amino acid residue at that position with G, or insert G or GG after the amino acid residue. In some embodiments, mutations located at amino acid positions 357, 394, 715, 719, 807, 844, 848, 857, or 861 are mutations that replace the amino acid residue at that position with a hydrophobic amino acid (e.g., A). Among these, the amino acid position numbers are defined by the corresponding amino acid positions shown in SEQ ID NO:1.

[0138] In some embodiments, the engineered Cas12i nucleases are N164, E176; K238; E323, D362, T447; E563, I926, I926, D958, L439, R857, N861, K807, N848, R715, R719, K394, H357, K844, E176+K238+T447+E563, E323+D362, E176+K238+T447+E563+N164, E176+K238+T447+E563+I926, E176+K238+T447+E563+E323+D362, N164+ I926, N164+E323+D362, E176+K238+T447+E563+N164+I926, E176+K238+T44 7+E563+N164+E323+D362, E176+K238+T447+E563+N164+I926+E323+D362, E1 76+K238+T447+E563+N164+E323+D362+I926, E176+K238+T447+E563+N164+E 323+D362+I926+L439, E176+K238+T447+E563+N164+E323+D362+I926+L439; E176+K238+T447+E563+N164+D958; E176+K238+T447+E563+I926+D958; E176+K238+T447+E563+E323+D362+D958; N164+I926+D958; N164+E323+D362+D958; E176+K238+T447+E563+N164+D958; E176+K238+T447+E563+N164+I926+D958; E176+K238+T447+E563+N164+E323+D362+D958; E176+K238+T447+E563+N164+I926+E323+D362+D958; E176+K238+T447+E563+N164+E323+D362+I926+D958; E176+K238+T447+E563+N164+E323+D362+I926+L439+D958; E176+K238+T447+E563+N164+E323+D362+I926+L439+D958; R719+K844; R857+K844; E176+K238+T447+E563+N164+E323+D362+R857; E176+K238+T447+E563+N164+E323+D362+N861; E176+K238+T447+E563+N164+E323+D362+K807; E176+K238+T447+E563+N164+E323+D362+N848; E176+K238+T447+E563+N164+E323+D362+R715; E176+K238+T447+E563+N164+E323+D362+R719; E176+K238+T447+E563+N164+E323+D362+K394; E176+K238+T447+E563+N164+E323+D362+H357; E176+K238+T447+E563+N164+E323+D362+K844; E176+K238+T447+E563+N164+E323+D362+R719+K844; E176+K238+T447+E563+N164+E323+D362+R857+K844; or Includes mutations in any of the amino acid residues or combinations of amino acid residues E176+K238+T447+E563+N164+E323+D362+K394; Of these, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0139] In some embodiments, the engineered Cas12i nuclease is N164Y;E176R;K238R;E323R;D362R;T447R;E563R;I926R;D958R;L439(L+G);L439(L+GG);R857A;N861A;K807A;N848A;R715A;R719A;K394A;H357A;K844A;R719A+K844A;N16 4Y+I926R;E323R+D362R;R857A+K844A;R857A+K844A;N164Y+I926R+D958R;N164Y+E323R+D362R;N164Y+E323R +D362R+D958R;E176R+K238R+T447R+E563R;E176R+K238R+T447R+E563R+N164Y;E176R+K238R+T447R+E563R+I 926R;E176R+K238R+T447R+E563R+E323R+D362R;E176R+K238R+T447R+E563R+N164Y+I926R;E176R+K238R+T44 7R+E563R+N164Y+E323R+D362R;E176R+K238R+T447R+E563R+N164Y+I926R+E323R+D362R;E176R+K238R+T447R +E563R+N164Y+E323R+D362R+I926G;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+L439(L+GG);E1 76R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+L439(L+G);E176R+K238R+T447R+E563R+N164Y+D958R; E176R+K238R+T447R+E563R+I926R+D958R; E176R+K238R+T447R+E563R+E323R+D362R+D958R;E176R+K238R+T447R+ E563R+N164Y+D958R;E176R+K238R+T447R+E563R+N164Y+I926R+D958R; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+D958R; E176R+K238R+T447R+E563R+N164Y+I926R+E323R+D362R+D958R; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+D958R; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+L439(L+GG)+D958R;E176R+K238R+T447R+E563R +N164Y+E323R+D362R+I926G+L439(L+G)+D958R;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R857A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+N861A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K807A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+N848A;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R715A;E17 6R+K238R+T447R+E563R+N164Y+E323R+D362R+R719A;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K394A;E176R+K 238R+T447R+E563R+N164Y+E323R+D362R+H357A;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K844A;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R719A+K844A;E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R857A+K844A;or Contains any of the following mutations or combinations of mutations: E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K394A; Of these, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1.

[0140] In some embodiments, the engineered Cas12i nuclease includes the mutant combination E176R+K238R+T447R+E563R+N164Y+E323R+D362R, of which the amino acid position numbers are defined as the corresponding amino acid positions shown in SEQ ID NO:1. This variant is hereafter named CasXX, and its sequence number is SEQ ID NO:8.

[0141] In some embodiments, the engineered Cas12i nuclease is E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R857A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R719A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K394A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K844A; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R719A+K844A; The Cas12i mutants include one of the following mutation combinations: E176R+K238R+T447R+E563R+N164Y+E323R+D362R+R857A+K844A; among these, the amino acid position number is defined as the corresponding amino acid position shown in SEQ ID NO:1. These Cas12i mutants are optimized mutants that have undergone further optimization in addition to CasXX (SEQ ID NO:8).

[0142] In some embodiments, the engineered Cas12i nuclease includes the following combination of mutations: E176R+K238R+T447R+E563R+N164Y+E323R+D362R+K394A, where the amino acid position numbers are defined as the corresponding amino acid positions shown in SEQ ID NO:1. This variant is hereafter named "HF-20", and its sequence is shown in SEQ ID NO:20.

[0143] In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, and D362R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, and I926R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, E323R, and D362R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, and I926G mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, and L439(L+GG) mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, and L439(L+G) mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, and D958R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, I926R, and D958R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, and D958R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, I926R, E323R, D362R, and D958R mutations.In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, and D958R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, L439(L+GG), and D958R mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, and K844A mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, R819a, and K844A mutations. In some embodiments, the engineered Cas12i nuclease includes the E176R, K238R, T447R, E563R, N164Y, E323R, D362R, R857A, and K844A mutations. Of these, the amino acid position numbers are defined by the corresponding amino acid positions shown in SEQ ID NO:1.

[0144] In some embodiments, to achieve the objective of improving gene editing efficiency, an engineered Cas12i nuclease having at least about 80% sequence identity (e.g., at least about 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) may also be used with the engineered Cas12i nuclease.

[0145] In some embodiments, an engineered Cas12i nuclease is provided, comprising an amino acid sequence having at least about 80% (e.g., at least about 81%, 82%, 83%, 84%, 85%, 86%, 87%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence identity with any of the amino acid sequences shown in SEQ ID NOs:2~24.

[0146] Reference Cas12i nuclease In some embodiments, the reference Cas12i nuclease is Cas12i1, Cas12i2, or an ortholog thereof. In some embodiments, the reference Cas12i nuclease is natural Cas12i1, or a variant thereof (e.g., a naturally occurring variant). In some embodiments, the reference Cas12i nuclease is natural Cas12i2 (as shown in SEQ ID NO: 1), or a variant thereof (e.g., a naturally occurring variant). In some embodiments, the reference Cas12i nuclease is a process-modified Cas12i nuclease (an engineered Cas12i nuclease containing one or more mutations from any of sections (1) to (5) described in the present invention). In some embodiments, the reference Cas12i nuclease is CasXX (SEQ ID NO: 8).

[0147] CRISPR-Cas12i type VI has been identified as an RNA-guided DNA endonuclease system. Unlike CRISPR-Cas systems such as Cas12b and Cas9, Cas12i-based CRISPR systems do not require a tracRNA sequence. In some embodiments, the RNA guide sequence includes crRNA. Generally, the crRNA described herein includes a direct repeat sequence and a spacer sequence. In some embodiments, the crRNA includes a direct repeat sequence ligated to a guide sequence or spacer sequence, and is essentially composed of or composed of direct repeat sequences. In some embodiments, the crRNA includes a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-spacer-DR), which are typical features of precursor crRNA (pre-crRNA) structures in other CRISPR systems. In some embodiments, the crRNA includes a truncated direct repeat sequence and a spacer sequence, which are typical features of processed or mature crRNA. In some embodiments, the CRISPR-Cas12i effector protein forms a complex with an RNA guide sequence, and the spacer sequence leads the complex to sequence-specific binding to a target nucleic acid that is at least 70% complementary to the spacer sequence (e.g., at least 70% complementary).

[0148] In some embodiments, the engineered Cas12i of this application is a nuclease that binds to a specific site on a target sequence, is cleaved under the guidance of a guide RNA, and possesses DNA and RNA endonuclease activity. In some embodiments, the Cas12i can perform autonomous crRNA biogeneration by processing a precursor crRNA array. Autonomous precursor crRNA processing can target two distinct genomic sites from a single crRNA transcript, thereby facilitating Cas12i delivery and enabling the application of double nicks. The Cas12i protein then processes the CRISPR array into two homologous crRNAs, thereby forming a paired nick complex. Multiplexing of the type VI (Cas12i) effector protein is performed using the effector protein's precursor crRNA processing ability, which can be programmed on a single RNA guide sequence for multiple targets with different sequences. This allows for the simultaneous manipulation of multiple genes or DNA targets for therapeutic applications. In some embodiments, the guide RNA includes a precursor crRNA expressed by a CRISPR sequence consisting of a target sequence that intersects with the raw DR sequence, and is repeated so that one, two, or more sites can be targeted simultaneously by endogenous precursor crRNA processing of the effector protein.

[0149] Cas12i nucleases from multiple biological sources can be used as reference Cas12i nucleases to provide the engineered Cas12i nucleases and effector proteins of this application. Exemplary Cas12i nucleases are described in WO2019 / 201331a1 and US2020 / 0063126A1, which are incorporated herein by general reference. In some embodiments, the reference Cas12i nuclease has enzymatic activity. In some embodiments, the reference Cas12i cleaves both strands of the nuclease, i.e., the target double-helix nucleic acid (e.g., double-helix DNA). In some embodiments, the reference Cas12i is a nickase that cleaves one strand of the target double-helix nucleic acid (e.g., double-helix DNA). In some embodiments, the reference Cas12i nuclease is enzymatically inactivated. In some embodiments, the reference Cas12i nuclease is Cas12i1, Cas12i2, or Cas12i-Phi. In some embodiments, the reference Cas12i nuclease contains the sequence of SEQID NO:1. Orthologs having sequence identity (e.g., at least about 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more) with Cas12i (e.g., Cas12i2) or its functional derivatives can be used as a basis for designing the engineered Cas12i nucleases or effector proteins of this application.

[0150] Functional variant of engineered Cas12i nuclease In some embodiments, the engineered Cas12i nuclease is a functional variant (or functional derivative) based on a naturally occurring Cas12i nuclease. In some embodiments, the amino acid sequence of the functional variant (or functional derivative) has at least one amino acid residue difference (e.g., deletion, insertion, substitution, and / or fusion) compared to the amino acid sequence of a corresponding engineered Cas12i nuclease, such as an engineered Cas12i nuclease containing one or more mutations from (1) to (7) above. In some embodiments, the functional variant has one or more mutations, such as amino acid substitutions, insertions, and deletions. For example, the functional variant may contain one substitution of any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids compared to a wild-type naturally occurring Cas12i nuclease or to the aforementioned engineered Cas12i nuclease. In some embodiments, one or more substitutions in the variant are conservative substitutions. In some embodiments, the functional variant has all domains of the naturally occurring Cas12i nuclease. In some embodiments, the functional variant does not have one or more domains of the naturally occurring Cas12i nuclease. In some embodiments, the functional variant has all domains of the engineered Cas12i nuclease. In some embodiments, the functional variant does not have one or more domains of the engineered Cas12i nuclease. In some embodiments, the biological activity of the functional variant of the Cas12i nuclease is altered by changes in its amino acids, such as conversion from the natural nuclease to an enzyme-inactive mutant.

[0151] For any of the Cas12i variant proteins described herein (e.g., nickase Cas12i protein, inactivated or catalytically inactivated Cas12i (dCas12i)), the Cas12i variant may include a Cas12i protein sequence having the same parameters as described above (e.g., domains present, identity percentage, etc.).

[0152] In some embodiments, the functional variant of the engineered Cas12i nuclease has enzymatic activity such as DNA double-strand or single-strand break activity, or at least about 60% (e.g., at least about 65%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) of the reference Cas12i nuclease (or its parent engineered Cas12i nuclease). In some embodiments, the enzymatic activity of the functional variant of the engineered Cas12i nuclease is at least 1.1 times (e.g., at least 1.2 times, 1.5 times, 2 times, 3 times, 4 times, 5 times, 10 times or more) of its parent engineered Cas12i nuclease.

[0153] In some embodiments, the functional variant of the engineered Cas12i nuclease has different catalytic activity from the variant form of the non-functional variant of the engineered Cas12i nuclease. In some embodiments, the functional variant mutation (e.g., amino acid substitution, insertion, and / or deletion) is located in the catalytic domain (e.g., the RuvC domain) of the Cas12i nuclease. In some embodiments, the functional variant of the engineered Cas12i nuclease comprises a mutation in one or more catalytic domains. A Cas12i nuclease that cleaves one strand of a double-stranded target nucleic acid but not the other is referred to herein as a “nickase” (e.g., “Cas12i nickase”). In this specification, a Cas12i nuclease that essentially lacks nuclease activity is referred to as an inactivated Cas12i protein ("dCas12i") (a heterologous polypeptide fused with it (fusion partner) can provide nuclease activity; see the chapter "Engineered Cas12i Effector Proteins" below for details). In some embodiments, a Cas12i nuclease functional variant is considered to essentially lack all DNA cleavage activity if the DNA cleavage activity of the functional variant mutant enzyme is approximately 25%, 10%, 5%, 1%, 0.1%, 0.01%, or lower compared to the mutant form of its non-functional variant.

[0154] Cas12i (dCas12i) whose enzyme activity is reduced or lost due to mutations in one or more amino acid residues in the Cas12i nuclease active site is also referred to in this invention as “nuclease activity-deficient Cas12i” or “enzyme-inactive mutant.” In some embodiments, the engineered Cas12i nucleases provided herein can be modified to have reduced or deleted nuclease activity, for example, nuclease activity-deficient Cas12i reduces nuclease activity (e.g., DNA double-strand or single-strand break activity) by at least about 50% (e.g., at least about 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) compared to wild-type Cas12i nuclease (or its parent engineered Cas12i nuclease). The Cas12i nuclease activity can be reduced by several methods, including introducing mutations into one or more domains of the Cas12i nuclease, such as the domain that interacts with PAM, the domain involved in DNA double-strand release, the RuvC domain, and the area that interacts with nucleic acids (DNA / RNA). In some embodiments, the catalytic residue of the Cas12i nuclease activity (e.g., a catalytic residue identified by any general identification method) can be replaced with a different amino acid residue (e.g., glycine or alanine) to reduce the nuclease activity. Examples of mutations such as Cas12i1 (shown in SEQ ID NO: 13) include D647A, E894A, and / or D948A. Examples of mutations such as Cas12i2 (shown in SEQ ID NO: 1) include D599a, E833A, S883A, H884A, D886A, R900a, and / or D1019a.In some embodiments, the engineered Cas12i nuclease or its functional derivative (e.g., an engineered Cas12i nuclease comprising one or more mutations from (1) to (7) above) comprises one or more nuclease activity deletion mutations of D599a, E833A, S883A, H884A, R900a, and D1019a; of which the amino acid positions are defined as the corresponding amino acid positions shown in SEQ ID NO:1.

[0155] Characteristics of engineered Cas12i nucleases or their functional variants In some embodiments, the engineered Cas12i nuclease (or its functional variant) has increased activity compared to the reference Cas12i nuclease. In some embodiments, the engineered Cas12i nuclease (or its functional variant) has decreased activity compared to the reference Cas12i nuclease. In some embodiments, the activity is targeted DNA binding activity. In some embodiments, the activity is site-specific nuclease activity. In some embodiments, the activity is double-strand DNA cleavage activity. In some embodiments, the activity is double-strand DNA cleavage activity. In some embodiments, the activity is single-strand DNA cleavage activity, including, for example, site-specific DNA cleavage activity or non-specific DNA cleavage activity. In some embodiments, the activity is single-strand RNA cleavage activity, including, for example, site-specific RNA cleavage activity or non-specific RNA cleavage activity. In some embodiments, the activity is measured in vitro. In some embodiments, the activity is measured in cells such as bacterial cells, plant cells, or eukaryotic cells. In some embodiments, the activity is measured in mammalian cells such as rodent cells or human cells. In some embodiments, the activity is measured in human cells such as 293T cells. In some embodiments, the activity is measured in mouse cells, e.g., Hepa1-6 cells. In some embodiments, the engineered Cas12i nuclease (or its functional variant) has any activity (one or more of the above-mentioned activities, e.g., site-specific nuclease activity) that is increased by at least about 10%, 20%, 30%, 40%, 50%, 60%, 80%, 90%, 95%, 1x, 1.1x, 1.5x, 2x, 3x, 4x, 5x, 10x or more compared to a reference Cas12i nuclease.In some embodiments, the engineered Cas12i nuclease (or its functional variant) has any activity (one or more of the above-mentioned activities, e.g., site-specific nuclease activity) that is reduced by at least about 10%, 20%, 30%, 40%, 50%, 60%, 80%, 90%, 95%, 1x, 1.1x, 1.5x, 2x, 3x, 4x, 5x, 10x or more compared to a reference Cas12i nuclease. The site-specific nuclease activity of the engineered Cas12i nuclease (or its functional variant) can be measured using methods known in the art, such as gel shift assays, including in vitro cleavage assays based on agarose gel electrophoresis as described in the embodiments provided herein.

[0156] In some embodiments, the activity is gene editing activity in cells. In some embodiments, the cells are bacterial cells, plant cells, or eukaryotic cells. In some embodiments, the cells are mammalian cells such as rodent cells or human cells. In some embodiments, the cells are 293T cells. In some embodiments, the activity is measured in mouse cells, such as Hepa1-6 cells. In some embodiments, the activity is insertion / deletion formation activity at a target genomic site in cells, such as site-specific cleavage of the target nucleic acid by the engineered Cas12i nuclease (or its functional variant), and DNA repair by a non-homologous end joining (NHEJ) mechanism. In some embodiments, the activity is activity to insert an exogenous nucleic acid sequence into a target genomic site in cells, such as site-specific cleavage of the target nucleic acid by the engineered Cas12i nuclease (or its functional variant), and DNA repair by a homologous recombination (HR, by further introduction of a repair template) mechanism. In some embodiments, the engineered Cas12i nuclease (or its functional variant) increases gene editing (e.g., cleavage, insertion / deletion formation, or repair) activity at a target genomic site in a cell (e.g., human cells such as 293T cells, or mouse Hepa1-6 cells) by at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 1x, 1.1x, 1.5x, 2x, 3x, 4x, 5x, 10x, or more, compared to a reference Cas12i nuclease. In some embodiments, the engineered Cas12i nuclease (or its functional variant) increases gene editing (e.g., cleavage, insertion / deletion formation, or repair) activity at multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) target genomic sites in human cells such as 293T cells or mouse Hepa1-6 cells by at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 1x, 1.1x, 1.5x, 2x, 3x, 4x, 5x, 10x, or more compared to a reference Cas12i nuclease.In some embodiments, the engineered Cas12i nuclease (or its functional variant) can edit a greater number of genomic sites (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70 or more) compared to the reference Cas12i nuclease, and can, for example, identify more PAM sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70 or more). In some embodiments, the shared PAM sequence of the engineered Cas12i nuclease (or its functional variant) is the same as that of the reference Cas12i nuclease.

[0157] The gene editing efficiency of engineered Cas12i nucleases (or their functional variants) in vitro or in cells can be determined using methods known in the art, including, for example, T7 endonuclease 1 (T7E1) assays, sequencing of target DNA (e.g., Sanger sequences and two-generation sequences), indel tracking by degradation (TIDE) assays, or indel detection by amplicon analysis (IDAA) assays. See, for example, Sentmanat MF et al., “A survey of validation strategies for CRISPR-Cas9 editing,” Scientific Reports, 2018, 8, article number 888, which is incorporated herein by general reference. In some embodiments, for example, as described in embodiments herein, targeted two-generation sequencing (NGS) is used to measure the gene editing efficiency of the engineered Cas12i nuclease in cells. Examples of genomic sites useful for determining the gene editing efficiency of the engineered Cas12i nuclease (or its functional variant) include, but are not limited to, CCR5, AAVS, CD34, RNF2, and EMX1. In some embodiments, the gene editing efficiency of the engineered Cas12i nuclease (or its functional variant) is the average gene editing efficiency of the engineered Cas12i nuclease at at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65 or more sites (e.g., human cell genome sites). In some embodiments, the gene editing efficiency (e.g., insertion / deletion rate) of the engineered Cas12i nuclease (or its functional variant) reaches at least 10%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or higher.

[0158] In some embodiments, the engineered Cas12i nuclease (or its functional variant) has increased target specificity, a reduced off-target rate (e.g., a reduction in identified off-target sites and / or a decrease in editing efficiency for one or more off-target sites) and / or increased target sequence editing efficiency compared to the reference Cas12i nuclease. In some embodiments, the engineered Cas12i nuclease (or its functional variant) reduces the off-target rate by at least about 5% (e.g., at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%) compared to the reference Cas12i nuclease. In some embodiments, the engineered Cas12i nuclease (or its functional variant) reduces at least one off-target site (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more) compared to the reference Cas12i nuclease. In some embodiments, the engineered Cas12i nuclease (or its functional variant) has the same or similar target sequence editing efficiency (e.g., within 1.1 times) or increased target sequence editing efficiency (e.g., at least 1.2 times, 1.5 times, 2 times, 3 times, 5 times, 10 times or more) compared to the reference Cas12i nuclease. In some embodiments, the engineered Cas12i nuclease (or its functional variant) reduces the off-target rate by at least about 5% compared to the reference Cas12i nuclease, while maintaining the same, similar (e.g., within 1.1 times), or increased editing efficiency for the target sequence.

[0159] Guide RNA or crRNA In some embodiments, the guide RNA or crRNA includes, or is composed of, direct repeat sequences and spacer sequences from 5' to 3'. In some embodiments, the guide RNA includes, or is composed of, direct repeat sequences, spacer sequences and nucleotide sequences constructed by ligation of direct repeat sequences from 5' to 3'. In some embodiments, the RNA guide includes crRNA. In some embodiments, the guide RNA does not include traccrRNA.

[0160] Generally, the crRNAs described herein include direct repeat sequences and spacer sequences. In some embodiments, the crRNA includes, and is essentially composed of, a direct repeat sequence ligated to a guide sequence or a spacer sequence. In some embodiments, the crRNA includes a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-spacer-DR), which is a typical precursor crRNA (pre-crRNA) structure for other CRISPR systems. In some embodiments, the crRNA includes a truncated direct repeat sequence and a spacer sequence, which is a typical processed or mature crRNA. In some embodiments, the CRISPR-Cas effector protein forms a complex with the RNA guide, and the spacer sequence guides this complex to sequence-specifically bind to a target nucleic acid complementary to the spacer sequence.

[0161] In some embodiments, the RNA guide includes direct repeats. In some embodiments, the RNA guide can form a secondary structure of a stem-loop structure, for example, as described herein.

[0162] In some embodiments, the CRISPR system described herein comprises multiple RNA guides (e.g., 2, 3, 4, 5, 10, 15 or more) or multiple nucleic acids encoding multiple RNA guides. In some embodiments, the CRISPR system described herein comprises a single RNA strand or a nucleic acid encoding a single RNA strand, of which the RNA guides are arranged in series. A single RNA strand may include multiple copies of the same RNA guide, multiple copies of different RNA guides, or a combination thereof. In some embodiments, each RNA guide is specific to a different target nucleic acid.

[0163] In some embodiments, the CRISPR system described herein includes an RNA guide or a nucleic acid encoding an RNA guide. In some embodiments, the RNA guide includes, or comprises, a direct repeat sequence and a spacer sequence that can hybridize with a target nucleic acid (e.g., hybridizes under appropriate conditions).

[0164] In some embodiments, the RNA guide sequence is modified to form a CRISPR effector complex, enabling successful binding to the target sequence, while simultaneously preventing successful nuclease activity (i.e., no nuclease activity / no insertion / deletion). These modified guide sequences are called “inactivation guides” or “inactivation guide sequences.” These inactivation guides or inactivation guide sequences may be catalytically inactivated or sterically inactivated with respect to nuclease activity. Inactivation guide sequences are typically shorter than the corresponding guide sequence that results in active RNA cleavage. In some embodiments, the inactivation guide is at least about 5%, 10%, 20%, 30%, 40%, or 50% shorter than the corresponding RNA guide with nuclease activity. The length of the RNA guide inactivation guide sequence may be 13–15 nucleotides (e.g., 13, 14, or 15 nucleotides), 15–19 nucleotides, or 17–18 nucleotides (e.g., 17 nucleotides). In some embodiments, the inactivating guide RNA can hybridize with a target sequence so that the CRISPR system directs to the target genomic locus in the cell without detectable cleavage activity.

[0165] The RNA guide and crRNA sequences and lengths described herein can be optimized. In some embodiments, the optimized length of the RNA guide can be determined by identifying the processed form of the crRNA or by studying the empirical length of the RNA guide of the crRNA. In some embodiments, the RNA guide sequence includes base modifications.

[0166] In some embodiments, the present invention also provides all possible variants of nucleic acids (e.g., cDNA), which can be prepared by selecting combinations based on possible codon selection. These combinations are based on a standard triplet genetic code applicable to polynucleotides encoding naturally occurring variants, and all of these variants are considered to be specifically disclosed.

[0167] Spacer array In some embodiments, the spacer sequence (or spacer, guide sequence) may be complementary to the target sequence of the target nucleic acid, such as DNA, for example, at least about 70% (e.g., at least about 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%). In some embodiments, the spacer sequence is complementary to the target sequence of the target nucleic acid, such as DNA, by at least 15 nucleotides (e.g., at least 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more).

[0168] Assuming sufficient complementarity functions, it is known in the art that perfect complementarity is not necessary. Adjustment of cleavage efficiency can be achieved by introducing mismatches (e.g., one or more mismatches, including mismatch locations along the spacer / target, or one or two mismatches between the spacer and target sequences). The more centrally located a mismatch, such as a double mismatch, is (i.e., not located at the 3' or 5' end), the greater the impact on cleavage efficiency. Therefore, cleavage efficiency can be adjusted by selecting mismatch locations along the spacer sequence. For example, if less than 100% cleavage of the target (e.g., within a cell population) is desired, one or two mismatches between the spacer and target sequences can be introduced into the spacer sequence.

[0169] The guide sequence may have an appropriate length. The spacer length of the RNA guide may be in the range of about 11 to 50 nucleotides (e.g., about 15 to 50 nucleotides). In some embodiments, the guide sequence is between about 18 and about 35 nucleotides, for example, including 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In some embodiments, the spacer length of the RNA guide is at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, or at least 22 nucleotides. In some embodiments, the spacer length is 15–17 nucleotides, 15–23 nucleotides, 16–22 nucleotides, 17–20 nucleotides, 20–24 nucleotides (e.g., 20, 21, 22, 23 or 24 nucleotides), 23–25 nucleotides (e.g., 23, 24 or 25 nucleotides), 24–27 nucleotides, 27–30 nucleotides, 30–45 nucleotides (e.g., 30, 31, 32, 33, 34, 35, 40 or 45 nucleotides), 30 or 35–40 nucleotides, 41–45 nucleotides, 45–50 nucleotides, or more. In some embodiments, the spacer length of the RNA guide is 31 nucleotides. In some embodiments, the direct repeat length of the RNA guide is at least 21 nucleotides, or 21–37 nucleotides (e.g., 23, 24, 25, 30, 35 or 36 nucleotides). In some embodiments, the spacer sequence contains or consists of approximately 15 to approximately 34 nucleotides (e.g., 16, 17, 18, 19, 20, 21, or 22 nucleotides). In some embodiments, the spacer sequence length is between 17 and 31 nucleotides. In some embodiments, the spacer sequence length is between 15 and 24 nucleotides. In some embodiments, the direct repeat length of the RNA guide is 23 or 20 nucleotides.

[0170] Direct Repeat (DR) Sequence Direct repeat sequences can guide Cas12i proteins (engineered Cas12i nucleases, functional variants, or effector proteins described in any of the present inventions) to a guide gRNA (or crRNA) to form a CRISPR-Cas complex that targets a target sequence. Any DR sequence that can be used in this application to form a CRISPR-Cas complex that targets a target sequence by binding an engineered Cas12i nuclease or effector protein to a guide gRNA (or crRNA), for example, the DR sequences described in U.S. 11168324 (all of which are incorporated herein by general reference), can be used in the present invention.

[0171] A direct repeat may contain two complementary nucleotide segments separated by an intervening nucleotide, thereby allowing the direct repeat to hybridize to form a double-stranded RNA (dsRNA) double helix, resulting in a stem-loop structure, where the two complementary nucleotide segments form the stem and the intervening nucleotide forms the loop or hairpin. For example, the intervening nucleotide forming the “loop” may have a length of about 6 to 8 nucleotides, or about 7 nucleotides. In different embodiments, the stem may contain at least 2, at least 3, at least 4, or 5 base pairs.

[0172] In some embodiments, a direct repeat may contain two nucleotide complementary segments of approximately 4 to approximately 7 (e.g., 4, 5, 6, 7) nucleotides in length, with approximately 5 to approximately 9 (e.g., 5, 6, 7, 8, 9) nucleotides in between. Those skilled in the art can simulate known direct repeat structures.

[0173] Direct repeats may contain or consist of approximately 13 to approximately 23 nucleotides, approximately 22 to approximately 40 nucleotides, approximately 23 to approximately 38 nucleotides, or approximately 23 to approximately 36 nucleotides.

[0174] In some embodiments, the direct repeat sequence includes a stem-loop structure near the 3' end (adjacent to the spacer sequence). In some embodiments, the direct repeat sequence includes a stem-loop near the 3' end with a stem length of 5 nucleotides. In some embodiments, the direct repeat sequence includes a stem-loop near the 3' end with a stem length of 5 nucleotides and a loop length of 7 nucleotides. In some embodiments, the direct repeat sequence includes a stem-loop near the 3' end with a stem length of 5 nucleotides and a loop length of 6, 7, or 8 nucleotides.

[0175] In some embodiments, the direct repeat sequence includes the sequence 5'-CCGUCNNNNNNUGACGG-3' (SEQ ID NO: 68) near the 3' end, where N refers to any nucleic acid base. In some embodiments, the direct repeat sequence includes the sequence 5'-GUGCCNNNNNNUGGCAC-3' (SEQ ID NO: 69) near the 3' end, where N refers to any nucleic acid base.

[0176] In some embodiments, the direct repeat sequence is the sequence 5'-GUGUCN near the 3' end. 5-6 Includes UGACAX1-3' (SEQ ID NO: 70 or 71), of which N 5-6X refers to a sequence of any 5 or 6 nucleic acid bases, where X1 refers to C, T, or U. In some embodiments, the direct repeat sequence includes a sequence near the 3' end 5'-UCX3UX5X6X7UUGACGG-3' (SEQ ID NO: 72), where X3 refers to C, T, or U, X5 refers to A, T, or U, X6 refers to A, C, or G, and X8 refers to A or G. In some embodiments, the direct repeat sequence includes a sequence near the 3' end 5'-CCX3X4X5CX7UUGGCAC-3' (SEQ ID NO: 73), where X3 refers to C, T, or U, X4 refers to A, T, or U, X5 refers to C, T, or U, and X7 refers to A or G. In some embodiments, the nucleotides encoding the direct repeat sequence contain or consist of a nucleotide sequence that is at least about 80% identical to SEQ ID NO: 59 (AGAAATCCGTCTTTCATTGACGG) or SEQ ID NO: 79 (GTTGCAAAACCCAAGAAAT CCGTCTTTCATTGACGG), for example, at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical. In some embodiments, the nucleotides encoding the direct repeat sequence contain at least 21 (e.g., 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or 36) nucleotides of SEQ ID NO: 79.

[0177] A “stem-loop structure” means a nucleic acid having a secondary structure that includes a nucleotide region known or predicted to form a double helix (stem), the double helix (stem) being linked on one side by a region (loop) that is primarily single-stranded nucleotides. The terms “hairpin” and “folded” structures are also used herein to refer to stem-ring structures. Such structures are known in the art, and these terms are used in accordance with their known meanings in the art. As is known in the art, stem-loop structures do not require precise base pairing. Therefore, the stem may contain one or more base mismatches. Alternatively, the base pairing may be precise, i.e., without mismatches. In some embodiments, the stem of a direct repeat consists of five complementary nucleic acid bases hybridized with each other, and the loop length is 6, 7, or 9 nucleotides.

[0178] In some embodiments, the sequence encoding the direct repeat includes or consists of the sequence shown in SEQ ID NO:79. In some embodiments, the sequence encoding the direct repeat includes or consists of a truncated nucleic acid having the first three 5' nucleotides of the SEQ ID NO:79 nucleic acid sequence. In some embodiments, the sequence encoding the direct repeat includes or consists of a truncated nucleic acid having the first four 5' nucleotides of the SEQ ID NO:79 nucleic acid sequence. In some embodiments, the sequence encoding the direct repeat includes or consists of a truncated nucleic acid having the first 5' nucleotides of the SEQ ID NO:79 nucleic acid sequence. In some embodiments, the sequence encoding the direct repeat includes or consists of a truncated nucleic acid having the first six 5' nucleotides of the SEQ ID NO:79 nucleic acid sequence. In some embodiments, the sequence encoding the direct repeat includes or consists of a truncated nucleic acid having the first seven 5' nucleotides of the SEQ ID NO:79 nucleic acid sequence. In some embodiments, the sequence encoding the direct repeat comprises or consists of a truncated nucleic acid having the first eight 5' nucleotides of the nucleic acid sequence SEQ ID NO:79. In some embodiments, the sequence encoding the direct repeat comprises or consists of the sequence indicated by SEQ ID NO:59.

[0179] In some embodiments, the direct repeat is a “functional variant” of the RNA sequence encoded by SEQ ID NO: 59 or 79, e.g., a “functional shortened version,” a “functional extended version,” or a “functional substitution version,” e.g., a portion of SEQ ID NO: 79 (shortened version), which still possesses DR function. The DR “functional variant” is a DR sequence that, after being extended at the 5' and / or 3' ends of a reference DR (e.g., parent DR) (functional extended version) or truncated (functional truncated version), and / or having one or more nucleotides inserted, deleted, and / or substituted into the reference DR sequence (functional substitution version), still possesses at least 20% of the function of the reference DR (e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more), i.e., mediates the function of binding the Cas12i protein to the corresponding crRNA. The DR functional variant generally retains a stem-loop-like secondary structure or a portion thereof that can be used for Cas12i protein binding. In some embodiments, the DR or its functional variant includes a stem-loop-like secondary structure or a portion thereof that can be used for Cas12i protein binding. In some embodiments, the DR or its functional variant includes at least two (e.g., 2, 3, 4, 5 or more) stem-loop-like secondary structures or portions thereof that can be used for Cas12i protein binding.

[0180] Prime Editing Guide: RNA (pegRNA) PEgRNA (relative to standard guide RNA) is modified to include an extension that provides a DNA synthesis template sequence encoding a single-stranded DNA flap, which is homologous to the strand of the target endogenous DNA sequence to be edited but contains one or more desired nucleotide changes, and is then incorporated into the target DNA molecule after being synthesized by a polymerase (e.g., reverse transcriptase). PEgRNA has already been described in WO2020191246, WO2021226558, etc., and is incorporated herein by general reference.

[0181] In different embodiments, the extended guide RNA includes (a) the guide RNA and (b) the RNA extension at the 5' or 3' end of the guide RNA, or at an intramolecular position of the guide RNA. Preferably, intramolecular localization of the extension portion does not impair the function of the protospacer region. The RNA extension may include (i) a reverse transcription template sequence containing the desired nucleotide changes, (ii) a reverse transcription primer binding site, and (iii) an optional linker sequence. In different embodiments, the reverse transcription template sequence may encode a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to the nick site, wherein the single-stranded DNA flap contains the desired nucleotide changes. The single-stranded DNA flap can replace the endogenous single-stranded DNA at the nick site. In different embodiments, the desired nucleotide changes incorporated into the target DNA may be single-nucleotide changes (e.g., transitions or transversions), insertions of one or more nucleotides, or deletions of one or more nucleotides.

[0182] In different embodiments, the desired nucleotide change may be a single nucleotide change (e.g., a transition or transversion change), a deletion, or an insertion. For example, the desired nucleotide change may be (1) a G-T substitution, (2) a G-A substitution, (3) a G-C substitution, (4) a T-G substitution, (5) a T-A substitution, (6) a T-C substitution, (7) a C-G substitution, (8) a C-T substitution, (9) a C-A substitution, (10) an A-T substitution, (11) an A-G substitution, or (12) an A-C substitution.

[0183] PAM The engineered Cas12i nuclease (or functional variant thereof) described in the present invention can identify a protospacer adjacent motif (PAM). In some embodiments, the target nucleic acid comprises a PAM. In some embodiments, the PAM is located at the 5' end of a sequence complementary to the target sequence that induces RNA in the target nucleic acid. In some embodiments, the PAM comprises or consists of the nucleic acid sequence 5'-TTN-3', 5'-TTH-3', 5'-TTY-3', or 5'-TTC-3'. In some embodiments, the PAM of the engineered Cas12i nuclease (or functional variant thereof) described in the present invention comprises the nucleic acid sequence 5'-N NN N-3'( N =A, T, G, C) For example, the PAM includes or consists of 5'-NTTN-3', 5'-NTAN-3', 5'-NTCN-3', 5'-NTGN-3', 5'-NATN-3', 5'-NAAN-3', 5'-NACN-3', 5'-NAGN-3', 5'-NCTN-3', 5'-NCAN-3', 5'-NCCN-3', 5'-NCGN-3', 5'-NGTN-3', 5'-NGAN-3', 5'-NGCN-3', 5'-NGGN-3'. In some embodiments, the PAM includes or consists of the nucleic acid sequence 5'-TTTN-3'. In some embodiments, the PAM includes or consists of the nucleic acid sequence 5'-TTN-3'.

[0184] In some embodiments, the PAM of the engineered and suitable Cas12i nuclease (or functional variant thereof) described in the present invention includes the nucleic acid sequences 5'-TTTA-3', 5'-CTTA-3', 5'-GTTA-3', 5'-ATTA-3', 5'-TTTC-3', 5'-CTTC-3', 5'-GTTC-3', 5'-ATTC-3', 5'-TTTG-3', 5'-CTTG-3', 5'-GTTG-3', 5'-ATTG-3', 5'-TTTT-3', 5'-CTTT-3', 5'-GTTT-3', and 5'-ATTT-3'.

[0185] Engineered Cas12i effector protein This application provides an engineered Cas12i (e.g., Cas12i2) effector protein comprising any engineered Cas12i nuclease or functional variant thereof as described in the present invention (e.g., CasXX shown in SEQID NO: 8), having improved activity such as target binding, double-strand cleavage activity, nickase activity, and / or gene editing activity. In some embodiments, an engineered Cas12i effector protein (e.g., Cas12i nuclease, Cas12i nickase, Cas12i fusion effector protein, or split Cas12i effector protein) comprising any of the engineered Cas12i nuclease or functional derivative thereof as described herein (e.g., dCas12i) is provided. In some embodiments, the engineered Cas12i effector protein comprises, or is composed primarily of, any engineered Cas12i nuclease or functional variant thereof as described in the present invention.

[0186] Further provided herein are engineered Cas12i effector proteins based on any of the engineered Cas12i2 nucleases or their functional variants (e.g., CasXX shown in SEQ ID NO: 8). In some embodiments, the engineered Cas12i effector protein has enzymatic activity such as DNA double-strand break activity. In some embodiments, the engineered Cas12i effector protein has nuclease activity that cleaves two strands of a target double-helix nucleic acid (e.g., double-helix DNA). In some embodiments, the engineered Cas12i effector protein has nickase activity that cleaves one strand of a target double-helix nucleic acid (e.g., double-helix DNA). In some embodiments, the engineered Cas12i effector protein includes an enzymatically inactive variant of the engineered Cas12i nuclease.

[0187] Split Cas12i effector protein This application also provides a split Cas12i effector protein based on one of the engineered Cas12i nucleases or functional variants thereof (e.g., CasXX shown in SEQ ID NO: 8) described herein. Split Cas12i effector proteins may be advantageous for delivery. In some embodiments, the engineered Cas12i effector protein can be split into two parts of the enzyme, and these two parts can be reconstituted together to provide a basically functional Cas12i effector protein. Cas effector proteins can be provided using known methods, for example, split forms of Cas12 and Cas9 proteins have already been described in WO2016 / 112242, WO2016 / 205749 and PCT / CN2020 / 11057, which are incorporated herein by general reference.

[0188] In some embodiments, a split Cas12i effector protein is provided, comprising a first polypeptide and a second polypeptide, wherein the first polypeptide comprises the N-terminal portion of any of the engineered Cas12i nucleases or functional derivatives described herein, and the second polypeptide comprises the C-terminal portion of any of the engineered Cas12i nucleases or functional derivatives thereof, wherein the first polypeptide and the second polypeptide can associate with each other in the presence of a guide RNA containing the guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid containing a target sequence complementary to the guide sequence. In some embodiments, the first polypeptide and the second polypeptide each contain a dimerization domain. In some embodiments, the first dimerization domain and the second dimerization domain associate with each other in the presence of an inducer (e.g., rapamycin). In some embodiments, the first polypeptide and the second polypeptide do not contain a dimerization domain. In some embodiments, the split Cas12i effector protein is self-inducible.

[0189] The engineered Cas12i effector proteins described in this invention can be cleaved in a manner that does not affect the catalytic domain. The Cas12i effector proteins can be used as nucleases (including nickase), or they can be inactivating enzymes that are RNA-induced DNA-binding proteins with essentially low or no catalytic activity (e.g., due to mutations in their catalytic domain).

[0190] In some embodiments, the nucleazerove and α-helix lobe of the engineered Cas12i effector protein are expressed as a separate polypeptide. While the nucleazerove and α-helix lobe do not self-interact, an RNA guide sequence recruits them into a single complex, which replicates the activity of the full-length Cas12i nuclease and catalyzes site-specific DNA cleavage. In some embodiments, a modified RNA guide sequence can be used to eliminate the activity of the split-Cas9 enzyme by preventing dimerization, enabling the development of an inducible dimerization system. This split-Cas9 enzyme is described, for example, in Wright, Addison V., et al. “Rational design of a split-Cas9 enzyme complex,” Proc. Nat'l. Acad. Sci., 112.10 (2015): 2984-2989, which is incorporated herein by general reference.

[0191] The split Cas12i effector protein moieties described herein may be designed to split (i.e., split) a reference-engineered Cas12i effector protein (e.g., a full-length engineered Cas12i nuclease) into two parts at a single splitting position, where the position separates the N-terminus from the C-terminus of the reference Cas12i effector protein. In some embodiments, the N-terminus comprises amino acid residues 1 to X of the reference Cas12i effector protein, and the C-terminus comprises amino acid residues X+1 to the C-terminus of the reference Cas12i effector protein. In this example, the numbering is consecutive but not required. This is because it is also taken into consideration that amino acids (or encoding nucleotides) may be trimmed in the internal region of the polypeptide chain from either a split end and / or mutation (e.g., insertion, deletion, substitution), and the condition is that the reconstituted engineered Cas12i effector protein retains sufficient DNA binding activity (if necessary), DNA nickase activity, or cutting activity, e.g., at least about 40% (e.g., at least about 50%, 60%, 70%, 80%, 90%, 95% or more) compared to the aforementioned reference Cas12i effector protein.

[0192] Splitting points can be designed in silico and cloned into constructs. In this process, mutations can be introduced into the split Cas12i effector protein, and non-functional domains can be removed. In some embodiments, the two parts or fragments of the split Cas12i effector protein (i.e., the N-terminal and C-terminal fragments) can form a complete Cas12i effector protein containing at least about 70% (e.g., at least about 80%, 90%, 95%, 96%, 97%, 98%, 99%, or more) of the complete Cas12i effector protein sequence.

[0193] Each of the split Cas12i effector proteins may contain one or more dimerization domains. In some embodiments, the first polypeptide comprises a first dimerization domain that partially fuses with the first split Cas12i effector protein, and the second polypeptide comprises a second dimerization domain that partially fuses with the second split Cas12i effector protein. The dimerization domains can be fused to the split Cas12i effector protein moiety by a peptide linker (e.g., a flexible peptide linker such as a GS linker) or by chemical bonding. In some embodiments, the dimerization domain fuses with the N-terminus of the split Cas12i effector protein moiety. In some embodiments, the dimerization domain fuses with the C-terminus of the split Cas12i effector protein moiety.

[0194] In some embodiments, the split Cas12i effector protein does not contain a dimerization domain.

[0195] In some embodiments, the dimerizing domain facilitates the association of two fragmented Cas12i effector protein moieties. In some embodiments, the fragmented Cas12i effector protein moieties are induced by an inducer and associated with or dimerized into a functional Cas12i effector protein. In some embodiments, the fragmented Cas12i effector protein contains an inducible dimerizing domain. In some embodiments, the dimerizing domain is not an inducible dimerizing domain; i.e., the dimerizing domain dimerizes in the absence of an inducer.

[0196] The inducer may be an inducer energy source or inducer molecule other than a guide RNA (e.g., crRNA). The inducer reconstitutes two fragmented Cas12i effector protein moieties into a functional Cas12i effector protein by dimerization of the induced dimerization domain. In some embodiments, the inducer links the two fragmented Cas12i effector protein moieties by the action of induceable association of the dimerization domain. In some embodiments, in the absence of an inducer, the two fragmented Cas12i effector protein moieties do not associate with each other and are reconstituted as a functional Cas12i effector protein. In some embodiments, in the absence of an inducer, the two separate Cas12i effector protein moieties can associate with each other in the presence of a guide RNA (e.g., crRNA) and be reconstituted as a functional Cas12i effector protein.

[0197] The inducer of this application may be heat, ultrasound, electromagnetic energy, or a compound. In some embodiments, the inducer is an antibiotic, a small molecule, a hormone, a hormone derivative, a steroid, or a steroid derivative. In some embodiments, the inducer is abscisic acid (ABA), doxycycline (DOX), isopropylbenzoic acid (cumate), rapamycin, 4-hydroxytamoxifen (4OHT), estrogen, or a molting hormone. In some embodiments, the segmented Cas12i effector system is an inducer-controlled system selected from the group consisting of antibiotic-based inducers, electromagnetic energy-based inducers, small molecule-based inducers, nuclear receptor-based inducers, and hormone-based inducers. In some embodiments, the segmented Cas12i effector system is an inducer-controlled system selected from the group consisting of tetracycline (Tet) / DOX induction systems, photoinduction systems, ABA induction systems, isopropylbenzoic acid (cumate) inhibitor / operon systems, 4OHT / estrogen induction systems, molting hormone-based induction systems, and FKBP12 / FRAP (FKBP12-rapamycin complex) induction systems. Such inducers are also discussed herein and in PCT / US2013 / 051418, which are incorporated herein by general reference. The FRB / FKBP / rapamycin system is described in Paulmurugan and Gambhir, Cancer Res, August 15, 2005 65; 7413, and Crabtree et al., Chemistry & Biology 13, 99-107, Jan 2006, which are incorporated herein by general reference.

[0198] In some embodiments, the paired split Cas12i effector proteins are separated and inactive until dimerization of the dimerizing domain (e.g., FRB and FKBP) is induced, resulting in the reconstruction of the functional Cas12i effector protein nuclease. In some embodiments, the first split Cas12i effector protein, containing the first half of the inducible dimer (e.g., FRB), is separated and delivered, and / or located separately from the second split Cas12i effector protein, containing the second half of the inducible dimer (e.g., FKBP).

[0199] Other exemplary FKBP-based induction systems for the segmented Cas12i effector systems that can be used for induction control as described herein include, but are not limited to, FKBP dimerized with calcineurin (CNA) in the presence of FK506, FKBP dimerized with CyP-Fas in the presence of FKCsA, FKBP dimerized with FRB in the presence of rapamycin, GyrB dimerized with GryB in the presence of coumarinmycin, GAI dimerized with GID1 in the presence of gibberellin, or Snap-tag dimerized with HaloTag in the presence of HaXS.

[0200] Alternatives within the FKBP family itself are also being considered. For example, in the presence of FK1012, FKBPs can be homodimerized (i.e., one FKBP can dimerize with another).

[0201] In some embodiments, the dimerizing domain is FKBP, and the inducer is FK1012. In some embodiments, the dimerizing domain is GryB, and the inducer is coumarinmycin. In some embodiments, the dimerizing domain is ABA, and the inducer is erythromycin.

[0202] In some embodiments, the fragmented Cas12i effector protein moiety can be automatically induced (i.e., auto-activated or self-induced) without an inducer to associate / dimerize with a functional Cas12i effector protein. While not bound by any theory or assumption, the auto-induction of the fragmented Cas12i effector protein moiety can be mediated by binding to a guide RNA such as crRNA. In some embodiments, the first and second polypeptides do not contain dimerization domains. In some embodiments, the first and second polypeptides contain dimerization domains.

[0203] In some embodiments, the reconfigured Cas12i effector protein of the fragmented Cas12i effector system described herein (including inducer control and auto-induction systems) has an editing efficiency of at least about 60% (e.g., at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or higher) relative to the reference Cas12i effector protein editing efficiency.

[0204] In some embodiments, the reconstituted Cas12i effector protein of the inducer-controlled split Cas12i effector system described herein has an editing efficiency of less than about 50% (e.g., about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or less) of the reference Cas12i effector protein editing efficiency in the absence of an inducer (i.e., by autoinduction).

[0205] Fusion Cas12i effector protein This application also provides an engineered Cas12i effector protein comprising additional protein domains and / or components such as a junction, nuclear localization / transport sequence, functional domain, and / or reporter protein.

[0206] In some embodiments, the engineered Cas12i effector protein is a protein complex comprising one or more heterogeneous protein domains (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains) and a nucleic acid target domain or functional derivative of any of the engineered Cas12i nucleases described in the present invention. In some embodiments, the engineered Cas12i effector protein is a fusion protein comprising one or more heterogeneous protein domains (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains) fused with the engineered Cas12i nuclease or a functional variant thereof (e.g., CasXX shown in SEQ ID NO: 8).

[0207] In some embodiments, the engineered Cas12i effector proteins of this application comprise one or more functional domains (e.g., by fusion of the protein with one or more peptide linkers such as GS peptide linkers), or one or more functional domains can be associated therewith (e.g., by co-expression of multiple proteins). In some embodiments, the one or more functional domains are enzyme domains. These functional domains can have many activities, such as DNA and / or RNA methylase activity, nucleotide deaminase activity (adenosine deaminase activity, cytidine deaminase activity), demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity (e.g., double-stranded endonuclease activity, nickase activity), nucleic acid binding activity, and switching activity (e.g., photo-induction or chemical induction). In some embodiments, the one or more functional domains are transcriptional activation domains (i.e., trans-activating domains) or repressor domains. In some embodiments, the one or more functional domains are histone modification domains. In some embodiments, the one or more functional domains are a transposase domain, a homologous recombination (HR) mechanism domain, a recombinase domain, and / or an integrase domain. In some embodiments, the functional domains are Kruppel correlation cassette (KRAB), VP64, VP16, Fok1, P65, HSF1, MyoD1, biotin-APEX, APOBEC1, AID, PmCDA1, Tad1, and M-MLV reverse transcriptase. In some embodiments, the functional domains are selected from the group consisting of a translation initiation domain, a transcriptional repression domain, a transactivation domain, an epigenetic modification domain, a nuclear base editing domain (e.g., a CBE or ABE domain), a reverse transcriptase domain, a reporter molecule domain (e.g., a fluorescence domain), and a nuclease domain.

[0208] In some embodiments, the functional domain has activity to modify target DNA or target DNA-related proteins, and such activity includes nuclease activity (e.g., HNH nuclease, RuvC nuclease, Trex1 nuclease, Trex2 nuclease), methylation activity, demethylation activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimerization activity, integrase activity, transposase activity, recombinase activity, polymerase activity, and ligase activity. The activity is selected from one or more of the following: helicase activity, photolyase activity, glycosylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenosine oxidation activity, deadenosine oxidation activity, SUMOylation activity, deSUMOylation activity, riboxylation activity, deriboxylation activity, myristylation activity, demyristylation activity, glycosylation activity (e.g., derived from O-GlcNAc transferase), deglycosylation activity, transcriptional repression activity, and transcriptional activation activity. Target DNA-related proteins refer to proteins that can bind to target DNA, such as histones, transcription factors, and mediators, or proteins that can bind to proteins that can bind to target DNA.

[0209] In some embodiments, the localization of one or more functional domains in the engineered Cas12i effector protein allows for a precise spatial orientation of the functional domain to affect a target having a given functional effect. For example, if the functional domain is a transcription activator (e.g., VP16, VP64, or p65), the transcription activator is positioned in a spatial orientation that can affect target transcription. Similarly, a transcription repressor is positioned to affect target transcription, and a nuclease (e.g., Fok1) is positioned to cleave or partially cleave a target. In some embodiments, the functional domain is located at the N-terminus of the engineered Cas12i effector protein. In some embodiments, the functional domain is located at the C-terminus of the engineered Cas12i effector protein. In some embodiments, the engineered Cas12i effector protein contains a first functional domain at the N-terminus and a second functional domain at the C-terminus. In some embodiments, the engineered Cas12i effector protein comprises a catalytically inactivated mutant (dCas12i) of any of the Cas12i nucleases engineered herein, fused with one or more functional domains.

[0210] In some embodiments, the engineered Cas12i effector protein is a transcription activator. In some embodiments, the engineered Cas12i effector protein comprises an enzymatic inactivation variant of any of the Cas12i nucleases engineered herein, fused with a transactivation domain. In some embodiments, the transactivation domain is selected from the group consisting of VP64, p65, HSF1, VP16, MyoD1, HSF1, RTA, SET7 / 9, and combinations thereof. In some embodiments, the transactivation domain comprises VP64, p65, and HSF1. In some embodiments, the engineered Cas12i effector protein comprises two split Cas12i effect polypeptides, each fused with a transactivation domain.

[0211] In some embodiments, the engineered Cas12i effector protein is a transcriptional repressor. In some embodiments, the engineered Cas12i effector protein comprises an enzymatically inactivated variant of any of the Cas12i nucleases engineered herein, fused with a transcriptional repression domain. In some embodiments, the transcriptional repressor domain is selected from the group consisting of Kruppel-correlated cassette (KRAB), EnR, NuE, NcoR, SID, SID4X, and combinations thereof. In some embodiments, the engineered Cas12i effector protein comprises two split Cas12i effect polypeptides, each fused with a transcriptional repression domain.

[0212] In some embodiments, the engineered Cas12i effector protein is a base editor, such as a cytosine editor or an adenosine editor. In some embodiments, the engineered Cas12i effector protein comprises an enzymatic inactivation variant of any of the engineered Cas12i nucleases described herein, fused with a nuclear base editing domain, such as a cytosine base editing (CBE) domain or an adenosine base editing (ABE) domain. In some embodiments, the nucleic acid base editing domain is a DNA editing domain. In some embodiments, the nucleic acid base editing domain has deaminase activity. In some embodiments, the nucleic acid base editing domain is a cytosine deaminase domain. In some embodiments, the nucleic acid base editing domain is an adenosine deaminase domain. Exemplary base editors based on Cas nucleases are described, for example, in WO2018 / 165629A1 and WO2019 / 226953A1, which are incorporated herein by general reference. Exemplary CBE domains include, but are not limited to, activation-inducible cytosine deaminase or AID (e.g., hAID), apolipoprotein B mRNA editing complex or APOBEC (e.g., rat APOBEC1, hAPOBEC3A / B / C / D / E / F / G), and PmCDA1. Exemplary ABE domains include, but are not limited to, TadA, ABE8, and their variants (see, for example, Gaudelli et al., 2017, Nature 551: 464-471; and Richter et al., 2020, Nature Biotechnology 38: 883-891). In some embodiments, the functional domain is an APOBEC1 domain, such as the rat APOBEC1 domain. In some embodiments, the functional domain is a TadA domain, such as the Escherichia coli (E. coli) TadA domain. In some embodiments, the engineered Cas12i effector protein further includes one or more nuclear localization sequences.

[0213] Adenosine deaminase As used herein, the terms “adenosine deaminase” or “adenosine deaminase protein” refer to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze a hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) to hypoxanthine (or the hypoxanthine portion of a molecule), as shown below. In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0214] Adenosine deaminases that can be used in combination with any of the engineerable Cas12i nuclease enzyme inactivation variants of the present invention include, but are not limited to, enzyme family members called adenosine deaminases that act on RNA (ADAR), enzyme family members called adenosine deaminases that act on tRNA (ADAT), and other family members containing adenosine deaminase domains (ADAD). Adenosine deaminases can target adenine in RNA / DNA and RNA double strands. In fact, Zheng et al. (Nucleic Acids Res. 2017, 45(6):3369-3377) confirmed that ADAR can perform adenosine-to-inosine editing reactions in RNA / DNA and RNA / RNA double strands. In some embodiments, adenosine deaminases can be modified to enhance their ability to edit DNA in RNA double strands and in heterologous RNA / DNA double strands.

[0215] In some embodiments, the adenosine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squid, fish, flies, and worms. In some embodiments, the adenosine deaminase is adenosine deaminase from humans, squid, or fruit flies.

[0216] In some embodiments, adenosine deaminase is human ADAR, comprising hADAR1, hADAR2, and hADAR3. In some embodiments, adenosine deaminase is Caenorhabditis elegans ADAR protein, comprising ADR-1 and ADR-2. In some embodiments, adenosine deaminase is Drosophila ADAR protein, comprising dAdar. In some embodiments, adenosine deaminase is Loligo pealeii ADAR protein, comprising sqADAR2a and sqADAR2b. In some embodiments, adenosine deaminase is human ADAT protein. In some embodiments, adenosine deaminase is Drosophila ADAT protein. In some embodiments, adenosine deaminase is human ADAD protein, comprising TENR(hADAD1) and TENRL(hADAD2). In some embodiments, the adenosine deaminase is TadA8e.

[0217] In some embodiments, adenosine deaminase is a TadA protein, such as E. coli TadA. See Kim et al., Biochemistry 45:6407-6416 (2006) and Wolf et al., EMBO J. 21:3841-3851 (2002). In some embodiments, adenosine deaminase is mouse ADA. See Grunebaum et al., Curr. Opin. Allergy Clin. Immunol. 13:630-638 (2013). In some embodiments, adenosine deaminase is human ADAT2. See Fukui et al., J. Nucleic Acids 2010:260512 (2010). In some embodiments, the deaminase (e.g., adenosine or cytidine deaminase) is one or more of those listed in the following literature: see Cox et al., Science. November 24, 2017, 358(6366):1019-1027; Komore et al., Nature. May 19, 2016, 533(7603):420-4; and Gaudelli et al., Nature. November 23, 2017, 551(7681):464-471.

[0218] In some embodiments, the adenosine deaminase protein comprises one or more deaminase domains. While it is undesirable to be bound by a specific theory, it is expected that the deaminase domains will be used to identify one or more target adenosine (A) residues contained in a double-stranded nucleic acid substrate and convert them to inosine (I) residues.

[0219] Cytidine deaminase In some embodiments, the deaminase is cytidine deaminase. As used herein, the terms “cytosine deaminase” or “cytosine deaminase protein” refer to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze a hydrolytic deamination reaction that converts cytosine (or the cytosine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytosine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0220] Cytosine deaminases that can be used in combination with any of the engineered Cas12i nuclease enzyme inactivation variants of the present invention include, but are not limited to, members of an enzyme family called apolipoprotein B mRNA editing complex (APOBEC) family deaminases, activation-inducing deaminases (AIDs), or cytosine deaminase 1 (CDA1). In some embodiments, these are deaminases in APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3E deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, or APOBEC4 deaminase.

[0221] In some embodiments, the cytosine deaminase can target cytosine in a single strand of DNA. In some embodiments, the cytosine deaminase can perform editing on a single strand located outside of a binding component. In some embodiments, the cytosine deaminase can perform editing in a local bubble, such as a local bubble formed by a mismatch of a guide sequence at a target editing site. In some embodiments, the cytosine deaminase can comprise a mutation that contributes to concentrated activity as described in Kim et al., Nature Biotechnology (2017) 35(4):371-377 (doi: 10.1038 / nbt.3803).

[0222] In some embodiments, the cytidine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squid, fish, flies, and worms. In some embodiments, the cytidine deaminase is a human, primate, bovine, canine, rat or mouse cytidine deaminase.

[0223] In some embodiments, the cytosine deaminase is a human APOBEC comprising hAPOBEC1 or hAPOBEC3. In some embodiments, the cytidine deaminase is human AID.

[0224] In some embodiments, the cytosine deaminase protein comprises one or more deaminase domains. While not wishing to be bound by theory, it is expected that the deaminase domain is used to identify one or more target cytosine (C) residues contained in a single-stranded bubble of an RNA duplex and convert them to uracil (U) residues.

[0225] In some embodiments, the engineered Cas12i effector protein is the main editor. A Cas9-based main editor is described, for example, in A. Anzalone et al., Nature, 2019, 576 (7785): 149-157, which is incorporated herein by general reference. In some embodiments, the engineered Cas12i effector protein comprises a nickase variant of one of the Cas12i nucleases engineered herein, fused with a reverse transcriptase domain. In some embodiments, the functional domain is a reverse transcriptase domain. In some embodiments, the reverse transcriptase domain is an M-MLV reverse transcriptase or a variant thereof, such as one or more mutants having D200N, T306K, W313F, T330P, and L603W. In some embodiments, an engineered CRISPR / Cas12i system comprising the main editor is provided. In some embodiments, the engineered CRISPR / Cas12i system also includes a second Cas12i nickase, for example, based on the same engineered Cas12i nuclease as the main editor. In some embodiments, the engineered CRISPR / Cas12i system includes a prime editing guide RNA (pegRNA) containing a primer binding site and a reverse transcriptase (RT) template sequence.

[0226] In some embodiments, the present application provides a split Cas12i effector system having one or more (e.g., 1, 2, 3, 4, 5, 6 or more) functional domains that associate (i.e., bind or fuse) with one or both of the split Cas12i effector protein portions. The functional domain can be provided as a fusion within a construct, as part of the first and / or second split Cas12i effector protein. The functional domain is generally fused to other portions in the split Cas12i effector protein (e.g., a split Cas12i effector protein portion) via a peptide linker, such as a GS linker. These functional domains can be used for the function of reconverting a split Cas12i effector system based on a catalytically inactivated Cas12i effector protein.

[0227] In some embodiments, the engineered Cas12i effector protein comprises one or more nuclear localization sequences (NLS) and / or one or more nuclear export sequences (NES). Exemplary NLS sequences include, for example, PKKKRKVPG (SEQ ID NO: 66) and ASPKKKRKV (SEQ ID NO: 67). The NLS and / or NES can be operably linked to the N-terminus and / or C-terminus of the engineered Cas12i effector protein, or to a polypeptide chain within the engineered Cas12i effector protein. In some embodiments, the NLS and / or NES can be linked to the N-terminus and / or C-terminus of any engineered Cas12i nuclease or functional variant thereof according to the present invention.

[0228] In some embodiments, the engineered Cas12i effector protein may encode additional components, such as a reporter protein. In some embodiments, the engineered Cas12i effector protein may include a fluorescent protein, such as GFP. Such a system can enable imaging of genomic sites (see, for example, "Dynamic Imaging of Genomic Loci in Living Human Cells by an Optimized CRISPR / Cas System," Chen B et al. Cell 2013). In some embodiments, the engineered Cas12i effector protein is an inducible, segmented Cas12i effector system useful for imaging genomic sites.

[0229] In some embodiments, we provide engineered Cas12i effector proteins that can induce double-strand or single-strand breaks in DNA molecules.

[0230] In some embodiments, engineered Cas12i effector proteins are provided, of which functional derivatives of the engineered Cas12i nuclease include Cas12i2 nuclease inactive variants comprising D599A, E833A, S883A, H884A, D886A, R900A, and / or D1019A (the amino acid positions are defined by the corresponding amino acid positions shown in SEQ ID NO:1), and Cas12i1 nuclease inactive variants comprising D647A, E894A, and / or D948A (the amino acid positions are defined by the corresponding amino acid positions shown in SEQ ID NO:13). Known enzyme-inactivating mutants of Cas12i2 nucleases, such as US10808245B2 and any Cas12i2 nuclease described in Huang X. et al., Nature Communications, 11, Article number: 5241 (2020), can be combined with the mutations in this application to provide functional derivatives of engineered Cas12i nucleases and their corresponding effector proteins.

[0231] Engineered CRISPR-Cas12i system In some embodiments, an engineered CRISPR-Cas12i system is provided comprising (a) any engineered Cas12i effector protein described herein (e.g., an engineered Cas12i nuclease or a functional variant thereof, e.g., CasXX shown in SEQ ID NO:8), and (b) a guide RNA comprising a guide sequence complementary to a target sequence, or one or more nucleic acids encoding the guide RNA; Among these, the engineered Cas12i effector protein and the guide RNA can form a CRISPR complex that includes the target nucleic acid of the target sequence and induces modification of the target nucleic acid (e.g., double-strand or single-strand breaks, base editing, etc.). In the context of this specification, the term “modification” includes nuclease cleavage, base editing, substitution, repair, etc., of target sites on the double-strand or single-strand of nucleic acid.

[0232] In some embodiments, the engineered CRISPR-Cas12i system comprises (a) any of the engineered Cas12i effector proteins described herein (e.g., any of the engineered Cas12i nucleases or their functional variants, or a nickas, split Cas12i, transcriptional repressor, transcriptional activator, base editor, or main editor based on the engineered Cas12i nuclease or its functional variant), and (b) a guide RNA containing a guide sequence complementary to the target sequence, or one or more nucleic acids encoding the guide RNA; of which the engineered Cas12i effector protein and the guide RNA can specifically bind to the target nucleic acid containing the target sequence and form a CRISPR complex that induces modification of the target nucleic acid (e.g., double-strand or single-strand breaks, base editing, etc.). In some embodiments, the engineered CRISPR-Cas12i system comprises one or more nucleic acids of the engineered Cas12i effector protein (e.g., an engineered Cas12i nuclease or a functional variant thereof, e.g., CasXX shown in SEQ ID NO:8) and / or the guide RNA. In some embodiments, the engineered CRISPR-Cas12i system comprises a precursor guide RNA array that can be processed into a plurality of crRNAs by, for example, the engineered Cas12i effector protein. In some embodiments, the engineered CRISPR-Cas12i system comprises one or more vectors encoding the engineered Cas12i effector protein and / or the guide RNA. In some embodiments, the engineered CRISPR-Cas12i system comprises a ribonucleoprotein (RNP) complex comprising the engineered Cas12i effector protein bound to the guide RNA.

[0233] The engineered CRISPR-Cas12i system of this application may include any suitable guide RNA. The guide RNA (gRNA) may include a guide sequence that can hybridize with a target sequence in a target nucleic acid, such as a target genomic region in a cell. In some embodiments, the gRNA includes a CRISPR RNA (crRNA) sequence containing the guide sequence.

[0234] Generally, the crRNAs described herein include direct repeat sequences and spacer sequences. In some embodiments, the crRNA includes, and is essentially composed of, a direct repeat sequence ligated to a guide sequence or a spacer sequence. In some embodiments, the crRNA includes a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-spacer sequence-DR), which is a typical feature of a precursor crRNA (pre-crRNA) structure. In some embodiments, the crRNA includes a truncated direct repeat sequence and a spacer sequence, which is a typical feature of a processed or mature crRNA. In some embodiments, the CRISPR-Cas12i effector protein forms a complex with an RNA guide sequence, and the spacer sequence induces the complex to sequence-specifically bind to a target nucleic acid complementary to the spacer sequence (e.g., at least 70% complementary).

[0235] In some embodiments, the guide RNA is a crRNA containing a guide sequence. In some embodiments, the engineered CRISPR-Cas12i system includes a precursor guide RNA array encoding multiple crRNAs. In some embodiments, the Cas12i effector protein cleaves the precursor guide RNA array to generate multiple crRNAs. In some embodiments, the engineered CRISPR-Cas12i system includes a precursor guide RNA array encoding multiple crRNAs, each crRNA containing a different guide sequence.

[0236] Structures and vectors This specification also provides constructs, vectors, and expression systems encoding any of the engineered Cas12i effector proteins described herein (e.g., engineered Cas12i nucleases or functional variants thereof). In some embodiments, the constructs, vectors, or expression systems further comprise one or more gRNA or crRNA arrays.

[0237] A “vector” is a composition of substances containing isolated nucleic acids that can be used to deliver the isolated nucleic acids into cells. Many vectors are known in the art and include, but are not limited to, linear polynucleotides, polynucleotides that associate with ions or amphiphilic compounds, plasmids, and viruses. Generally, a suitable vector includes a replication origin that acts in at least one organism, a promoter sequence, a convenient restrictional nucleotidase site, and one or more selective markers. The term “vector” should also be interpreted to include non-plasmids and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as polylysine compounds and liposomes.

[0238] In some embodiments, the vector is a viral vector. Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, lentivirus vectors, retrovirus vectors, vaccinia virus vectors, herpes simplex virus vectors, and their derivatives. In some embodiments, the vector is a phage vector. Viral vector technology is well known in the art and is described, for example, in Sambrook et al. (2001, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York) and other virology and molecular biology manuals.

[0239] Many virus-based systems have been developed to introduce genes into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. Heterogeneous nucleic acids may be inserted into vectors using techniques known in the art and packaged in retroviral particles. Recombinant viruses can then be isolated and delivered to mammalian cells in vitro or ex vivo. Many retroviral systems are known in the art. In some embodiments, adenovirus vectors are used. Many adenovirus vectors are known in the art. In some embodiments, lentiviral vectors are used. In some embodiments, self-inactivating lentiviral vectors are used.

[0240] In one embodiment, the vector is an adeno-associated virus (AAV) vector, such as AAV2, AAV8, or AAV9, which is at least 1 × 10⁻¹⁶ 5 Adenovirus or adeno-associated virus may be administered by a single dose containing particles (also called particle units or pu). In some embodiments, the dose is at least about 1 × 10⁻⁶. 6 The particles, at least about 1 × 10⁻¹⁶ 7 The particles, at least about 1 × 10⁻¹⁶ 8 A number of particles, or at least about 1 × 10⁶ 9 Adeno-associated virus particles. The delivery method and dosage are described in WO2016205764 and U.S. Patent No. 8454972, which are incorporated herein by general reference.

[0241] In some embodiments, the vector is a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, delivery by a modified AAV vector can be used. The modified AAV vector can be based on one or more of several capsid types, including AAV1, AV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6) and pseudo-AAVs (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6). Exemplary AAV vectors and techniques that can be used to generate rAAV particles are known in the art (see, for example, Aponte-Ubillus et al. (2018) Appl. Microbiol. Biotechnol. 102(3):1045-54; Zhong et al. (2012) J. Genet. Syndr. Gene Ther. S1: 008; West et al. (1987) Virology 160:38-47(1987); Tratschin et al. (1985) Mol. Cell. Biol. 5:3251-60; U.S. Patent Nos. 4797368 and 5173414, International Publication Nos. WO2015 / 054653 and WO93 / 24641, which are incorporated herein by reference, respectively).

[0242] Any known AAV vector for delivering Cas9 and other Cas proteins can be used to deliver the engineered Cas12i system of this application.

[0243] In some embodiments, the rAAV construct can be administered to a subject enterally. In some embodiments, the rAAV construct can be administered to the subject parenterally. In some embodiments, the rAAV particles can be injected into one or more cells, tissues or organs via subcutaneous, intraocular, intravitreal, subretinal, intravenous (IV), intraventricular, intramuscular, intrathecal (IT), intracisternal, intraperitoneal, inhalation, topical or direct injection. In some embodiments, the rAAV particles can be administered to a subject by injection into the hepatic artery or portal vein.

[0244] Methods for introducing vectors into mammalian cells are known in the art. Vectors can be transferred into host cells by physical, chemical, or biological methods.

[0245] Physical methods for introducing a vector into a host cell include calcium phosphate precipitation, lipofection, particle bombardment, microinjection, electroporation, and the like. Methods for producing cells comprising vectors and / or exogenous nucleic acids are well-known in the art. See, for example, Sambrook et al. (2001) Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York. In some embodiments, the vector is introduced into the cell by electroporation.

[0246] Biological methods for introducing a heterologous nucleic acid into a host cell include the use of DNA and RNA vectors. Viral vectors have become the most widely used method for inserting genes into mammals such as human cells.

[0247] Chemical methods for introducing vectors into host cells include colloidal dispersion systems such as polymer complexes, nanocapsules, microspheres, beads, and lipid-based systems, the lipid-based systems including oil-in-water emulsions, micelles, mixed micelles, and liposomes. A typical colloidal system used as an in vitro delivery vector is liposomes (e.g., artificial membrane vesicles). In some embodiments, the engineered CRISPR-Cas12i system is delivered in nanoparticles in the form of RNPs.

[0248] In some embodiments, the CRISPR-Cas12i system or the vector or expression system encoding its components includes one or more selectable or detectable markers that provide means for isolating or efficiently selecting cells modified by the CRISPR-Cas12i system (e.g., initial and large scale).

[0249] Reporter genes can be used to identify cells that may be transfected and to evaluate the function of regulatory sequences. Generally, a reporter gene is a gene that is not present or expressed in the receptor organism or tissue, and the expression of the encoded polypeptide is demonstrated by easily detectable properties such as enzymatic activity. The expression of the reporter gene is determined at an appropriate time after the DNA has been introduced into the receptor cell. Suitable reporter genes may include luciferase, β-galactosidase, chloramphenicol acetyltransferase, genes encoding secreted alkaline phosphatase, or green fluorescent protein genes (e.g., Ui-Tei et al. FEBS Letters 479:79-82 (2000)).

[0250] Other methods for confirming the presence of heterologous nucleic acids in host cells include molecular biological assays well known to those skilled in the art, such as Southern and Northern blotting, RT-PCR, and PCR, and biochemical assays that detect the presence or absence of specific peptides by immunological methods such as ELISA and Western blotting.

[0251] In some embodiments, the nucleic acid sequence encoding the engineered Cas12i effector protein and / or the guide RNA is operably ligated to a promoter. In some embodiments, the promoter is the endogenous promoter of the cell being engineered using the engineered CRISPR-Cas12i system. For example, the nucleic acid encoding the engineered Cas12i effector protein can be knocked in to be located downstream of the endogenous promoter in the genome of the engineered mammalian cell using any method known in the art. In some embodiments, the endogenous promoter is a promoter of an abundant protein (e.g., β-actin). In some embodiments, the endogenous promoter is an inducible promoter that can be induced, for example, by an endogenous activation signal in the engineered mammalian cell. In some embodiments, the engineered mammalian cell is a T cell, and the promoter is a T cell activation-dependent promoter (e.g., IL-2 promoter, NFAT promoter, or NFκB promoter).

[0252] In some embodiments, the promoter is a heterologous promoter for cells that are engineered and modified using the engineered CRISPR-Cas12i system. Multiple promoters have been explored for gene expression in mammalian cells, and any promoter known in the art may be used in this application. Promoters can be broadly classified into constitutive promoters or controlled promoters, such as inducible promoters.

[0253] In some embodiments, the nucleic acid sequence encoding the engineered Cas12i effector protein and / or guide RNA is operably ligated to a constitutive promoter. The constitutive promoter enables the constitutive expression of heterologous genes (also called transgenic genes) in host cells. Exemplary constitutive promoters considered herein include, but are not limited to, the cytomegalovirus (CMV) promoter, human elongation factor-1α (hEF1α), ubiquitin C promoter (UbiC), phosphoglycerol kinase promoter (PGK), Simiammavirus 40 early promoter (SV40), and the chicken β-actin promoter ligated to the CMV earliest enhancer (CAG). In some embodiments, the promoter is the CAG promoter, which includes the cytomegalovirus (CMV) earliest enhancer element, the promoter, the first exon and first intron of the chicken β-actin gene, and the splicing receptor of the rabbit β-globin gene.

[0254] In some embodiments, the engineered CRISPR-Cas12i effector protein and / or the nucleic acid sequence encoding the RNA is operably ligated to an inducible promoter. Inducible promoters belong to the controlled promoter type. Inducible promoters can be induced by one or more conditions, such as physical conditions, the microenvironment, or the physiological state of the host cell, an inducer (i.e., an inducer), or a combination thereof. In some embodiments, the inducing conditions are selected from the group of inducers, irradiation (e.g., ionizing radiation, light), temperature (e.g., heat), redox state, tumor environment, and the activation state of the cell to be engineered and modified by the engineered CRISPR-Cas12i system. In some embodiments, the promoter may be induced by a small molecule inducer, such as a compound. In some embodiments, the small molecule is selected from the group of doxycyclines, tetracyclines, alcohols, metals, or steroids. Chemically induced promoters are the most widely studied. Such promoters include small molecule chemicals whose transcriptional activity is modulated by the presence or absence of doxycyclines, tetracyclines, alcohols, steroids, metals, and other compounds. Doxycycline induction systems with inverse tetracycline-regulating transactivators (rtTAs) and tetracycline-responsive element promoters (TREs) are currently the most mature systems. WO9429442 describes the strict regulation of gene expression in eukaryotic cells by tetracycline-responsive promoters. WO9601313 discloses tetracycline-regulated transcription regulators. Furthermore, Tet technologies, such as the Tet-on system, are described, for example, on the TetSystems.com website. In this application, any known chemical regulatory promoter may be used to drive the expression of the engineered CRISPR-Cas12i protein and / or the guide RNA.

[0255] In some embodiments, the nucleic acid sequence encoding the engineered Cas12i effector protein is optimized by codons.

[0256] In some embodiments, an expression construct is provided that includes a codon-optimized sequence encoding the engineered Cas12i effector protein, which is conjugated to a BPK2104-ccdB vector. In some embodiments, the expression construct encodes a label (e.g., a 10×His label) operably ligated to the C-terminus of the engineered Cas12i effector protein.

[0257] In some embodiments, each engineered fragmented Cas12i construct encodes a fluorescent protein such as GFP or RFP. The reporter protein can be used to evaluate the co-localization and / or dimerization of the engineered Cas12i protein, for example, by microscopy. The nucleic acid sequence encoding the engineered Cas12i effector protein can be fused with a nucleic acid sequence encoding an additional component using a sequence encoding a self-cleaving peptide such as a T2A, P2A, E2A, or F2A peptide.

[0258] In some embodiments, an expression construct for use in mammalian cells (e.g., human cells) is provided, comprising a nucleic acid sequence encoding the engineered Cas12i effector protein. In some embodiments, the expression construct comprises a codon-optimized sequence encoding the engineered Cas12i effector protein inserted into a pCAG-2A-eGFP vector, thereby operably ligating the Cas12i protein to eGFP. In some embodiments, a second vector is provided for the expression of guide RNA (e.g., crRNA or a precursor crRNA array) in mammalian cells (e.g., human cells). In some embodiments, the sequence encoding the guide RNA is expressed in the backbone of the pUC19-U6-i2-cr RNA vector.

[0259] In some embodiments, one or more vectors expressing one or more elements of the CRISPR-Cas12i system are introduced into host cells, and the expression of the CRISPR-Cas12i system elements guides the formation of nucleic acid targeting complexes at one or more target sites. For example, the Cas12i nucleic acid targeting effector enzyme and the nucleic acid targeting guide RNA can be operably ligated to separate regulatory elements on separate vectors. The RNA of the nucleic acid targeting system can be delivered to transgenic Cas12i nucleic acid targeting effector protein animals or mammals, for example, animals or mammals that constitutively, inductively, or conditionally express the nucleic acid targeting effector protein, or can be expressed by other means, for example, by pre-administering one or more vectors encoding and expressing the nucleic acid targeting effector protein in vivo to animals or mammals having cells containing the nucleic acid targeting effector protein. Alternatively, two or more elements expressed by the same or different regulatory elements can be incorporated into a single vector, while one or more additional vectors provide any components of the nucleic acid targeting system not included in the first vector. Nucleic acid targeting elements incorporated into a single vector can be aligned in any suitable direction, for example, such that one element is located 5′ ("upstream") relative to a second element, or 3′ ("downstream") relative to a second element. The coding sequence of one element can be located on the sense or antisense strand of the coding sequence of a second element and can be aligned in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding Cas12i nucleic acid targeting effector protein and nucleic acid targeting guide RNA, the transcripts being incorporated within one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the nucleic acid targeting effector protein and nucleic acid targeting guide RNA can be operably ligated to the same promoter and expressed with the same promoter.Delivery media, vectors, particles, nanoparticles, formulations, and components for expressing one or more elements of a nucleic acid targeting system are as used in the aforementioned documents, such as WO2014 / 093622 (PCT / US2013 / 074667). In some embodiments, the vector includes one or more insertion sites, e.g., restriction endonuclease recognition sequences (also called “cloning sites”). In some embodiments, one or more insertion sites (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. When multiple different guide sequences are used, a single expression construct can be used to target nucleic acid targeting activity to multiple different corresponding target sequences in cells. For example, a single vector may contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences. In some embodiments, a vector containing about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more such guide sequences is provided and can optionally be delivered to cells. In some embodiments, the vector includes a regulatory element operably linked to an enzyme coding sequence encoding a nucleic acid-targeting effector protein. The Cas12i nucleic acid-targeting effector protein or one or more nucleic acid-targeting guide RNAs can be delivered separately, and advantageously, at least one of these can be delivered via a particle complex. To allow time for the expression of the Cas12i nucleic acid-targeting effector protein, the nucleic acid-targeting effector protein mRNA can be delivered before the nucleic acid-targeting guide RNA. The nucleic acid-targeting effector protein mRNA can be administered 1 to 12 hours (preferably about 2 to 6 hours) before the administration of the nucleic acid-targeting guide RNA. Alternatively, the nucleic acid-targeting effector protein mRNA and the nucleic acid-targeting guide RNA can be administered together. Advantageously, a second enhancement dose of guide RNA can be administered 1 to 12 hours (preferably about 2 to 6 hours) after the initial administration of nucleic acid-targeted effector protein mRNA + guide RNA.Directing other administrations of nucleic acid-targeted effector protein mRNA and / or guide RNA may be useful in achieving the most effective level of genomic modification.

[0260] In some embodiments, a CRISPR-Cas12i system is provided comprising (1) an engineered Cas12i effector protein (an engineered Cas12i nuclease or a functional variant thereof) as described in any of the embodiments, or a polynucleotide encoding an engineered Cas12i effector protein as described in any of the embodiments, and (2) a crRNA or a polynucleotide encoding a crRNA, wherein the crRNA comprises (i) a spacer sequence that can hybridize with a target sequence of target DNA, and (ii) a direct repeat sequence ligated to the spacer sequence that can induce the engineered Cas12i effector protein to bind to the crRNA in order to form a CRISPR-Cas12i complex targeting the target sequence.

[0261] In some embodiments, a CRISPR-Cas12i system is provided, comprising one or more vectors, the one or more vectors comprising (1) a first regulatory element operably ligated to a nucleotide sequence encoding one of the engineered Cas12i effector proteins (e.g., engineered Cas12i nuclease or a functional variant thereof), and (2) a second regulatory element operably ligated to a polynucleotide encoding a crRNA. The crRNA comprises (i) a spacer sequence that can hybridize with a target sequence of target DNA, and (ii) a direct repeat sequence ligated to the spacer sequence that can induce the binding of the engineered Cas12i effector protein to the crRNA to form a CRISPR-Cas12i complex targeting the target sequence, wherein the first regulatory element and the second regulatory element are located on the same or different vectors of the CRISPR-Cas12i system.

[0262] In some embodiments, the first and second control elements are located on different vectors of the CRISPR-Cas12i system. In some embodiments, the first and second control elements are located on the same vector of the CRISPR-Cas12i system. In some embodiments, the first control element and the nucleotide sequence encoding the engineered Cas12i effector protein are upstream of the second control element and the polynucleotide encoding the crRNA. In some embodiments, the first control element and the nucleotide sequence encoding the engineered Cas12i effector protein are downstream of the second control element and the polynucleotide encoding the crRNA. In some embodiments, the first and second control elements are the same. In some embodiments, the first and second control elements are different.

[0263] In some embodiments, a CRISPR-Cas12i system is provided which comprises a vector comprising (1) a polynucleotide encoding one of the engineered Cas12i effector proteins (either an engineered Cas12i nuclease or a functional variant thereof), and (2) a polynucleotide encoding a crRNA, wherein the crRNA comprises (i) a spacer sequence that can hybridize with a target sequence of target DNA, (ii) a direct repeat sequence ligated to the spacer sequence that can induce the engineered Cas12i effector protein to bind to the crRNA to form a CRISPR-Cas12i complex that targets the target sequence, and (3) a control element operably ligated to the polynucleotide encoding the engineered Cas12i effector protein and the polynucleotide encoding the crRNA.

[0264] In some embodiments, the vector comprises, from 5' to 3', the control element, a polynucleotide encoding the engineered Cas12i effector protein, and a polynucleotide encoding the crRNA. In some embodiments, the vector comprises, from 5' to 3', the control element, a polynucleotide encoding the crRNA, and a polynucleotide encoding the engineered Cas12i effector protein. In some embodiments, the polynucleotide encoding the crRNA and the polynucleotide encoding the engineered Cas12i effector protein are linked by a linker sequence, for example, a polynucleotide sequence encoding any of P2A, T2A, E2A, F2A, BmCPV2A, BmIFV2A, (GS)n (SEQ ID NO: 74), (GGS)n (SEQ ID NO: 75), and (GGGS)n (SEQ ID NO: 76) (where n is at least an integer of 1), or any polynucleotide sequence of IRES, SV40, CMV, UBC, EF1α, PGK, and CAGG, or any combination thereof.

[0265] In some embodiments, the components of the CRISPR-Cas12i system described in the present invention may be delivered in various forms, such as DNA / RNA, RNA / RNA, or protein-RNA combinations. For example, the engineered Cas12i effector protein (either the engineered Cas12i nuclease or its functional variant) may be delivered as a DNA-coding polynucleotide, or as an RNA-coding polynucleotide, or as a protein. The guide may be delivered as a DNA-coding polynucleotide or RNA. A mixed delivery method may also be used.

[0266] In some embodiments, the present invention provides a method for delivering one or more polynucleotides, for example, one or more vectors described herein, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell.

[0267] III.How to use This application provides methods for detecting target nucleic acids or modified nucleic acids in vitro, ex vivo, or in vivo using either the engineered Cas12i effector protein (either the engineered Cas12i nuclease or a functional variant thereof) or the CRISPR-Cas12i system described herein, and methods for therapeutic (e.g., gene editing) or diagnostics using the engineered Cas12i effector protein or the CRISPR-Cas12i system described herein. It also provides the engineered Cas12i effector protein or the CRISPR-Cas12i system described herein for detecting or modifying nucleic acids in cells, and applications for treating or diagnosing diseases or conditions in subjects, and applications for compositions comprising any of the engineered Cas12i effector proteins or one or more components of the engineered Cas12i effector protein in the manufacture of pharmaceuticals for detecting or modifying nucleic acids in cells, and for treating or diagnosing diseases or conditions in subjects.

[0268] Method for detecting target nucleic acids in a sample This application also provides a method for detecting target nucleic acids using either the engineered Cas12i effector protein or the CRISPR-Cas12i system having improved activity. Using the Cas12i effector protein as a detection reagent allows the V-type CRISPR / Cas protein (e.g., Cas12i), once activated by detecting target DNA, to intermittently cleave non-target single-stranded DNA (ssDNA or RNA, i.e., single-stranded nucleic acids whose guide sequence does not hybridize with the guide RNA). Therefore, if target DNA (double-stranded or single-stranded) is present in the sample (e.g., exceeding a threshold amount in some cases), the result is cleavage of single-stranded nucleic acids in the sample, which can be detected using any convenient detection method (e.g., detection using labeled single-stranded DNA or RNA). Cas12i can cleave both ssDNA and ssRNA. For example, methods using the Cas protein as a detection reagent are described in US10253365 and WO2020 / 056924, which are incorporated herein by general reference.

[0269] In some embodiments, provided is a method of detecting a target DNA (e.g., double-stranded or single-stranded) in a sample, comprising: (a) contacting the sample with any of: (i) an engineered Cas12i effector protein described herein (any engineered Cas12i nuclease or a functional variant thereof), (ii) a guide RNA comprising a guide sequence for hybridizing to the target DNA, and (iii) a detection nucleic acid that is single-stranded (i.e., a "single-stranded detection nucleic acid") and does not hybridize to the guide sequence of the guide RNA; and (b) measuring a detectable signal generated by cleavage of the single-stranded detection nucleic acid by the engineered Cas12i effector protein. In some cases, the single-stranded detection nucleic acid comprises a dye pair that emits fluorescence (e.g., the fluorescence-emitting dye pair is a fluorescence resonance energy transfer (FRET) pair, a quencher / fluorophore pair). In some cases, the target DNA is viral DNA (e.g., papillomavirus, hepadnavirus, herpesvirus, adenovirus, poxvirus, parvovirus, etc.). In some embodiments, the single-stranded detection nucleic acid is DNA. In some embodiments, the single-stranded detection nucleic acid is RNA.

[0270] The method for detecting target DNA (single-stranded or double-stranded) in a sample of the present disclosure can detect target DNA with high sensitivity. In some cases, the method of the present disclosure can be used to detect target DNA present in a sample comprising a plurality of DNAs (including said target DNA and a plurality of non-target DNAs), wherein the target DNA is present at 10 non-target DNAs 7 present in one or more copies per (e.g., 10 non-target DNAs 6 one or more copies per, 10 non-target DNAs 5 one or more copies per, 10 non-target DNAs 4 one or more copies per, 10 non-target DNAs 3 one or more copies per, 10 non-target DNAs 2(One or more copies per individual, one or more copies per 50 non-target DNA molecules, one or more copies per 20 non-target DNA molecules, one or more copies per 10 non-target DNA molecules, or one or more copies per 5 non-target DNA molecules). In some embodiments, the engineered Cas12i effector proteins described herein can detect target DNA with higher sensitivity than the reference Cas12i nuclease. In some embodiments, the engineered Cas12i effector proteins can detect target DNA with sensitivity of 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or higher compared to the reference Cas12i nuclease.

[0271] Method of modification In some embodiments, the application provides a method for modifying a target nucleic acid comprising a target sequence, the method comprising contacting the target nucleic acid with one of the engineered CRISPR-Cas12i systems described herein. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is present in a cell. In some embodiments, the cell is a bacterial cell, yeast cell, mammalian cell, plant cell, or animal cell. In some embodiments, the method is performed ex vivo. In some embodiments, the method is performed in vivo. Modification of the target nucleic acid includes, but is not limited to, single-strand breaks, double-strand breaks, base substitutions, base insertions, base deletions, mutations (e.g., pathogenic mutations), and sequence repair of the target nucleic acid.

[0272] In some embodiments, the engineered CRISPR-Cas12i system cleaves the target nucleic acid or alters the target sequence within the target nucleic acid. In some embodiments, the expression of the target nucleic acid is altered by the engineered CRISPR-Cas12i system. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target sequence is associated with a disease or condition based on, for example, the misexpression (overexpression, nonexpression, or expression of pathogenic RNA or protein) of the target sequence. In some embodiments, the engineered CRISPR-Cas12i system includes a precursor guide RNA array encoding multiple crRNAs, each crRNA containing a different guide sequence.

[0273] In some embodiments, this application provides a method for treating a disease or condition associated with a target nucleic acid in an individual cell, comprising modifying the target nucleic acid in the individual cell for the purpose of treating the disease or condition using any of the methods described herein. In some embodiments, the disease or condition is selected from the group of cancer, cardiovascular disease, genetic disorders (such as sickle cell anemia (SCD) or β-thalassemia (TDT)), autoimmune diseases, metabolic diseases, neurodegenerative diseases, ophthalmic diseases, bacterial infections, and viral infections.

[0274] The engineered CRISPR-Cas12i systems described herein can modify target nucleic acids in cells in a variety of ways, depending on the type of Cas12i effector protein engineered in the CRISPR-Cas12i system. In some embodiments, the method induces site-specific cleavage in the target nucleic acid. In some embodiments, the method cleaves genomic DNA in cells, e.g., bacterial cells, plant cells, or animal cells (e.g., mammalian cells). In some embodiments, the method kills cells by cleaving genomic DNA in them. In some embodiments, the method cleaves viral nucleic acids in cells.

[0275] In some embodiments, the method alters (e.g., increases or decreases) the expression level of the target nucleic acid in cells. In some embodiments, the method enhances the expression level of the target nucleic acid in cells by using an engineered Cas12i effector protein based on, for example, a non-enzymatic Cas12i protein fused to a transactivation domain. In some embodiments, the method reduces the expression level of the target nucleic acid in cells by using an engineered Cas12i effector protein based on, for example, a non-enzymatic Cas12i protein fused to a transcriptional repression domain. In some embodiments, the method introduces epigenetic modifications to the target nucleic acid or a target nucleic acid-related protein (e.g., a protein bound to or near the target nucleic acid, e.g., a transcription factor or histone) in cells by using an engineered Cas12i effector protein based on, for example, a non-enzymatic Cas12i protein fused to an epigenetic modification domain. The engineered Cas12i systems described herein can be used to introduce other modifications to the target nucleic acid depending on the functional domains contained in the engineered Cas12i effector protein.

[0276] In some embodiments, the method alters the target sequence in the target nucleic acid in a cell. In some embodiments, the method introduces a mutation into the target nucleic acid in a cell, for example, by changing a pathogenic mutation to a non-pathogenic sequence (e.g., by either the dCas12i of the present invention fused to adenosine deaminase or cyticine deaminase). In some embodiments, the method repairs double-strand breaks induced in the target DNA in a cell using one or more endogenous DNA repair pathways, such as non-homologous end joining (NHEJ) or homology-oriented repair (HDR), as a result of sequence-specific cleavage by the CRISPR complex. Exemplary mutations include, but are not limited to, insertions, deletions, substitutions, and frameshifts. In some embodiments, pegRNA is used to provide a DNA repair template by simultaneously cleaving the target sequence. In some embodiments, the method inserts donor DNA into the target site. In some embodiments, the insertion of donor DNA results in the introduction of a selection marker or reporter protein into the cell. In some embodiments, the insertion of donor DNA results in gene knock-in. In some embodiments, the insertion of donor DNA results in a knockout mutation. In some embodiments, the insertion of donor DNA results in a substitutional mutation, such as a mononucleotide substitution. In some embodiments, the method induces a phenotypic change in the cell.

[0277] In some embodiments, the engineered CRISPR-Cas12i system is used as part of a genetic circuit, or to insert a genetic circuit into the genomic DNA of a cell. The engineered, fragmented Cas12i effector proteins controlled by the inducers described herein can be used, in particular, as components of a genetic circuit. The genetic circuit can be used in gene therapy. Methods and techniques for designing and using genetic circuits are known in the art. Furthermore, see, for example, Brophy, Jennifer AN, and Christopher A. Voigt. “Principles of genetic circuit design.” Nature methods 11.5 (2014): 508.

[0278] The engineered CRISPR-Cas12i systems described herein can be used to modify multiple target nucleic acids. In some embodiments, the target nucleic acid is located in a cell. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is extrachromosomal DNA. In some embodiments, the target nucleic acid is exogenous to the cell. In some embodiments, the target nucleic acid is viral nucleic acid, such as viral DNA. In some embodiments, the target nucleic acid is a plasmid in a cell. In some embodiments, the target nucleic acid is a horizontally transferred plasmid. In some embodiments, the target nucleic acid is RNA.

[0279] In some embodiments, the target nucleic acid is an isolated nucleic acid, such as isolated DNA. In some embodiments, the target nucleic acid is present in a cell-free environment. In some embodiments, the target nucleic acid is a vector, such as an isolated plasmid. In some embodiments, the target nucleic acid is an isolated linear DNA fragment.

[0280] The methods described herein are suitable for any suitable cell type. In some embodiments, the cells are bacterial, yeast, fungal, algal, plant, or animal cells (e.g., mammalian cells such as human cells). In some embodiments, the cells are naturally occurring cells, such as cells isolated by tissue biopsy. In some embodiments, the cells are cells isolated from an in vitro cultured cell line. In some embodiments, the cells are derived from a primary cell line. In some embodiments, the cells are derived from an immortal cell line. In some embodiments, the cells are genetically engineered cells.

[0281] In some embodiments, the cells are animal cells of organisms selected from a group of cattle, sheep, goats, horses, pigs, deer, chickens, ducks, geese, rabbits, and fish.

[0282] In some embodiments, the cells are plant cells of organisms selected from the group consisting of corn, wheat, barley, oats, rice, soybeans, oil palm, safflower, sesame, tobacco, flax, cotton, sunflower, pearl millet, foxtail millet, sorghum, canola, cannabis, vegetable crops, fodder crops, industrial crops, woody crops, and biomass crops.

[0283] In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the human cells are human embryonic kidney 293T (HEK293T or 293T) cells or HeLa cells. In some embodiments, the cells are human embryonic kidney (HEK293T) cells. In some embodiments, the cells are mouse Hepa1-6 cells. In some embodiments, the mammalian cells are selected from the group consisting of immune cells, hepatocytes, tumor cells, stem cells, blood cells, nerve cells, zygotes, muscle cells (e.g., cardiomyocytes), and skin cells.

[0284] In some embodiments, the cells are immune cells selected from a group of T cells that are activated by cytotoxic T cells, helper T cells, natural killer (NK) T cells, iNK-T cells, NK-T-like cells, γδ T cells, tumor-infiltrating T cells, and dendritic cells (DCs). In some embodiments, the method produces modified immune cells such as CAR-T cells or TCR-T cells.

[0285] In some embodiments, the cells are embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, gamete progenitor cells, gametes, zygotes, or cells within an embryo.

[0286] The methods described herein can be used to modify target cells in vivo, ex vivo, or in vitro, and the modification can be carried out in a manner that alters the cells so that the offspring or cell lines of the modified cells retain the modified phenotype. The modified cells and offspring may be part of a multicellular organism such as a plant or animal having ex vivo or in vivo applications such as genome editing and gene therapy.

[0287] In some embodiments, the method is performed ex vivo. In some embodiments, after introducing the engineered CRISPR-Cas12i system into cells, the modified cells (e.g., mammalian cells) are grown ex vivo. In some embodiments, the modified cells are cultured and grown for at least about 1, 2, 3, 4, 5, 6, 7, 10, 12, or 14 days. In some embodiments, the modified cells are cultured for no more than about 1, 2, 3, 4, 5, 6, 7, 10, 12, or 14 days. In some embodiments, the modified cells are further evaluated or screened to select one or more cells having a desired phenotype or characteristics.

[0288] In some embodiments, the target sequence is a sequence associated with a disease or condition. Examples of diseases or conditions include, but are not limited to, cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, ophthalmological disease, bacterial infection, and viral infection. In some embodiments, the disease or condition is a genetic disease. In some embodiments, the disease or condition is a monogenic disease or condition. In some embodiments, the disease or condition is a polygenic disease or condition.

[0289] In some embodiments, the target sequence has mutations compared to the wild-type sequence. In some embodiments, the target sequence has a single nucleotide polymorphism (SNP) associated with a disease or condition.

[0290] In some embodiments, the donor DNA inserted into the target nucleic acid encodes a biological product selected from the group consisting of reporter proteins, antigen-specific receptors, therapeutic proteins, antibiotic-resistant proteins, RNAi molecules, cytokines, kinases, antigens, antigen-specific receptors, cytofactor receptors, and suicide polypeptides. In some embodiments, the donor DNA encodes a therapeutic protein. In some embodiments, the donor DNA encodes a therapeutic protein that can be used in gene therapy. In some embodiments, the donor DNA encodes a therapeutic antibody. In some embodiments, the donor DNA encodes an engineered receptor such as a chimeric antigen receptor (CAR) or an engineered TCR. In some embodiments, the donor DNA encodes a therapeutic RNA such as a small RNA (e.g., siRNA, shRNA, or miRNA) or a long non-coding RNA (lincRNA).

[0291] The methods described herein can be used for multiple gene editing or regulation at two or more (e.g., 2, 3, 4, 5, 6, 8, 10 or more) different target sites. In some embodiments, the methods detect or modify multiple target nucleic acids or target nucleic acid sequences. In some embodiments, the methods involve contacting target nucleic acids with a guide RNA containing multiple (e.g., 2, 3, 4, 5, 6, 8, 10 or more) crRNA sequences, each containing a different target sequence.

[0292] Engineered cells containing modified target nucleic acids, generated using any of the methods described herein, are also provided. These engineered cells can be used in cell therapy. Engineered cells can be prepared from autologous or allogeneic cells using the cell therapy methods described herein.

[0293] The methods described herein can also be used to generate isogenic lines of cells (e.g., mammalian cells) for studying genetic variants.

[0294] Engineered non-human animals, including the engineered cells described herein, are also provided. In some embodiments, the engineered non-human animals are genome-edited non-human animals. The engineered non-human animals can be used as disease models.

[0295] Techniques for generating non-human genome-edited or transgenic animals are well known in the art, including, but are not limited to, pronuclear microinjection, viral infection, and transformation into embryonic stem cells and induced pluripotent stem (iPS) cells. Detailed methods that can be used include, but are not limited to, those described by Sundberg and Ichiki (2006, Genetically Engineered Mice Handbook, CRC Press) and those described by Gibson (2004, A Primer Of Genome Science 2nd ed. Sunderland, Mass.: Sinauer).

[0296] The engineered animals may be, but are not limited to, any suitable species, including cattle, horses, sheep, dogs, deer, felines, goats, pigs, primates, and less understood mammals such as elephants, deer, zebras, or camels.

[0297] Treatment method In some embodiments, applications of the engineered CRISPR-Cas12i system are provided in the manufacture of pharmaceuticals for treating diseases or disorders related to target nucleic acids in the cells of an individual. In some embodiments, methods are provided for treating diseases or disorders related to target nucleic acids in the cells of an individual using the engineered CRISPR-Cas12i system.

[0298] In some embodiments, the present invention provides a method for treating a disease in a subject (e.g., a human) that is required, the method comprising administering the CRISPR-Cas12i system of the present invention to the subject (e.g., by intravenous injection or infusion), the CRISPR-Cas12i system comprising (1) any engineered Cas12i effector protein (e.g., any engineered Cas12i nuclease or any functional variant thereof), or a polynucleotide encoding an engineered Cas12i effector protein, and (2) a crRNA, or a polynucleotide encoding the crRNA, the crRNA having (i) a target nucleic acid related to the disease (ii) a spacer sequence that can hybridize with a target sequence, and (ii) a direct repeat sequence linked to the spacer sequence that can induce the engineered Cas12i effector protein to bind to the crRNA to form a CRISPR-Cas12i complex that targets the target sequence, wherein the hybridization of the spacer sequence and the target sequence mediates contact between the engineered Cas12i effector protein and the target sequence, resulting in the engineered Cas12i effector protein modifying (e.g., cleaving or base editing) the target sequence, thereby treating the subject's disease. Any of the components of the CRISPR-Cas12i system may be delivered simultaneously, sequentially, or in any form of DNA / RNA, DNA / DNA, RNA / RNA, protein / RNA, or protein / DNA.

[0299] Further, the present invention provides therapeutic methods using any of the methods for modifying a target nucleic acid in a cell described herein. In some embodiments, the present invention provides a method for treating a disease or condition associated with a target nucleic acid in an organismal cell, comprising contacting the target nucleic acid with any of the engineered CRISPR-Cas12i systems described herein, wherein the guide sequence of the guide RNA is complementary to the target sequence of the target nucleic acid, and the engineered Cas12i effector protein and the guide RNA associate with each other to bind to the target nucleic acid in order to modify (e.g., cleave or substitute) the target nucleic acid, thereby treating the disease or condition. In some embodiments, a mutation (e.g., knockout or knockout mutation) is introduced into the target nucleic acid. In some embodiments, the expression of the target nucleic acid is enhanced. In some embodiments, the expression of the target nucleic acid is suppressed. In some embodiments, the present application provides a method for treating a disease or condition in an individual, comprising administering to the individual an effective amount of one of the engineered CRISPR-Cas12i systems described herein and donor DNA encoding a therapeutic agent (e.g., using a non-pathogenic natural sequence as a repair template), wherein the guide sequence of the guide RNA is complementary to the target sequence of the target nucleic acid of the individual, and the engineered Cas12i effector protein and the guide RNA bind to each other to bind to the target nucleic acid and insert the donor DNA into the target sequence, thereby treating the disease or condition.

[0300] In some embodiments, the present application provides a method for treating a disease or condition in an individual, comprising administering an effective dose to the individual engineered engineered cells containing a modified target nucleic acid, which are produced by contacting the cells with one of the CRISPR-Cas12i systems engineered herein, wherein the guide sequence of the guide RNA is complementary to the target sequence of the target nucleic acid, and the engineered Cas12i effector protein and the guide RNA associate with each other to bind to the target nucleic acid in order to modify the target nucleic acid. In some embodiments, the engineered cells are immune cells. In some embodiments, the engineered cells are stem cells such as hematopoietic stem cells and neural stem cells. In some embodiments, the engineered cells are nerve cells. In some embodiments, the individual is a human. In some embodiments, the individual is a model animal such as a rodent, pet, or farm animal. In some embodiments, the individual is a mammal such as a cat, dog, rabbit, or hamster.

[0301] In some embodiments, the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, ophthalmological disease, bacterial infection, and viral infection. In some embodiments, the target nucleic acid is PCSK9. In some embodiments, the disease or condition is cardiovascular disease. In some embodiments, the disease or condition is coronary artery disease. In some embodiments, the method lowers the cholesterol levels of an individual. In some embodiments, the method treats diabetes in an individual.

[0302] Disorders or diseases that can be treated using the CRISPR-Cas12i system of the present invention include, but are not limited to, cystic fibrosis, hereditary angioedema, diabetes mellitus, progressive pseudohypertrophic muscular dystrophy, Baker muscular dystrophy, α-1-antitrypsin deficiency, Pompeii disease, tonic muscular dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell anemia, or β-thalassemia. In some embodiments, the disorder or disease is a transthyretin amyloidosis such as wild-type transthyretin amyloidosis (ATTRwt), hereditary transthyretin amyloidosis (ATTRm), familial amyloid polyneuropathy (FAP, ATTR-PN), or familial amyloid cardiomyopathy (FAC, ATTR-CM). In some embodiments, the disorder or disease is instability of the transthyroxine protein due to mutation or abnormal expression (e.g., high expression) of the TTR gene. In some embodiments, the disorder or disease is another disorder or disease resulting from TTR gene mutation or abnormal expression (e.g., high expression), or a derivative disorder or disease.

[0303] Delivery method In some embodiments, the engineered CRISPR-Cas12i system or its components, the nucleic acid molecules thereof, or the nucleic acid molecules encoding or providing its components can be delivered to host cells (e.g., any of the vectors described in the “Constructions and Vectors” section above) via various delivery systems such as plasmids or viruses. In some embodiments or methods, the engineered CRISPR-Cas12i system can be delivered by other methods such as nucleofection or electroporation of a ribonucleoprotein complex comprising the engineered Cas12i effector protein and one or more homologous RNA guide sequences thereof.

[0304] In some embodiments, delivery is via nanoparticles or exosomes.

[0305] In some embodiments, paired Cas12i nickase complexes can be delivered directly using nanoparticles or other direct protein delivery methods so that the complexes containing the two paired crRNA elements are delivered together. Furthermore, proteins can be delivered via viral vectors or directly to cells, and then a CRISPR array containing two pairs of spacer regions for double nick can be delivered directly. In some cases, for direct RNA delivery, RNA can be conjugated to at least one sugar moiety, such as N-acetylgalactosamine (GalNAc) (particularly triangle GalNAc).

[0306] In some embodiments or methods, the engineered CRISPR-Cas12i system can be delivered using any therapeutically appropriate delivery method, such as delivery by intravenous injection or infusion, or local delivery to the disease site, such as intratumoral delivery. Suitable methods for administering the engineered CRISPR-Cas12i system described herein include, but are not limited to, local, subcutaneous, percutaneous, intradermal, intrafocal, intraarticular, intraperitoneal, intrabladderal, transmucosal, gingival, intradental, intracochlear, transtympanic, intraorgan, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseous, periocular, intratumoral, intracerebral, and intraventricular delivery. In some embodiments, the engineered CRISPR-Cas12i system described herein is administered to a subject via injection, catheter, suppository, or implant, the implant being a porous, non-porous, or gel-like material containing a film such as a sialic acid membrane, or fibers.

[0307] IV. Kits and Products Compositions, kits, unit drugs, and products comprising one or more components of the engineered Cas12i nuclease or its functional variant, the engineered Cas12i effector protein, or the engineered CRISPR-Cas12i system described herein are also provided.

[0308] In some embodiments, a kit is provided which comprises one or more AAV vectors encoding either an engineered Cas12i nuclease or a functional variant thereof, an engineered Cas12i effector protein, or an engineered CRISPR-Cas12i system as described herein. In some embodiments, the kit further comprises one or more guide RNAs (or DNA or vectors encoding them). In some embodiments, the kit further comprises donor DNA. In some embodiments, the kit further comprises cells such as human cells.

[0309] The kit may include one or more additional components, such as containers, reagents, culture media, cytokines, buffers, and antibodies, to enable the proliferation of engineered cells. The kit may also include a device for administering the composition.

[0310] The kit may also include protocols for using the engineered CRISPR-Cas12i system described herein, such as methods for detecting or modifying target nucleic acids. In some embodiments, the kit includes protocols for treating or diagnosing a disease or condition. Protocols relating to the use of the kit's components typically include information on intentionally treated doses, administration schedules, and routes of administration. The containers may be unit doses, bulk packaging (e.g., multi-dose packaging), or subunit doses. For example, such a kit (providing a sufficient quantity of the compositions disclosed herein) can provide effective treatment to an individual over an extended period. The kit may also include multiple unit dose compositions and instructions for use, with packaging quantities sufficient for storage and use in a pharmacy (e.g., hospital pharmacies and compound pharmacies).

[0311] The kit of the present invention is contained in a suitable package. Suitable packaging includes, but is not limited to, vials, bottles, wide-mouth bottles, and flexible packaging (e.g., sealed polyester film or plastic bags). The kit may optionally provide additional components such as buffer solutions and explanatory information. Therefore, this application also provides products including vials (such as sealed vials), bottles, wide-mouth vials, flexible packaging, etc.

[0312] The product may include a container and a label or packaging insert attached to or on the container. Suitable containers include, for example, bottles, vials, and syringes. The container may be made of various materials such as glass or plastic. Generally, the container may contain a composition for effectively treating a disease or disorder described herein and may have a sterile entrance (for example, the container may be an intravenous solution bag or a vial with a stopper that can be punctured with a subcutaneous needle). The label or packaging insert indicates that the composition is used to treat a specific condition in an individual. The label or packaging insert may further include a protocol for administering the composition to the individual.

[0313] A package insert typically refers to a description included in the commercial packaging of a therapeutic product that contains information regarding the indications, usage, dosage, administration, contraindications, and / or warnings for use of such therapeutic product.

[0314] Furthermore, the product may include a second container containing pharmaceutically acceptable buffers such as bacteriostatic water for injection (BWFI), phosphate-buffered saline, Ringer's solution, and glucose solution. From a business and user perspective, other materials such as other buffers, diluents, filters, needles, and syringes may also be included. If necessary, solubilizers and local anesthetics (e.g., lidocaine) may also be included to alleviate pain at the injection site.

[0315] Generally, the components are provided individually or as a mixture in unit dosage forms such as dry lyophilized powder or anhydrous concentrate in a sealed container indicating an active dose, such as an ampoule or pouch. When the drug is administered by intravenous fluid, it can be prepared in an infusion bottle containing sterile, pharmaceutically acceptable water or saline solution. When the pharmaceutical composition is administered by injection, ampoules of sterile water for injection or saline solution can be provided so that the components can be mixed before administration. Example section

[0316] Specific embodiments of the present invention will be described in more detail below with reference to the drawings. While the drawings illustrate specific embodiments of the present invention, it should be understood that the present invention can be realized in various forms, without being limited to the embodiments described herein. In contrast, these embodiments are provided to allow for a more complete understanding of the present invention and to fully convey its scope to those skilled in the art.

[0317] Example 1: The gene editing efficiency was verified by substituting the amino acid that interacts with PAM in the reference Cas12i2 enzyme with a positively charged amino acid. Plasmid construction The Cas12i2 coding sequence was synthesized using codon optimization (human). Variants of the Cas protein were generated by PCR-based site-directed mutagenesis. Specifically, the Cas12i2 protein DNA sequence was divided into two parts around the mutation site, two pairs of primers were designed to amplify the DNA sequences of these two parts, and the necessary mutation sequence was introduced into the primers simultaneously. Finally, the two fragments were incorporated into a pCAG-2A-eGFP vector (penicillin-resistant) using the Gibson cloning method. The mutant combinations were constructed by multi-stage splitting of the Cas12i2 protein DNA and using PCR and Gibson cloning. The location of the mutants was determined using protein visualization software commonly used in this field (e.g., PyMol, Chimera) to analyze the structural information of Cas12i2. Structural information for Cas12i2 was referenced from PDB: 6LTU, 6LTR, 6LU0, 6LTP). The Cas12i2 effector protein was expressed in human 293T cells via a pCAG-2A-eGFP vector. The DNA encoding the Cas12i2 protein was inserted between XmaI and NheI. A crRNA vector expressing Cas12i2 in 293T cells was constructed by annealing an oligonucleotide containing the target sequence to a pUC19-U6-i2-crRNA backbone digested with BasI. The nucleotide sequence encoding DR is SEQ ID NO: 59.

[0318] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well cell dishes (Corning) and cultured for 16 hours until cell density reached 70%. Each 24-well cell dish was transfected with 600 ng of a plasmid encoding Cas12i2 protein and 300 ng of a plasmid encoding crRNA using Lipofectamine 3000 (Invitrogen). After approximately 72 hours of transfection, HEK293T cells to be subjected to fluorescence-activated cell sorting (FACS) were digested with trypsin-EDTA (0.05%) (Gibco). Cell sorting was performed using MoFlo XDP (Beckman Coulter) with a GFP channel.

[0319] Targeted deep sequencing analysis for genome modification GFP-positive HEK293T cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of targets, cell degradation saturates were used directly as templates, and target sites were directly amplified by barcoded PCR. PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. In this application, the metric of frequency (%) of insertions and deletions is used consistently for the comparison and analysis of gene editing efficiency. Reads less than 0.05% of complete reads are discarded.

[0320] Example 1-A Four engineered Cas12i2 molecules with a single amino acid substitution were selected. Engineered Cas12i2 enzymes having a single amino acid sequence mutation were expressed according to the method described in Example 1, and preferred amino acid substitution methods and their corresponding gene editing efficiencies are shown in Figures 1a, 1b, and Table 1. We first selected 10 amino acids E176, E178, Y226, A227, N229, E237, K238, K264, T447, and E563, where Cas12i2 is within 9 Å of the PAM DNA, and performed a point mutation test for arginine(R). As shown in Figure 1a and Table 1, by comparing the gene editing efficiency of these Cas12i mutants and wild-type Cas12i2 (SEQ ID NO:1) at two genomic sites: CCR5-3 and RNF2-7 in 293T cells, Cas12i mutants with the following amino acid substitutions—E176R, K238R, T447R, and E563R—were able to effectively enhance gene editing efficiency (as shown in Figure 1a, all four monoamino acid substitutions in E176R, K238R, T447R, and E563R yielded insertion / deletion rates higher than approximately 10% (this is the gene editing efficiency of the reference enzyme for CCR5-3) and higher than approximately 12% (this is the gene editing efficiency of the reference enzyme for RNF2-7), which are clearly superior to other single-amino acid substitution schemes). The remaining six mutants did not contribute to improving gene editing efficiency, and in fact, it decreased it. The coding sequences for the crRNA spacer sequences used here for CCR5-3 and RNF2-7 are shown in SEQ ID NO: 60 and 61, respectively.

[0321] Example 1-B Comparison of engineered Cas12i2 having multiple preferred amino acid substitutions simultaneously. Following the method described in Example 1, engineered Cas12i2 enzymes having two or more preferred amino acid substitutions in their amino acid sequences were expressed, and the combinations and their gene editing efficiencies are shown in Table 1 and Figure 1b. Point mutations in four efficiency-enhancing mutants obtained from the screening in Example 1-A—E176R, K238R, T447R, and E563R—were combined. As shown in Figure 1b and Table 1, by comparing the gene editing efficiencies of these mutants and wild-type Cas12i2 in three genomic sites, CCR5-3, CCR5-5, and RNF2-7, in 293T cells, it was found that mutants with further improved efficiency could be obtained after combining point mutations. In particular, the combination of the four types of mutations (E176R+K238R+T447R+E563R) yielded the most efficient combination of mutants. The coding sequences for the crRNA spacer sequences used for CCR5-3, RNF2-7, and CCR5-5 are shown in SEQ ID NO: 60, 61, and 62, respectively.

[0322] [Table 1]

[0323] Example 2: The amino acids involved in DNA double-strand uncoupling in the reference Cas12i2 enzyme are replaced with aromatic ring-containing amino acids, and the efficiency of each gene editing process is verified. Plasmid construction A variant of the Cas12i2 protein was produced by PCR-based site-directed mutagenesis. Specifically, the Cas12i2 protein DNA sequence was divided into two parts around the mutation site, two pairs of primers were designed to amplify these two partial DNA sequences, and the sequence requiring the mutation was simultaneously introduced into the primers. Finally, the two fragments were incorporated into a pCAG-2A-eGFP vector using the Gibson clone method. The amino acid substitution sites can be determined by analyzing the structural information of Cas12i2 using common protein structure visualization software (e.g., PyMol or Chimera). Structural information for Cas12i2 is referenced from PDB:6LTU, 6LTR, 6LU0, and 6LTP. The Cas12i2 effector protein is expressed in human 293T cells via the pCAG-2A-eGFP vector. The DNA encoding the Cas12i2 protein was inserted between XmaI and NheI. A vector expressing Cas12i2crRNA at 293T was constructed by binding an annealed oligonucleotide containing the target sequence to a pUC19-U6-i2-crRNA backbone digested with BasI. The nucleotide sequence encoding DR is SEQ ID NO:59.

[0324] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well cell dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas12i2 protein and 300 ng of a plasmid encoding crRNA were transfected into each 24-well cell dish. After 68 hours of transfection, HEK293T cells to be subjected to fluorescence-activated cell sorting (FACS) were digested with trypsin-EDTA (0.05%) (Gibco). Cell sorting was performed using MoFlo XDP (Beckman Coulter) with a GFP channel.

[0325] Targeted deep sequencing analysis for genome modification GFP-positive HEK293T cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of targets, the cell degradation sorbate was used directly as a template, and the target sites were directly amplified by barcoded PCR. PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. Reads with less than 0.05% of complete reads were discarded.

[0326] First, we selected the amino acids Q163 and N164 involved in DNA double-strand release in the Cas12i2 enzyme and performed point mutation tests on aromatic ring-attached amino acids (Y, F, W). From Figure 2 and Table 2, we compared the gene editing efficiency of these mutants with that of wild-type Cas12i2 at three genomic sites: CCR5-3, CCR5-5, and RNF2-7 in 293T cells. We discovered that five mutants—Q163W, Q163Y, Q163F, N164Y, and N164F—can effectively enhance gene editing efficiency (at least one genomic site). In particular, N164Y and N164F showed superior gene editing efficiency at three genomic sites. N164W did not improve gene editing efficiency compared to the reference enzyme.

[0327] [Table 2]

[0328] Example 3: The efficiency of gene editing was verified by substituting a positively charged amino acid located in the RuvC domain of the reference Cas12i2 enzyme that interacts with a single-stranded DNA substrate. Plasmid construction Variants of the Cas12i2 protein were produced by PCR-based site-directed mutagenesis. Specifically, the Cas12i2 protein DNA sequence was divided into two parts around the mutation site, two pairs of primers were designed to amplify the DNA sequences of these two parts, and the necessary mutation sequence was introduced into the primers simultaneously. Finally, the two fragments were incorporated into a pCAG-2A-eGFP vector using the Gibson clone method. The mutant combinations were constructed by multi-stage splitting of the Cas12i2 protein DNA and using PCR and Gibson cloning. The location of the mutants was determined by analyzing the structural information of Cas12i2 using common protein structure visualization software (e.g., PyMol, Chimera, etc.). Structural information for Cas12i2 is referenced from PDB IDs: 6LTU, 6LTR, 6LU0, and 6LTP. The ssDNA substrate shown in these Cas12i2 structures is only 5nt. To obtain information on the interaction between longer ssDNA and Cas12i2, the structure of Cas12i1 (PDB ID: 6W5C, 6W62 and 6W64, Zhang H. et al. Nature Structural & Molecular Biology 27, 1069-1076 (2020)) was homologously aligned with the structure of Cas12i2. The ssDNA substrate (9nt) in the Cas12i1 structure was placed in the RuvC catalytic pocket of Cas12i2, and amino acids within 9a were further searched through this model. The Cas12i2 effector protein is expressed in human 293T cells via a pCAG-2A-eGFP vector. The DNA encoding the Cas12i2 protein was inserted between XmaI and NheI. A vector expressing Cas12i2 crRNA in 293T cells was constructed by binding an annealed oligonucleotide containing the target sequence to a pUC19-U6-i2-crRNA backbone digested by BasI. The nucleotide sequence encoding DR is SEQ ID NO:59. The coding sequences for the crRNA spacer sequences for CCR5-3 and RNF2-7 are shown as SEQ ID NO:60 and 61, respectively.

[0329] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well cell dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas12i2 protein and 300 ng of a plasmid encoding crRNA were transfected into each 24-well cell dish. After 68 hours of transfection, HEK293T cells to be subjected to fluorescence-activated cell sorting (FACS) were digested with trypsin-EDTA (0.05%) (Gibco). Cell sorting was performed using MoFlo XDP (Beckman Coulter) with a GFP channel.

[0330] Targeted deep sequencing analysis for genome modification GFP-positive HEK293T cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of the targets, the cell degradation sorbate was used directly as a template, and the target sites were directly amplified by barcoded PCR. The PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. Reads less than 0.05% of complete reads were discarded.

[0331] In the reference Cas12i2 enzyme, an amino acid located in the RuvC domain that interacts with single-stranded DNA substrates was replaced with a positively charged amino acid. The gene editing efficiency of these mutants and wild-type Cas12i2 at the CCR5-3 and / or RNF2-7 genomic sites was compared in 293T cells. As shown in Figure 3a and Table 3, N391R, I926R, and G929R can effectively improve gene editing efficiency (at least one genomic site). Among them, I926R shows superior gene editing efficiency at two genomic sites. As shown in Figures 3b, 3c, and Table 3, many mutants can effectively improve gene editing efficiency (at least one genomic site). Of these, the ranking of single-amino acid substitution schemes with improved gene editing efficiency is D362R > E323R > Q425R > N925R > other mutants with improved efficiency.

[0332] As shown in Figures 3a, 3b, and 3c, we combined point mutations from four efficiency-enhancing mutants: E323R, D362R, Q425R, and I926R. By comparing the gene editing efficiency of these mutants and wild-type Cas12i2 in two genomic sites: CCR5-3 and RNF2-7, in 293T cells, we found that we could obtain mutants (at least one site) with improved gene editing efficiency after combining point mutations.

[0333] Point mutations or combinations from some of the mutants that could improve efficiency, obtained through screening as shown in Figures 3a, 3b, 3c, and 3d, were combined with flexible region mutations I926G and L439(L+GG). As shown in Figure 3e and Table 3, by comparing the gene editing efficiency of these mutants and wild-type Cas12i2 at the RNF2-7 site in 293T cells, it was found that more efficient mutants, such as N925R+I926G and I926R+L439(L+GG), could be obtained after combining point mutations. This suggests that further introduction of flexible region mutations can improve the gene editing efficiency of Cas12i mutants.

[0334] [Table 3-1]

[0335] [Table 3-2]

[0336] Example 4: The efficiency of gene editing was verified by substituting positively charged amino acids in the reference Cas12i2 enzyme that interact with the DNA-RNA double helix. Plasmid construction Variants of the Cas12i2 protein were produced by PCR-based site-directed mutagenesis. Specifically, the Cas12i2 protein DNA sequence was divided into two parts around the mutation site, two pairs of primers were designed to amplify these two partial DNA sequences, and the mutation-requiring sequence was introduced into the primers simultaneously. Finally, the two fragments were incorporated into the pCAG-2A-eGFP vector using the Gibson clone method. The mutant combinations were constructed by splitting the Cas12i2 protein DNA into multiple stages and using PCR and Gibson clones. The location of the mutants was obtained by analyzing the structural information of Cas12i2 using common protein structure visualization software (e.g., PyMol, Chimera, etc.). Structural information of Cas12i2 is referenced from PDB:6LTU,6LTR,6LU0,6LTP). The Cas12i2 effector protein is expressed in human 293T cells via the pCAG-2A-eGFP vector. The DNA encoding the Cas12i2 protein was inserted between XmaI and NheI. A vector expressing Cas12i2 crRNA at 293T was constructed by annealing an oligonucleotide containing the target sequence to a pUC19-U6-i2-crRNA backbone digested with BasI. The nucleotide sequence encoding DR is SEQ ID NO:59. The coding sequences for crRNA spacer sequences for CCR5-3 and RNF2-7 are shown as SEQ ID NO:60 and 61, respectively.

[0337] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well cell dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas12i2 protein and 300 ng of a plasmid encoding crRNA were transfected into each 24-well cell dish. After 68 hours of transfection, HEK293T cells to be subjected to fluorescence-activated cell sorting (FACS) were digested with trypsin-EDTA (0.05%) (Gibco). Cell sorting was performed using MoFlo XDP (Beckman Coulter) with a GFP channel.

[0338] Targeted deep sequencing analysis for genome modification GFP-positive HEK 293 T cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of targets, cell degradation sorbents were used directly as templates, and target sites were amplified directly by barcoded PCR. PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. Reads with less than 0.05% of complete reads were discarded.

[0339] Figure 4 and Table 4 summarize the comparison of gene editing efficiency at two genomic sites: CCR5-3 and RNF2-7, between Cas12i2 mutants and wild-type Cas12i2 in this embodiment using 293T cells. We found that seven mutants, G116R, E117R, T159R, S161R, E319R, E343R, and D958R, can effectively enhance gene editing efficiency (at least in one genomic site). Of these, D958R showed superior gene editing efficiency at two genomic sites.

[0340] [Table 4]

[0341] Example 5: A portion of the engineered amino acid mutations of Cas12i2, which were screened in Examples 1-4 to improve gene editing efficiency, are combined, and the efficiency of their gene editing is verified. Plasmid construction The mutant combinations are constructed by splitting the Cas12i2 protein DNA into multiple segments and using PCR and Gibson cloning. The mutant locations are determined by analyzing the Cas12i2 structural information using common protein visualization software (e.g., PyMol, Chimera, etc.). The Cas12i2 structural information is referenced from PDB (6LTU, 6LTR, 6LU0, 6LTP). The Cas12i2 effector protein is expressed in human 293T cells via a pCAG-2A-eGFP vector. The DNA encoding the Cas12i2 protein was inserted between XmaI and NheI. The vector expressing Cas12i2 crRNA in 293T cells was constructed by binding an annealed oligonucleotide containing the target sequence to a pUC19-U6-i2-crRNA backbone digested by BasI. The nucleotide sequence encoding DR is SEQ ID NO: 59.

[0342] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293 T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well cell dishes (Corning) and cultured for 16 hours until cell density reached 70%. Each 24-well cell dish was transfected with 600 ng of a plasmid encoding Cas12i2 protein and 300 ng of a plasmid encoding crRNA using Lipofectamine 3000 (Invitrogen). After approximately 68 hours of transfection, HEK 293 T cells were digested with trypsin-EDTA (0.05%) (Gibco) for fluorescence-activated cell sorting (FACS). Cell sorting was performed using MoFlo XDP (Beckman Coulter) with a GFP channel.

[0343] Targeted deep sequencing analysis for genome modification GFP-positive HEK293T cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of targets, the cell degradation sorbate was used directly as a template, and the target sites were directly amplified by barcoded PCR. PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. Reads with less than 0.05% of complete reads were discarded.

[0344] The amino acid mutations or combinations of amino acid mutations screened in Examples 1-4 were further combined: E176R+K238R+T447R+E563R, N164Y, E323R+D362R, I926R, E323R+D362R+I926R, E323R+D362R+I926G, E323R+D362R+I926G+L439(L+G), and E323R+D362R+I926G+L439(L+GG). As shown in Figure 5 and Table 5, by comparing the gene editing efficiency of these mutants and wild-type Cas12i acting on five genomic sites: CCR5-3, CCR5-5, CD34-8, CD34-9, and RNF2-14 in 293T cells, we found that improved mutants could be obtained more efficiently after combining different types of point mutations. At the same time, the mutant considered to be the most efficient (E176R+K238R+T447R+E563R+N164Y+E323R+D362R) was named CasXX. Of these, the coding sequences for the crRNA spacer sequences used for CD34-8, CD34-9, and RNF2-14 are shown in SEQ ID NO: 63, 64, and 65, respectively.

[0345] [Table 5]

[0346] Furthermore, the following mutation combination was constructed: E176R+K238R+T447R+E563R+N164Y+D958R; E176R+K238R+T447R+E563R+I926R+D958R; E176R+K238R+T447R+E563R+E323R+D362R+D958R; N164Y+I926R+D958R;N164Y+E323R+D362R+D958R; E176R+K238R+T447R+E563R+N164Y+I926R+D958R; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+D958R; E176R+K238R+T447R+E563R+N164Y+I926R+E323R+D362R+D958R; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+D958R; E176R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+L439(L+GG)+D958R; and E176R+K238R+T447R+E563R+N164Y+E323R+D362R+I926G+L439(L+G)+D958R. The gene editing efficiency can be detected by T77 endonuclease 1 (T7E1) assay and targeted deep sequencing.

[0347] Example 6: Verification of gene editing efficiency by comparing CasXX with conventional gene editing tools. Plasmid construction To compare the gene-editing activity of CasXX with other Cass proteins, the coding sequences of AsCas12a, BhCas12b v4, SpCas9, SaCas9, and SaCas9-KKH were synthesized by codon optimization (human). The Cas effector proteins were expressed in human 293T cells via a pCAG-2A-eGFP vector. The DNA encoding the Cas protein was inserted between XmaI and NheI. By annealing oligonucleotides containing the target sequences to a pUC19-U6-i2-crRNA backbone digested by BasI, vectors of sgRNA or crRNA expressing AsCas12a, BhCas12b v4, SpCas9, SaCas9, SaCas9-KKH, and Cas12i2 were constructed in 293T cells. The nucleotide sequence of the DR encoding CasXX is SEQ ID NO: 59.

[0348] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK 293 T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well cell dishes (Corning) and cultured for 16 hours until cell density reached 70%. Each 24-well cell dish was transfected with 600 ng of a plasmid encoding Cas12i 2 protein and 300 ng of a plasmid encoding crRNA using Lipofectamine 3000 (Invitrogen). After approximately 68 hours of transfection, HEK293 T cells to be subjected to fluorescence-activated cell sorting (FACS) were digested with trypsin-EDTA (0.05%) (Gibco). Cell sorting was performed using MoFlo XDP (Beckman Coulter) with a GFP channel.

[0349] Targeted deep sequencing analysis for genome modification GFP-positive HEK293T cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of targets, the cell degradation sorbate was used directly as a template, and the target sites were directly amplified by barcoded PCR. PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. Reads with less than 0.05% of complete reads were discarded.

[0350] We first tested the gene editing efficiency of CasXX at 62 personal genome sites containing different PAM sequences. The designed spacer sequence was 20 nucleotides. As shown in Figure 6a, CasXX demonstrated extremely potent gene editing capabilities, with an average gene editing efficiency exceeding 60% and an efficiency exceeding 50% at all tested sites. It also showed high gene editing efficiency for any NTTN PAM (N=A, T, G, C).

[0351] To further demonstrate the gene-editing capabilities of engineered CasXX, CasXX was compared with AsCas12a at the TTTN PAM site, and then with BhCas12b v4 at the TTN PAM site. As shown in Figure 6b, CasXX exhibits higher average gene-editing efficiency at target sites with two types of PAM.

[0352] Furthermore, CasXX was compared with SpCas9, SaCas9, and SaCas9-KKH at the same site. As shown in Figure 6c, CasXX showed a higher average gene editing efficiency at the test site.

[0353] To test the gene-editing activity of CasXX in vivo, mouse Hepa1-6 hepatocellular carcinoma cell lines were transfected with the CasXX-encoding pCAG-2A-eGFP vector and the crRNA-encoding pUC19-U6-i2-crRNA vector using liposome transfection. From these, 65 crRNAs corresponding to endogenous gene sites were designed. The designed spacer sequences were 20 nucleotides long. Insertion and deletion frequencies were obtained from PCR amplification sequencing, similar to methods used for analyzing exogenous gene editing.

[0354] As shown in Figure 6d, CasXX demonstrated powerful gene editing capabilities for 65 endogenous gene sites in the mouse Hepa1-6 cell line, with an average gene editing efficiency exceeding 60%.

[0355] Example 7: Gene editing using CasXX at genomic sites containing different PAMs. Plasmid construction A plasmid expressing CasXX was constructed by incorporating the DNA sequence encoding CasXX into the pCAG-2A-EGFP plasmid. The vector for crRNA expression in HEK293T was constructed by annealing oligonucleotides containing the target sequence to a BasI-digested pUC19-U6-crRNA backbone. Of these, 64 different crRNAs were designed to target 64 human endogenous sites. The 5' end of the endogenous target nucleic acids had different PAM5'-NNNN-3' (N=A, T, G, or C) as shown in Figure 7 to detect CasXX's recognition ability to different PAMs. Of these, the PAM sequences contained in these 64 sites cover all NNNN combinations: NTTN, NTAN, NTCN, NTGN, NATN, NAAN, NACN, NAGN, NCTN, NCAN, NCCN, NCGN, NGTN, NGAN, NGCN, NGGN. The nucleotide sequence encoding DR is SEQ ID NO: 59. The designed spacer sequence is 20 nucleotides.

[0356] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-cell petri dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas protein and a plasmid encoding crRNA were transfected into cells cultured in 24-well cell petri dishes with 300 ng. After 72 hours of transfection, cells were digested with trypsin-EDTA (0.05%) (Gibco), and then GFP fluorescence-activated cell sorting (FACS) was performed.

[0357] Targeted deep sequencing analysis for genome modification GFP-positive 293FT cells sorted by FACS were lysed in buffer L and incubated at 55°C for 3 hours, followed by incubation at 95°C for 10 minutes. dsDNA fragments containing target sites at different genomic locations were PCR amplified using corresponding primers. For deep sequencing of the target, the cell degradation sorbate was used directly as a template, and the target site was directly amplified by barcoded PCR. The PCR products were purified and compiled into several libraries for high-throughput sequencing. The frequency (%) of insertions and deletions was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. Reads with less than 0.05% of complete reads were discarded. The experimental results are shown in Figure 7.

[0358] As shown in Figure 7, CasXX showed efficient gene editing efficiency at target sites with PAMs (NTTN, NTAN, NTCN, NTGN, NATN, NAAN, NCTN, NCAN, and NGTN) at the 5' end, with an average gene editing efficiency exceeding 40% at these sites.

[0359] Example 8: In vitro cleavage of double-stranded DNA containing different PAMs using CasXX Construction of mutant expression plasmids The DNA sequences encoding wild-type Cas12i2 (SEQ ID NO:1) and CasXX (SEQ ID NO:8) were incorporated into a BPK2014 plasmid (which is chloramphenicol-resistant) to construct prokaryotic expression plasmids for Cas12i2 and CasXX proteins.

[0360] Protein purification The BPK2014 prokaryotic expression plasmid was used to transform E. coli strain BL21(λDE3) (TransGen Biotech). The transformed bacterial suspension was spread onto chloramphenicol-containing solid LB. Three clones were collected in 5 ml of liquid LB and incubated overnight. The bacteria were then transferred to 3 L of liquid LB and incubated until the OD600 reached 0.6–0.8. Subsequently, they were induced with IPTG (0.5 mM) at 16°C for 20 hours. Ultracentrifugation was performed to harvest bacteria expressing Cas12i, which were resuspended in degradation buffer (50 mM Tris-HCl, pH 7.5, 300 mM NaCl) and sonicated. After centrifugation, the Cas12i protein in the supernatant was purified using a Ni column. In short, after incubation with the supernatant, the Ni column was sequentially washed with lysis buffers containing 0 mM, 20 mM, and 50 mM imidazole. Subsequently, the Cas12i protein was eluted with 500 mM imidazole-supplemented lysis buffer. The collected samples were then loaded onto an ion-exchange column (CM Sepharose Fast Flow, GE). Wild-type Cas12i2 and CasXX proteins were eluted with storage buffer (20 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 10% glycerin, pH 7.5). The proteins were sterilized through filtration and stored at -80°C.

[0361] crRNA in vitro transcription The nucleotide sequence encoding DR is SEQ ID NO: 59. Oligonucleotides containing the T7 promoter sequence (named T7-F) and oligonucleotides containing crRNA and the T7 promoter complementary sequence (named T7-12i-crRNA-R) were synthesized and annealed in 1x NEBuffer™2 (NEB). The sequences of these oligonucleotides are shown in Table 6. Using the annealing product as a template, crRNA was produced using the HiScribe™ T7 Quick High Yield RNA Synthesis Kit (NEB). The transcribed crRNA was purified using the Monarch® RNA Cleanup Kit (NEB).

[0362] [Table 6]

[0363] In vitro enzymatic cleavage of the Cas12i protein To prepare linear dsDNA substrates for in vitro cleavage, targets containing the same prototype spacer region and different PAMs were first cloned into pUC19 (penicillin-resistant) treated with EcoR1 and HindIII. The target sequences containing the 5' PAM are shown in Table 7. Next, the target-supported pUC19 plasmid was linearized with SacI and purified using DNA Clean & Concentrator (Zymo Research). For the in vitro cleavage experiment, 400 nM Cas12i protein was first incubated with 2 μM crRNA at 37°C for 15 minutes. Then, in a 10 μl reaction system containing 1x NEBuffer™3.1 (NEB), Cas12i-crRNA RNP and 150 ng of linearized target DNA were reacted at 37°C for 40 minutes. After that, the reaction was terminated with 50 mM EDTA, and the RNA was digested with an RNase cocktail (Invitrogen) at 37°C for 15 minutes. Finally, the samples were treated with protease (NEB) at 37°C for 15 minutes. The reaction products were separated by electrophoresis on a 1.2% agarose gel.

[0364] As shown in Figure 8, the wild-type Cas12i2 protein exhibits partial cleavage efficiency only for double-stranded DNA containing 5'-NTTN-3'PAM, but has little cleavage activity for the remaining PAMs. However, the CasXX protein shows efficient cleavage efficiency for double-stranded DNA containing NTTN, NTAN, NTCN, NTGN, NATN, NAAN, NACN, NCTN, NCAN, NGTN, and NGAN PAMs. CasXX can almost completely cleave double-stranded DNA containing NTTN, NTAN, NTCN, NATN, NAAN, NACN, NCTN, NCAN, and NGTN PAMs. Therefore, the engineered Cas12i nucleases provided by this invention can be used for a wider range of gene editing or therapeutic applications.

[0365] [Table 7]

[0366] Example 9: Detection of off-target effects of CasXX using GUIDE-Seq Plasmid construction A plasmid expressing CasXX was constructed by incorporating the DNA sequence encoding CasXX into the pCAG-2A-EGFP plasmid. A vector expressing the Cas protein crRNA in 293T (human renal epithelial cell line) was constructed by annealing oligonucleotides containing the target sequence to a BasI-digested pUC19-U6-crRNA backbone. The crRNA contains a spacer sequence (SEQ ID NO: 77) that can target the endogenous sites of EMX1-7. The nucleotide sequence encoding DR is SEQ ID NO: 59.

[0367] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). The cells were seeded in 24-cell petri dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas protein, 300 ng of a plasmid encoding crRNA, and 10 pmol of annealed double-stranded DNA tag (see Table 8 for sequence) were transfected into cells cultured in 24-well petri dishes. After 72 hours of transfection, the cells were digested with trypsin-EDTA (0.05%) (Gibco), followed by GFP fluorescence-activated cell sorting (FACS). Cells successfully expressing the Cas enzyme were selected.

[0368] [Table 8]

[0369] Detection of off-target effects of CasXX across the entire genome. The genome for GFP-positive 293T cells sorted by FACS was extracted using the EZNA(R) MicroElute Genomic DNA Kit (Omega). The purified genome was quantified using Qubit. The genome was fragmented to approximately 500 bp using a Covaris S220 instrument according to the instrument's recommended program. Subsequently, a DNA library was constructed using the VAHTS Universal Pro DNA Library Prep Kit for Illumina (Vazyme). The library construction process was based on the reference (Tsai, SQ et al. “GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases,” Nat Biotechnol. 2015, 33(2):187-197, the contents of which are incorporated herein by reference). The library product was sequenced using high-throughput sequencing. The sequencing results were analyzed to search for potential off-target sites. A systematic analysis of the off-target effects of CasXX at the EMX1-7 sites was performed; see Figure 10 for specific analysis results. The sequence in the first row is the reference target sequence, and the sequences below represent the target sequence, the off-target sequence, and the number of reads in which the double-stranded DNA tag is enriched in the target sequence and the off-target sequence, respectively. As evidenced by the higher number of off-target events, the editing specificity of CasXX enzymes in mammalian cells is not ideal and needs further optimization.

[0370] Example 10: A new amino acid mutation was introduced based on the CasXX sequence (SEQ ID NO: 8), and the characteristics of improved specificity of this mutant were verified (mutant screening experiments were performed by selecting two target sites, namely the EMX1-7 site and the RNF2-1 site). Screening experiment preparation: Construction of mutant expression plasmids Based on CasXX (including the N164Y+E176R+K238R+E323R+D362R+T447R+E563R mutation based on SEQ ID NO:1), we further introduced novel amino acid point mutations to enhance the specificity of CasXX gene editing and reduce off-target effects. Based on the CasXX sequence (SEQ ID NO:8), we performed PCR on the CasXX DNA sequence using primers containing mutant bases, and incorporated the purified PCR product into the pCAG-2A-EGFP plasmid using the NEBuilder(R) HiFi DNAAssembly Master Mix (NEB) kit, thereby constructing a CasXX plasmid expressing the introduced related point mutations. We obtained 26 variants, including single amino acid mutations, based on the CasXX sequence, and named them CasXX-HF-1 to HF-26, respectively. The specific mutation methods are shown in Table 9. In this case, CasXX-HF-26 restored the N164Y point mutation in CasXX to the original amino acid N of wild-type Cas12i2, i.e., Y164N (i.e., deleted the N164Y mutation).

[0371] [Table 9]

[0372] Construct crRNA expression plasmids targeting EMX1-7 sites. A vector expressing Cas protein crRNA at 293T was constructed by ligating an annealed oligonucleotide containing the target sequence to a pUC19-U6-crRNA backbone digested with BasI(NEB) using T4ligase(NEB). The crRNA was designed to target the EMX1-7 sites. The specific sequence is shown in Table 10. The nucleotide sequence encoding DR is SEQ ID NO: 59.

[0373] [Table 10]

[0374] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-cell petri dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas protein and 300 ng of a plasmid encoding crRNA were transfected into cells cultured in 24-well petri dishes. After 72 hours of transfection, cells were digested with trypsin-EDTA (0.05%) (Gibco), and then GFP fluorescence-activated cell sorting (FACS) was performed.

[0375] Detection of Cas protein editing efficiency at EMX1-7 target and off-target sites using T7E1 enzyme digestion. GFP-positive HEK293T cells sorted by FACS were degraded with 40 μL of buffer (bimake), incubated at 55°C for 3 hours, and then incubated at 95°C for 10 minutes. dsDNA fragments containing target or off-target sites at different genomic locations were PCR-amplified using primers for the EMX1-7 target site, EMX1-7 off-target site 1, EMX1-7 off-target site 2, and EMX1-7 off-target site 3 (Figure 10). See Table 11 for the sequences of the target or off-target sites. Subsequently, 10 μL of the PCR product was re-annealed to form heterologous double-stranded dsDNA. Next, the mixture was incubated in 1 / 10 volume of NE Buffer. TM2.1 and 0.2 μL of T7endonuclease I (NEB) were applied at 37°C for 50 minutes. The digestion products were analyzed by ~2.5% agarose gel electrophoresis. The percentage of insertions and deletions (Indel,%) was calculated based on the grayscale values ​​of the bands. The percentage of insertions and deletions at target or off-target sites for each Cas12i mutant is shown in Table 12, corresponding to Figure 11. Experimental results showed that single-point mutants based on CasXX sequences (SEQ ID NO: 8), namely R857A, N861A, K807A, N848A, R715A, R719A, K394A, H357A, and K844A (corresponding to engineered Cas12i nucleases with SEQ ID NO: 14-22, see Table 9), were able to effectively reduce the insertion / deletion rate at off-target sites EMX1-7-OT-1, EMX1-7-OT-2, and EMX1-7-OT-3, demonstrating higher specificity. In the experimental results, improved editing specificity was defined as a reduction in the insertion / deletion rate at off-target sites and a fundamental invariance (or increase) in editing efficiency at target sites.

[0376] [Table 11]

[0377] [Table 12-1]

[0378] [Table 12-2]

[0379] Construct a crRNA expression plasmid that targets the RNF2-1 site. A vector expressing Cas protein crRNA in HEK293T was constructed by ligating an annealed oligonucleotide containing the target sequence to a pUC19-U6-crRNA backbone digested with BasI(NEB) using T4ligase(NEB). All final crRNAs target the same RNF2-1 site, but have different spacer sequences. The specific sequences encoding these spacers are shown in Table 13. Of these, RNF2-1-FM indicates that the spacer sequence in the crRNA is a perfect match with the RNF2-1 site. RNF2-1-Mis-1 / 2 indicates that the spacer sequence in the crRNA does not match with the RNF2-1 site at the 1st and 2nd base positions, but matches at the remaining positions. RNF2-1-Mis-5 / 6 indicates that the spacer sequence in the crRNA does not match with the RNF2-1 site at the 5th and 6th base positions, but matches at the remaining positions. RNF2-1-Mis-17 / 1 indicates that the spacer sequence in the crRNA does not match the RNF2-1 site at base positions 17 and 18, but matches at the remaining positions. RNF2-1-Mis-19 / 20 indicates that the spacer sequence in the crRNA does not match the RNF2-1 site at base positions 19 and 20, but matches at the remaining positions. The purpose of setting these base mismatches is to simulate off-target effects. The nucleotide sequence encoding DR is SEQ ID NO: 59.

[0380] [Table 13]

[0381] Cell culture, transfection, and fluorescence-activated cell sorting (FACS) HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). The cells were seeded in 24-cell petri dishes (Corning) and cultured for 16 hours until the cell density reached 70%. Using Lipofectamine 3000 (Invitrogen), 600 ng of a plasmid encoding Cas protein and 300 ng of a plasmid encoding crRNA were transfected into the cells cultured in the 24-well petri dishes. After 72 hours of transfection, the cells were digested with trypsin-EDTA (0.05%) (Gibco), and then GFP fluorescence-activated cell sorting (FACS) was performed. Cells that successfully expressed the Cas enzyme were selected.

[0382] Detection of editing efficiency of Cas12i mutants in RNF2-1 and off-target sites using T7E1 enzyme digestion. GFP-positive HEK293T cells sorted by FACS were degraded with 40 μL of buffer (bimake), incubated at 55°C for 3 hours, and then incubated at 95°C for 10 minutes. RNF2-1 primers were used to PCR amplify the RNF2-1 site dsDNA fragment. Subsequently, 10 μL of the PCR product was re-annealed to form heterologous double-stranded dsDNA. The mixture was then incubated in 1 / 10 volume of NE Buffer. TM2.1 and 0.2 μL of T7endonuclease I (NEB) were applied at 37°C for 50 minutes. The digestion products were analyzed by 2.5% agarose gel electrophoresis. The percentage of insertions and deletions (Indel,%) was calculated based on the grayscale values ​​of the bands. Table 14 shows the percentage of insertions and deletions at the target site for each Cas12i mutant using different crRNAs, corresponding to Figure 12. Experimental results showed that single-point mutants based on CasXX sequences, namely R857A, N861A, K807A, N848A, R715A, R719A, K394A, H357A, and K844A (corresponding to engineered Cas12i nucleases with SEQ ID NOs: 14-22; see Table 9), could effectively reduce the insertion / deletion rate even when using crRNA-RNF2-1-Mis-1 / 2, crRNA-RNF2-1-Mis-5 / 6, crRNA-RNF2-1-Mis-17 / 18, and crRNA-RNF2-1-Mis-19 / 20, which mimic off-target sites. In the experimental results, improved editing specificity was defined as a reduction in the insertion / deletion rate at off-target sites and a fundamental invariance (or improvement) in editing efficiency at target sites.

[0383] [Table 14-1]

[0384] [Table 14-2]

[0385] Example 11: Mutants with novel amino acid mutations introduced based on the CasXX sequence For mutants in which new amino acid mutations were introduced into the CasXX sequence, those with improved specificity were screened using a fluorescence reporting system. The difference between this example and Example 10 is that the analytical method for indicating editing efficiency was changed from agarose gel electrophoresis to fluorescence.

[0386] Construct a crRNA expression plasmid that targets the mCherry gene. A vector expressing Cas protein crRNA in HEK293T was constructed by ligating an annealed oligonucleotide containing the target sequence to a pUC19-U6-crRNA backbone digested with BasI(NEB) using T4ligase(NEB). All final crRNAs target the same mCherry site, but have different spacer sequences. Table 15 shows the specific seque...

Claims

1. An engineered Cas12i nuclease comprising one or more mutations in a reference Cas12i nuclease, wherein the reference Cas12i nuclease is a wild-type Cas12i2 nuclease with amino acid sequence SEQ ID NO:1, and the mutations are: (1) Substitution of one or more amino acids that interact with PAM in the reference Cas12i nuclease with positively charged amino acids (wherein this substitution of one or more amino acids that interact with PAM in the reference Cas12i nuclease with positively charged amino acids is one or more of the E176R, K238R, T447R, and E563R substitutions in SEQ ID NO: 1); (2) Substituting one or more amino acids involved in the release of the DNA double helix in the reference Cas12i nuclease with aromatic ring amino acids (where, the substitution of one or more amino acids involved in the release of the DNA double helix in the reference Cas12i nuclease with aromatic ring amino acids is one or more of the Q163F, Q163Y, Q163W, N164Y, and N164F substitutions in SEQ ID NO: 1); (3) Substitution of one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with the single-stranded DNA substrate with a positively charged amino acid (wherein this substitution of one or more amino acids located in the RuvC domain of the reference Cas12i nuclease that interact with the single-stranded DNA substrate with a positively charged amino acid is one or more of the E323R, D362R, N391R, Q424R, Q425R, N925R, I926R, and G929R substitutions in SEQ ID NO: 1); (4) Substituting one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with positively charged amino acids (wherein the substitution of one or more amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with positively charged amino acids is one or more of the G116R, E117R, T159R, S161R, E319R, E343R and D958R substitutions in SEQ ID NO: 1); and (5) Substituting one or more polar or positively charged amino acids that interact with the DNA-RNA double helix in the reference Cas12i nuclease with hydrophobic amino acids (wherein this substitution is one or more of the H357A, K394A, R715A, R719A, K807A, K844A, N848A, R857A and N861A substitutions in SEQ ID NO: 1); Engineered Cas12i nuclease.

2. The engineered Cas12i nuclease according to claim 1, wherein the engineered Cas12i nuclease comprises any of the following mutations or combinations of mutations: (1) E563R; (2) E176R, T447R, E176R and E563R; (3) K238R and E563R; (4) E176R, K238R and T447R; (5) E176R, K238R and E563R; (6) E176R, T447R and E563R; and (7) E176R, K238R, T447R and E563R.

3. The engineered Cas12i nucleases are (1) E323R; (2) D362R; (3) Q425R; (4) N925R; (5) I926R; (6) E323R and D362R; (7) E323R and Q425R; (8) E323R and I926R; (9) Q425R and I926R; (10) D362R and I926R; (11) N925R and I926R; (12) E3 An engineered Cas12i nuclease according to claim 1, comprising any of the mutations or combinations of mutations of 23R, D362R and Q425R; (13) E323R, D362R and I926R; (14) E323R, Q425R and I926R; (15) D362R, N925R and I926R; and (16) E323R, D362R, Q425R and I926R.

4. The invention further comprises one or more flexible region mutations, the mutations increasing the flexibility of the flexible region in the reference Cas12i nuclease, the flexible region being selected from amino acid residues 439-443 or amino acid residues 925-929; The aforementioned flexible region mutation is the mutation at L439 and / or I926; Of these, the aforementioned amino acid position is defined as the amino acid position in SEQ ID NO: 1, The one or more flexible region mutations involve substituting the flexible region amino acid with G and / or inserting one or two Gs after it; The one or more flexible region mutations include I926G, L439(L+G), or L439(L+GG); The engineered Cas12i nuclease according to claim 1, wherein the amino acid position number is defined as the amino acid position in SEQ ID NO:

1.

5. An engineered Cas12i nuclease, comprising one or more mutations in a reference Cas12i nuclease, wherein the reference Cas12i nuclease is a wild-type Cas12i2 nuclease with amino acid sequence SEQ ID NO: 1, and the mutations are: (1) E563R; (2) E176R and T447R; (3) E176R and E563R; (4) K238R and E563R; (5) E176R, K238R and T447R; (6) E176R, T447R and E563R; (7) E176R, K238R and E563R; (8) E176R, K238R, T447R and E563R; (9) N164Y; (1 0) N164F; (11) E323R; (12) D362R; (13) Q425R; (14) N925R; (15) I926R; (16) D958R; (17) E323R and D362R; (18) E323R and Q425R; (19) E323R and I926R; (20) Q425R and I926R; (21) D362R and I926R; (22) N925R and I 926R; (23) E323R, D362R and Q425R; (24) E323R, D362R and I926R; (25) E323R, Q425R and I926R; (26) D362R, N925R and I926R; (27) E323R, D362R, Q425R and I926R; (28) D362R and I926G; (29) N925R and I926G; ( 30) D362R, N925R and I926G; (31) I926R and L439(L+G); (32) I926R and L439(L+GG); (33) E323R, D362R and I926G; (34) R719A and K844A; and (35) an engineered Cas12i nuclease comprising one or more of R857A and K844A.

6. An engineered Cas12i nuclease; comprising one or more mutations in a reference Cas12i nuclease, wherein the amino acid sequence of the reference Cas12i nuclease is SEQ ID NO:1 wild-type Cas12i2 nuclease, with the mutations being: (1) E176R, K238R, T447R, E563R and N164Y; (2) E176R, K238R, T447R, E563R and I926R; (3) N164Y, E323R and D362R; (4) E176R, K238R, T447R, E563R, E323R and D362R; (5) N164Y and I926R; (6) E176R, K238R, T447R, E563R, N164Y and I926R; (7) E176 R, K238R, T447R, E563R, N164Y, E323R and D362R; (8) E176R, K238R, T447R, E563R, N164Y, I926R, E323R and D362R; (9) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and I926G; (10) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G and L439 (L+GG); (11) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G and L439 (L+G); (12) E176R, K238R, T447R, E563R, N164Y and D958R; (13) E176R, K238R, T447R, E563R, I926R and D958R; (14) E176R, K238R, T447R, E563R, E323R, D362R and D958R; (15) N164Y, I926R and D958R; (16) N164Y, E323R, D362R and D 958R; (17) E176R, K238R, T447R, E563R, N164Y, I926R and D958R; (18) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and D958R; (19) E176R, K238R, T447R, E563R, N164Y, I926R, E323R, D362R and D958R; (20) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G and D958R;(21) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, L439 (L+GG) and D958R; (22) E1 76R, K238R, T447R, E563R, N164Y, E323R, D362R, I926G, L439 (L+G) and D958R; (23) E176R, K23 8R, T447R, E563R, N164Y, E323R, D362R and R857A; (24) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and N861A; (25) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and K807A; (26) E1 76R, K238R, T447R, E563R, N164Y, E323R, D362R and N848A; (27) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and R715A; (28) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and R719A (29) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and K394A; (30) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and H357A; (31) E176R, K238R, T447R, E563R, N164Y, E323R, D362R and K844A; (32) E176R, K238R, T447R, E563R, N164Y, E323R, D362R, R719A and K844A; or (33) Engineered Cas12i nucleases containing any of the mutations of E176R, K238R, T447R, E563R, N164Y, E323R, D362R, R857A and K844A.

7. The engineered Cas12i nuclease according to claim 1, comprising an engineered Cas12i nuclease having an amino acid sequence shown in any of SEQ ID NOs: 2-12 and 14-24.

8. An engineered Cas12i effector protein comprising the engineered Cas12i nuclease or an enzyme-inactive variant thereof as described in claim 1; The engineered Cas12i nuclease enzyme-inactive mutants are enzyme-inactive mutants comprising one or more mutations of D599A, E833A, S883A, H884A, R900A, and D1019A; among which the amino acid position is defined as the amino acid position in SEQ ID NO: 1, the engineered Cas12i effector protein.

9. The engineered Cas12i effector protein according to claim 8, further comprising a functional domain that fuses with the engineered Cas12i nuclease or an enzyme-inactive variant thereof.

10. The engineered Cas12i effector protein according to claim 9, wherein the functional domain is one or more selected from the group consisting of a translation initiation domain, a transcription repression domain, a transactivation domain, an epigenetic modification domain, a nucleic acid base editing domain, a reverse transcriptase domain, a reporter molecule domain, and a nuclease domain.

11. The engineered Cas12i effector protein comprises a first polypeptide comprising the N-terminal portion of the engineered Cas12i nuclease or an enzyme-inactive variant thereof, and a second polypeptide comprising the C-terminal portion of the engineered Cas12i nuclease or an enzyme-inactive variant thereof, wherein the first polypeptide and the second polypeptide can associate with each other in the presence of a guide RNA containing a guide sequence to form a clustered, regularly spaced, short palindromic repeat sequence (CRISPR) complex that specifically binds to a target nucleic acid, the target nucleic acid comprising a target sequence complementary to the guide sequence; The first polypeptide comprises N-terminal partial amino acid residues 1 to X of the engineered Cas12i nuclease described in any of claims 1 to 7, and the second polypeptide comprises amino acid residues X+1 to the C-terminus of the engineered Cas12i nuclease described in any of claims 1 to 7; The first polypeptide and the second polypeptide each contain a dimerization domain; The engineered Cas12i effector protein according to claim 8, wherein the dimer domain of the first polypeptide and the dimer domain of the second polypeptide associate with each other in the presence of an inducer.

12. An engineered CRISPR-Cas12i system, (a) an engineered Cas12i nuclease according to claim 1, or an engineered Cas12i effector protein comprising the said Cas12i nuclease; and (b) A guide RNA comprising a guide sequence complementary to the target sequence, or one or more nucleic acids encoding the guide RNA, Among these, the engineered Cas12i nuclease or engineered Cas12i effector protein and the guide RNA can form a CRISPR complex, the CRISPR complex specifically binds to a target nucleic acid containing the target sequence, and induces modification of the target nucleic acid, thereby creating an engineered CRISPR-Cas12i system.

13. (a) Contacting the sample with the engineered CRISPR-Cas12i system and tagged detection nucleic acid according to claim 12, wherein the detection nucleic acid is single-stranded and does not hybridize with the guide sequence of the guide RNA; and (b) A method for detecting a target nucleic acid in a sample, comprising measuring a detectable signal generated by the engineered Cas12i nuclease or engineered Cas12i effector protein cleaving the tagged detection nucleic acid, thereby detecting the target nucleic acid.

14. This includes contacting the target nucleic acid with the engineered CRISPR-Cas12i system described in claim 12; The above method is performed in vitro or ex vivo; The aforementioned target nucleic acid is present in cells, and the method involves modifying a target nucleic acid containing a target sequence.

15. Use in the manufacture of a drug for treating a disease or disorder related to a target nucleic acid in the cells of an organism, as described in claim 12, wherein the disease or disorder is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, ophthalmological disease, bacterial infection, and viral infection.

16. A method for modifying a target nucleic acid, including a target sequence, which is performed in vitro or ex vivo, comprising contacting the target nucleic acid with the engineered CRISPR-Cas12i system described in claim 12.

17. A composition or kit comprising the engineered Cas12i nuclease described in claim 1, and an engineered Cas12i effector protein containing the Cas12i nuclease.

Citation Information

Patent Citations

  • Engineered Cas effector protein and using method thereof

    CN112195164A