Cas12 protein and application thereof

CN120958124APending Publication Date: 2025-11-14GUANGZHOU REFORGENE MEDICINE CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202480023951.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2024-09-19
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

The prior art has not yet found effective new Cas12 protein and CRISPR-Cas12 gene editing systems in the field of CRISPR gene editing, limiting the flexibility and efficiency of gene editing.

Method used

A Cas12 protein is provided whose amino acid sequence has 50% to 100% sequence identity to a known sequence and is able to form a complex with a guide polynucleotide, specifically bind and cleave the target nucleic acid.

Benefits of technology

By providing Cas12 protein with high sequence identity, efficient binding to guide polynucleotides and precise cleavage of target nucleic acids are achieved, improving the flexibility and efficiency of CRISPR gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000243_0000
    Figure 00000243_0000
  • Figure 00000244_0000
    Figure 00000244_0000
  • Figure 00000245_0000
    Figure 00000245_0000
Patent Text Reader

Abstract

Relates to a Cas12 protein, a guide polynucleotide, a Cas12 inactivated variant, a fusion protein or conjugate comprising the Cas12 protein, an isolated nucleic acid, a CRISPR-Cas12 system, a carrier system, a delivery system, a cell, a pharmaceutical composition and a kit, and applications thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Cas12 protein and its applications

[0001] This application claims priority to Chinese Patent Application No. 2023112143306 filed on September 19, 2023, and Chinese Patent Application No. 2024103885922 filed on April 1, 2024. This application incorporates the entirety of the aforementioned Chinese patent applications. Technical Field

[0002] The present disclosure relates to the field of CRISPR gene editing, and specifically to a Cas12 protein and its applications. Background Art

[0003] The CRISPR-Cas system is an adaptive immune defense developed by bacteria and archaea over a long period of evolution, used to combat invading viruses and foreign DNA. The clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated protein system (CRISPR-Cas system) allows for direct modification of genetic sequences within cells, providing a rapid and effective method.

[0004] Many researchers in this field are working to find new Cas12 proteins and CRISPR-Cas12 gene editing systems.

[0005] Summary of the Invention

[0006] The present disclosure provides Cas12 proteins and their applications.

[0007] On the one hand, a technical solution provided by the present disclosure is: a Cas12 protein, wherein the Cas12 protein is CLUSTER1 protein, CLUSTER2 protein, CLUSTER3 protein, CLUSTER4 protein, CLUSTER5 protein, CLUSTER6 protein, CLUSTER7 protein, CLUSTER8 protein, CLUSTER9 protein, CLUSTER10 protein, CLUSTER11 protein, CLUSTER12 protein or CLUSTER13 protein.

[0008] In another aspect, the present disclosure provides a Cas12 protein, the amino acid sequence of which comprises or is an amino acid sequence having at least 50% sequence identity with any one of SEQ ID NOs: 1-53, 696, and 728.

[0009] In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to any one of SEQ ID NOs: 1-53, 696, 728.

[0010] In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% sequence identity compared to any one of SEQ ID NOs: 1-53, 696, 728.

[0011] In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 85% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 90% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 95% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 97% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 98% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99.5% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99.7% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In specific embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99.8% identity compared to any one of SEQ ID NOs: 1-53, 696, 728. In a specific embodiment of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is 100% identical to any one of SEQ ID NOs: 1-53, 696, and 728.

[0012] In a specific embodiment of the present disclosure, the Cas12 protein retains the function of the protein shown in any one of SEQ ID NOs: 1-53, 696, and 728.

[0013] In the specific embodiments of the present disclosure, the Cas12 protein can form a complex with the guide polynucleotide. In the specific embodiments of the present disclosure, the Cas12 protein can specifically bind to the target nucleic acid with the guide polynucleotide.

[0014] In the specific embodiments of the present disclosure, the Cas12 protein can form a complex with the guidance polynucleotide, and the complex can specifically bind to the target nucleic acid. In the specific embodiments of the present disclosure, the Cas12 protein can form a complex with the guidance polynucleotide, and the complex can specifically bind to the target DNA.

[0015] In the specific embodiments of the present disclosure, the Cas12 protein can specifically bind to the guide polynucleotide and cut the target nucleic acid. In the specific embodiments of the present disclosure, the Cas12 protein can specifically bind to the guide polynucleotide and cut the target DNA. In the specific embodiments of the present disclosure, the Cas12 protein can form a complex with the guide polynucleotide, and the complex can specifically bind to and cut the target nucleic acid. In the specific embodiments of the present disclosure, the Cas12 protein can form a complex with the guide polynucleotide, and the complex can specifically bind to and cut the target DNA.

[0016] In the present disclosure, retaining the function of the protein as shown in any one of SEQ ID NOs: 1-53, 696, and 728 refers to retaining the ability to form a complex with a guide polynucleotide, retaining the ability to bind to a target nucleic acid complementary to the guide sequence of the guide polynucleotide, retaining the ability to target and cleave the target nucleic acid with the guide polynucleotide, and / or retaining the ability to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.

[0017] In a specific embodiment of the present disclosure, the function of the protein represented by any one of SEQ ID NOs: 1-53, 696, and 728 is retained by the protein to form a complex with the guide polynucleotide.

[0018] In a specific embodiment of the present disclosure, the function of the protein shown in any one of SEQ ID NOs: 1-53, 696, and 728 is retained by retaining the ability to bind to a target nucleic acid complementary to the guide sequence of the guide polynucleotide.

[0019] In a specific embodiment of the present disclosure, the function of retaining the protein as shown in any one of SEQ ID NOs: 1-53, 696, and 728 is to retain the ability to guide the polynucleotide to target and cleave the target nucleic acid.

[0020] In a specific embodiment of the present disclosure, the function of the protein represented by any one of SEQ ID NOs: 1-53, 696, and 728 is retained to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.

[0021] In a preferred embodiment of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence as shown in any one of SEQ ID NOs: 1-53, 696, and 728.

[0022] In a specific embodiment of the present disclosure, the PAM sequence (5'→3') recognizable by the Cas12 protein is selected from any one or more of the following:

[0023] A, C, T, G,

[0024] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,

[0025] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,

[0026] <h2 style=";text-align:left;direction:ltr">NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA,CANC,ACGA,CCAC,CCGG,CTNG,CNGN,GGTA,NGNC,GTTT,CTAA,TNCT,CTGN,NGAC,TGTA,TANN,GCNT,GCTC,CNCG,AAAN,CCNT,GANA,CACA,CTNA,ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG,GTGG,GNGN,ACCA,NTAA,ACTN,NCTG,NCTA,TTTT,GCNG,NTAG,CAAA,GGNA,CNTN,TTAG,TCTG,NCTN,TATG,GCGT,TANT,GGGT,NACN,ACTG,CCNG,GNNT<h2 style=";text-align:left;direction:ltr">CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN<h2 style=";text-align:left;direction:ltr">GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA,AAGA,CTTT,AAAC,AGGN,ACNT,NTGT,CTTN,ATCA,NACT,NNAG,NGTN,NAAC,TGCG,GGNT,ATAN,TTGC,ANCN,CCCC,ANGA,NGCG,TCTC,CTCG,ATNA,AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,

[0027] The N is A, T, C or G.

[0028] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-T-3'.

[0029] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-G-3'.

[0030] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-A-3'.

[0031] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-C-3'.

[0032] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TA-3'.

[0033] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TC-3'.

[0034] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TG-3'.

[0035] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TT-3'.

[0036] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TN-3'.

[0037] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TTN-3'.

[0038] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TTT-3'.

[0039] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TTG-3'.

[0040] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TTC-3'.

[0041] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-TTA-3'.

[0042] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-WTN-3'.

[0043] In a specific embodiment of the present disclosure, the Cas12 protein can recognize a PAM with a sequence of 5'-ATN-3'.

[0044] The N can be any one of A, T, C and G. The W can be A or T.

[0045] In some embodiments of the present disclosure, the Cas12 protein is an inactivated variant of Cas12. In some embodiments of the present disclosure, the Cas12 protein is a variant in which nuclease activity is inactivated. In some embodiments of the present disclosure, the Cas12 protein is a dead Cas12 inactivated variant or a nickase Cas12 inactivated variant. Alternatively, the RuvC domain of the Cas12 protein is inactivated.

[0046] In some embodiments of the present disclosure, the Cas12 protein is selected from the active fragments constituting any one of the Cas12 proteins of the present disclosure.

[0047] In some embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to any one of SEQ ID NOs: 46, 696, 52, 728.

[0048] Optionally, the Cas12 protein can form a complex with the guide polynucleotide; further, the complex can specifically bind to the target nucleic acid; further, the complex can cut the target nucleic acid, modify the target nucleic acid and / or regulate the expression of the target nucleic acid.

[0049] Optionally, the Cas12 protein can form a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence, and the backbone sequence can interact with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the direct repeat sequence comprises or is a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to any one of SEQ ID NO: 704, 529, and 534.

[0050] Optionally, the backbone sequence does not include a tracrRNA sequence.

[0051] Alternatively, the Cas12 protein can recognize a PAM sequence of 5'-TTN-3' and / or 5'-TTNC-3'. The N can be any one of A, T, C and G.

[0052] In some embodiments of the present disclosure, a Cas12 protein is provided, the amino acid sequence of the Cas12 protein comprising or being an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 46.

[0053] Optionally, the Cas12 protein can form a complex with the guide polynucleotide; further, the complex can specifically bind to the target nucleic acid; further, the complex can cleave the target nucleic acid, modify the target nucleic acid and / or regulate the expression of the target nucleic acid;

[0054] Optionally, the Cas12 protein can form a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence, and the backbone sequence can interact with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96% or at least 97% sequence identity with the sequence shown in SEQ ID NO: 704.

[0055] In some embodiments of the present disclosure, a Cas12 protein is provided, the amino acid sequence of the Cas12 protein comprising or being an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 696.

[0056] Optionally, the Cas12 protein can form a complex with the guide polynucleotide; further, the complex can specifically bind to the target nucleic acid; further, the complex can cut the target nucleic acid, modify the target nucleic acid and / or regulate the expression of the target nucleic acid.

[0057] Optionally, the Cas12 protein can form a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence, and the backbone sequence can interact with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96% or at least 97% sequence identity with the sequence shown in SEQ ID NO: 704.

[0058] In some embodiments of the present disclosure, a Cas12 protein is provided, the amino acid sequence of the Cas12 protein comprising or being an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 52.

[0059] Optionally, the Cas12 protein can form a complex with the guide polynucleotide; further, the complex can specifically bind to the target nucleic acid; further, the complex can cut the target nucleic acid, modify the target nucleic acid and / or regulate the expression of the target nucleic acid.

[0060] Optionally, the Cas12 protein can form a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence, and the backbone sequence can interact with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96% or at least 97% sequence identity with the sequence shown in SEQ ID NO: 534.

[0061] In some embodiments of the present disclosure, a Cas12 protein is provided, the amino acid sequence of the Cas12 protein comprising or being an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 728.

[0062] Optionally, the Cas12 protein can form a complex with the guide polynucleotide; further, the complex can specifically bind to the target nucleic acid; further, the complex can cut the target nucleic acid, modify the target nucleic acid and / or regulate the expression of the target nucleic acid.

[0063] Optionally, the Cas12 protein can form a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence, and the backbone sequence can interact with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96% or at least 97% sequence identity with the sequence shown in SEQ ID NO: 534.

[0064] In some embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 46.

[0065] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid and a direct repeat sequence.

[0066] Optionally, the direct repeat sequence comprises or is a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to SEQ ID NO: 704.

[0067] Optionally, the complex binds to the target nucleic acid under the guidance of the guide sequence.

[0068] Optionally, the Cas12 protein can recognize a PAM with a sequence of 5'-TTN-3';

[0069] The N may be any one of A, T, C and G.

[0070] In some embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 696.

[0071] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid and a direct repeat sequence.

[0072] Optionally, the direct repeat sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to SEQ ID NO: 704.

[0073] Optionally, the complex binds to the target nucleic acid under the guidance of the guide sequence;

[0074] Optionally, the Cas12 protein can recognize a PAM with a sequence of 5'-TTN-3'.

[0075] In some embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 52.

[0076] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid and a direct repeat sequence.

[0077] Optionally, the direct repeat sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to SEQ ID NO: 534.

[0078] Optionally, the complex binds to the target nucleic acid under the guidance of the guide sequence;

[0079] Optionally, the Cas12 protein can recognize a PAM with a sequence of 5'-TTN-3'.

[0080] In some embodiments of the present disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to SEQ ID NO: 728.

[0081] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid and a direct repeat sequence.

[0082] Optionally, the direct repeat sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to SEQ ID NO: 534.

[0083] Optionally, the complex binds to the target nucleic acid under the guidance of the guide sequence.

[0084] Optionally, the Cas12 protein can recognize a PAM with a sequence of 5'-TTN-3';

[0085] The N may be any one of A, T, C and G.

[0086] Alternatively, the reverse complement is a partial complement or a complete complement. In some embodiments, the guide sequence hybridizes to the target nucleic acid.

[0087] In some embodiments, the Cas12 protein is a mutant of the Cas protein shown in any one of SEQ ID NOs: 1-53, 696, and 728.

[0088] In some embodiments, the Cas12 protein is an inactivated variant of the Cas protein shown in any one of SEQ ID NOs: 1-53, 696, and 728.

[0089] In some embodiments, the Cas12 protein provided herein comprises one, two or more mutations, such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof, compared to the Cas12 protein shown in any one of SEQ ID NOs: 1-53, 696, and 728. In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 4, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 128, 129 or 130 amino acid changes (e.g., insertions, deletions or substitutions) in the guide polynucleotide, but retain the ability to bind to a target nucleic acid molecule that is complementary to the guide sequence of the guide polynucleotide and / or retain the ability to process an RNA transcript comprising the guide sequence into a guide polynucleotide molecule.In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191,

[0090] In some embodiments of the present disclosure, the Cas12 protein has a mutation at any amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. Alternatively, the mutation is to any other natural amino acid residue. Further alternatively, the mutation is to residue R, H, K, or A. In some embodiments, the mutation is to residue R. In some embodiments, the mutation is to residue A. In some embodiments, the mutation is to residue H. In some embodiments, the mutation is to residue K.

[0091] In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 1-41, 42-195, 196-290, 291-358, 359-479, 480-636, 637-689, 690-846, 847-884, 885-959, 960-1080, or 1081-1139 corresponding to the sequence shown in SEQ ID NO: 696.

[0092] In some embodiments of the present disclosure, the Cas12 protein has a mutation in the RuvC domain corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in amino acid residues 637-689, 885-959, or 1081-1139 corresponding to the sequence shown in SEQ ID NO: 696.

[0093] In some embodiments of the present disclosure, the Cas12 protein has a mutation in the helical domain corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in amino acid residues 42-195, 291-479, or 690-846 corresponding to the sequence shown in SEQ ID NO: 696.

[0094] In some embodiments of the present disclosure, the Cas12 protein has a mutation in the 1st to 41st amino acid residues corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in the 42nd to 195th amino acid residues corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in the 196th to 290th amino acid residues corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in the 291st to 358th amino acid residues corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in the 359th to 479th amino acid residues corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation in the 480th to 636th amino acid residues corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 637-689 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 690-846 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 847-884 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 885-959 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 960-1080 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residues 1081-1139 corresponding to the sequence shown in SEQ ID NO: 696.

[0095] In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 696 at positions corresponding to amino acid residues 1081-1139 of the sequence set forth in SEQ ID NO: 696.

[0096] In some embodiments of the present disclosure, the Cas12 protein corresponds to SEQ ID 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 19, 21, 22, 23, of the sequence shown in NO:696 24, 25, 26, 27, 28, 29, 30, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 46, 47, 48, 49, 50, 51, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 88, 8 9, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 101, 102, 103, 104, 105, 106, 108, 109, 110, 111, 112, 114, 115, 116, 117, 118, 119, 120, 121, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 151, 152, 153, 154, 155, 156, 1 57, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 169, 170, 171, 172, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 194, 195, 196, 197, 198, 199, 200, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222 2, 223, 224, 225, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 242, 243, 244, 245, 247, 248, 249, 250, 251, 252, 253, 255, 256, 257, 258, 259, 260, 261, 262, 263, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 281, 282, 283, 284, 285, 286, 287, 288, 289,290、291、292、293、294、295、296、297、298、299、300、301、302、303、305、306、308、309、310、313、315、316、317、318、319、320、321、322、323、324、325、327、328、329、330、331、332、333、334、335、336、337、339、340、341、342、343、344、345、346、347、348、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、431、432、433、435、436、437、439、440、441、442、443、444、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、467、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、496、497、499、500、501、502、503、504、506、507、508、509、510、511、512、513、514、515、516、517、518、519、520、521、522、523、524、525、526、527、528、529、531、532、533、534、535、536、537、538、539、540、541、542、543、544、545、546、547、548、549、550、552、553、555、556、557、558、559、560、561、562、563、564、565、566、567、568、569、570、571、572、573、574、575、576、577、578、579、580、581、582、583、584、585、586、587、589、590、592、593、594、595、596、597、598、599、601、602、603、604、605、606、607、608、609、610、611、612、613、614、615、616、618、619、620、621、622、623、624、625、626、627、628、630、631、632、633、634、635、636、637、638、639、640、641、642、643、644、645、646、647、648、649、650、651、652、653、654、655、656、657、658、659、660、661、662、663、664、665、666、667、668、669、670、671、672、673、674、675、676、678、679、680、681、683、684、685、686、688、689、691、692、693、694、695、696、697、698、699、700、701、702、703、704、705、706、707、708、709、710、711、712、713、715、716、717、719、720、721、722、723、724、725、727、728、729、730、731、732、733、734、736、737、738、739、740、741、742、743、744、745、746、747、748、749、751、752、753、754、755、756、758、759、760、761、762、764、765、766、767、768、769、771、772、773、774、775、776、779、780、781、782、783、784、785、786、787、789、790、791、792、794、795、797、798、800、801、802、804、805、806、807、808、809、810、811、812、813、814、815、817、818、819、821、822、823、824、825、826、827、828、829、830、831、832、833、834、835、836、837、838、839、840、841、842、844、845、846、847、848、849、850、851、852、853、854、855、856、857、858、859、860、862、863、864、865、866、867、868、870、872、873、874、875、876、877、879、880、881、882、883、884、885、886、887、888、890、891、892、893、894、895、896、897、898、899、900、901、902、903、904、905、906、909、910、911、912、913、914、916、917、918、919、920、921、922、923、924、925、926、927、928、930、931、932、933、934、935、936、937、938、939、940、941、942、943、944、945、946、947、948、949、950、951、952、953、954、955、956、957、958、959、960、961、963、964、965、966、967、968、969、970、971、972、973、974、975、976、977、979、980、981、982、983、985、986、987、988、989、990、991、992、993、994、995、996、997、998、999、1000、1001、1002、1003、1004、1005、1006、1007、1008、1009、1010、1011、1012、1013、1014、1015、1016、1017、1018、1021、1023、1024、1025、1026、1027、1028、1029、1030、1031、1032、1033、1035、1036、1037、1038、1039、1040、1041、1042、1043、1044、1045、1046、1047、1048、1049、1050、1052、1053、1055、1056、1057、1058、1059、1060、1061、1062、1063、1064、1065、1066、1067、1068、1069、1070、1071、1072、1073、1074、1075、1076、1077、1078、1079、1080、1081、1082、1083、1084、1085、1086, 1087, 1088, 1089, 1090, 1091, 1092, 1093, 1094, 1095, 1096, 1097, 1098, 1100, 1101, 1102, 1103, 1104, 1105, 1107, 1108, 1109, 1110, 1111, 1112, 1113, 1115, 1116, 1117 At least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, or at least twelve of the amino acid residues at positions 17, 1120, 1122, 1123, 1124, 1125, 1128, 1129, 1130, 1131, 1132, 1133, 1134, and 1135 have mutations. Optionally, the mutations are to any other naturally occurring amino acid residues. Further optionally, the mutations are to residues R, H, K, or A. In some embodiments, the mutations are to residues R. In some embodiments, the mutations are to residues A. In some embodiments, the mutations are to residues H. In some embodiments, the mutations are to residues K.

[0097] In some embodiments of the present disclosure, the Cas12 protein is at positions 1, 2, 3, 4, 5, 7, 10, 24, 30, 48, 51, 55, 58, 59, 66, 108, 118, 138, 141, 175, 178, 185, 186, 257, 333, 352, 356, 375, 376, 378, 379, 383, 397, 400, 416, 426, 443, 449, 456, 459, 462, 469, 484, 485, 509, 561, 597, 607, 609, 623, 638, 639, 640, 697, 722, 731, 733, 755, 758, 771, 773, 779, 781, 784, 785, 786, 789, 792, 794, 798 , 822, 823, 825, 826, 829, 830, 833, 834, 836, 842, 845, 846, 847, 850, 851, 853, 855, 856, 858, 859, 860, 866, 884, 892, 893, 900, 904, 926, 956, 985, 988, 989, 992, 99 or K. Further alternatively, the mutation is to an R residue.

[0098] In some embodiments of the present disclosure, the Cas12 protein is at positions 12, 29, 35, 36, 40, 53, 57, 60, 64, 71, 72, 73, 75, 94, 95, 96, 97, 99, 137, 148, 149, 153, 164, 167, 171, 172, 174, 177, 190, 192, 194, 199, 204, 207, 208, 211, 215, 228, 232, 236, 238, 244, 248, 253, 256, 258, 261, 262, 275, 282, 286, 292, 298, 300, 302, 320, 324, 328, 332, 336, 339, 366, 373, 374, 384, 389, 393, 395, 415, 432, 436, 440, 453, 458, 460, 471, 472, 474, 506, 508, 510, 519, 523, 526, 528, 531, 534, 544, 550, 570, 571, 573, 592, 594, 596 , 613, 615, 616, 619, 622, 624, 626, 648, 650, 651, 652, 658, 661, 663, 684, 686, 689, 693, 704, 744, 783, 787, 800, 849, 854, 876, 888, 890, 891, 894, 912, 940, 941, 942, 945, 946, 948, 961, 971, 979, 987, 998, 1000, 1002, 1006, 10 or a combination thereof. The present invention further comprises a mutation in at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven or at least twelve of amino acid residues at positions 1090, 1091, 1092, 1093, 1094, 1095, 1096, 1097, 1098, 1099, 1100, 1111 or 112 of amino acid residues at positions 1090, 1091, 1092, 1093, 1094, 1095, 1101, 1102, 1096, 1097, 1103, 1104, 1098, 1099, 1111 or 112 of amino acid residues at positions 1090, 1091, 1092, 1093, 1094, 1095, 1096, 1097, 1098, 1099, 1101, 1102, 1099, 1103, 1104, 1099, 1104, 1099, 1105

[0099] In some embodiments of the present disclosure, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16 or more amino acid residues corresponding to the sequence shown in SEQ ID NO: 696:

[0100] I33, G184, S185, Q186, G194, N195, G196, G197, N245, G256, L260, Y278, S2 85. Y316, H350, D352, A355, A356, C385, P386, H387, G390, K391, N392, D42 9. Q461, Q462, Q469, E485, S491, K521, P525, L611, K629, K631, N633, D841 , N898, K987, A988, G989, Q990, T991, D1010, E1013, A1136, K1138, T1139.

[0101] In some embodiments of the present disclosure, the Cas12 protein has any one, any two, any three, any four, any five, any six, any seven, any eight, any nine, any ten, any eleven, any twelve, any thirty-three, any fourteen, any fifteen, or any sixteen or more amino acid mutations as shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 696:

[0102] I33R, G184R, S185R, Q186R, G194R, N195R, G196R, G197R, N245R, G256R, L260R, Y278R, S2 85R, Y316R, H350R, D352R, A355R, A356R, C385R, P386R, H387R, G390R, K391R, N392R, D42 9R, Q461R, Q462R, Q469R, E485R, S491R, K521R, P525R, L611R, K629R, K631R, N633R, D841 R, N898R, K987R, A988R, G989R, Q990R, T991R, D1010R, E1013R, A1136R, K1138R, T1139R.

[0103] In some embodiments of the present disclosure, the Cas12 protein has any one, any two, any three, any four, any five, or more amino acid mutation combinations of the sequence shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 696 (various mutation combinations are separated by commas): 186+352+1+426+846+858+860, 186+352+1+426+846+860, 186+352+1+5+426+858+860, 186+352+1+3+426+858+860, 186+352+3+426+860, 186+352+1+333+426+858+860, 186+352+1+426+485+858+860, 186+ 352+1+5+426+860、186+352+3+426+858+860、186+352+5+426+858+860、186+352+333+426+858+860、186+352+333+426+860、186+352+426+846+860、186+352+5+426+860、186+352+1+3+426+860、5+426+860、186+352+426+846+858+860、186+352+1+426+485+860、426+858+860、426+ 858、186+352+426+860、186+352+426+485+858+860、186+352+426+485+860、426+846+858+860、7+426+846、186+352+860、184+186+352+376+1132、5+333+426、5+426+858、186+352+1+333+426+860、184+186+352+3+107+426、186+352+5+426、333+376+426、186+352+5、2+5+846+858、 186+352+3+426、5+846+858、186+352+426+846、186+352+7、186+352+3+376+426、186+352+376+426+860+865、186+352+376+426、3+846+860、846+858+988、3+426+858、3+846+858、3+858+860、186+352+858、184+186+352+860、333+426、184+186+352+3+376、3+426+860、846+860+988、3+860、184+186+352+3+639、186+352+333+426、846+585+860、186+352+426、333+426+846、333+426+485、5+846+860、3+846+988、184+186+352+376、3+333+426、186+352+426+485、333+426+858、333+426+1132、428+485、186+352+333+352+376+426、858+860、186+352+376+426+485+ 860, 184+186+352+426, 186+352+639, 5+858+988, 3+858+988, 5+858, 3+858, 184+186+352+846, 184+186+352+639, 858+988, 184+186+352+426+1132, 186+352+1132, 184+186+352+5, 184+186+352+858, 858+860+1132, 3+5, 426+649, 186+352+426+485+860, 186+352+333, 184+186+35 2. 186+376, 846+860, 858+988+1132, 846+858, 333+376, 376+426, 184+186+352+3, 3+846+1132, 5+846+1132, 186+352+426+1132, 376+426+485+660, 426, 5+846, 846+860+1132, 333+376+485, 184+186+352+639+1132, 352+426, 333+485, 184+186+352+333, 846+858+1132, 333+426+86 0, 186+352+988, 5+860, 846+988, 186+352, 3+846, 846+1132, 184+186+352+1132, 186+485, 988+1132, 184+186+352+485, 376+485, 5+1132, 3+7, 186+352+485, 184+186+352+7, 184+186+352+333+336, 3+1132, 426+858+988, 186+352+376, 186+352+3 and 186+352+333+336+352+376+426. Further optionally, the mutation is to residue R or A.

[0104] In some embodiments of the present disclosure, the Cas12 protein has any one of the following amino acid mutation combinations of the sequence shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 696 (various mutation combinations are separated by commas): 186+352+1+426+846+858+860, 186+352+1+426+846+860, 186+352+1+5+426+858+860, 186+352+1+3+426+858+860, 186+352+1+3+426+858+860, 186+352+1+426+846+860 2+3+426+860、186+352+1+333+426+858+860、186+352+1+426+485+858+860、186+352+1+5+426+860、186+352+3+426+858+860、186+352+5+426+858+860、186+352+333+42 6+858+860、186+352+333+426+860、186+352+426+846+860、186+352+5+426+860、186+352+1+3+426+860、5+426+860、186+352+426+846+858+860、186+352+1+426+485+86 0, 426+858+860, 7+426+858, 186+352+426+860, 186+352+426+485+858+860, 186+352+426+485+860, 426+846+858+860, 7+426+846, 186+352+860; further optionally, the mutation is to residue R or A.

[0105] In some embodiments of the present disclosure, the Cas12 protein undergoes any one, any two, any three, any four, any five, or more amino acid mutations of the sequence shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 696:

[0106] D352R+Q186R、D352R+L260R、D352R+A355R、A355R+L260R、P386R+C385R、E485R+Q462R、D352R+Q186R、A355R+L260R、G184R+R186Q、D352R+Q186R+I33R、D352R+Q186R+G184R、D352R+Q186R+S185R、D352R+Q186R+G256R、D352R+Q186R+Y278R、D352R+Q186R+S285R、D352R+Q186R+Y316R、D352R+Q186R+H350R、D352R+Q186R+A356R、D352R+Q186R+Q469R、D352R+Q186R+S491R、D352R+Q186R+K521R、D352R+Q186R+P525R、D352R+Q186R+K629R、D352R+Q186R+N633R、D352R+Q186R+D841R、D352R+Q186R+N898R、D352R+Q186R+K987R、D352R+Q186R+T991R、D352R+Q186R+D1010R、D352R+Q186R+E1013R。

[0107] In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 2nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 3rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 4th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 5th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 7th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 10th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 24th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 30th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 48th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 51st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 55th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 58th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 59th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue position 66 corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue position 108 corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue position 118 corresponding to the sequence as shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 138th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 141st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 175th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 178th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 185th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 186th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 257th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 333rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 352nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 356th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 375th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 376th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 378th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 379th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 383 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 397 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 400th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 416th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 426th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 443rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 449th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 456th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 459th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 462nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 469th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 484th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 485th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 509th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 561st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 597th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 607 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 609 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 623rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 638th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 639th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 640th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 697th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 722nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 731st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 733rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 755th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 758th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 771st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 773rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 779th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 781st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 784 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 785 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 786th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 789th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 792th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 794th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 798th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 822nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 823rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 825th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 826th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 829th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 830th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 833rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 834th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 836th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 842 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 845 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 846th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 847th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 850th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 851st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 853rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 855th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 856th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 858th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 859th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 860th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 866th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 884th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 892th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 893rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 900th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 904th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 926th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 956th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 985th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 988th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 989th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 992th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 993rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 996th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1016th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1033rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1045th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1050th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1073rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 1074th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 1095 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 1100 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 1124 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 1129 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 1132 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is to residue R.

[0108] In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 651 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is to residue A.

[0109] In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 891 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is to residue A.

[0110] In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 1082 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is to residue A.

[0111] In some embodiments of the present disclosure, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16 or more amino acid residues corresponding to the sequence shown in SEQ ID NO: 52:

[0112] V15, Q172, A173, G182, E183, G184, K185, K186, G239, V243, D264, E271, Y295, L297, N317, T329, E331, I335, K339, K339, N347, E363, H366, V426, K429, S430, L433, S452, S455, S465, E493, P497, K587, E768, T825, A911, G914, K915, I916, K918, T919, T920, A922, E940.

[0113] In some embodiments of the present disclosure, the Cas12 protein has any one, any two, any three, any four, any five, any six, any seven, any eight, any nine, any ten, any eleven, any twelve, any thirty-three, any fourteen, any fifteen, or any sixteen or more amino acid mutations as shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 52:

[0114] V15R, Q172R, A173W, G182R, E183R, G184R, K185R, K186R, G239R, V243R, D264R, E271R, Y295R, L297R, N317R, T329R, E331R, I335R, K339R, K339R, N347E, E363R, H366R, V426R, K429R, S430R, L433R, S452R, S455R, S465R, E493R, P497R, K587R, E768R, T825R, A911R, G914R, K915R, I916R, K918R, T919R, T920R, A922R, E940R.

[0115] In some embodiments of the present disclosure, the Cas12 protein undergoes any one, any two, any three, any four, any five, or more amino acid mutations of the sequence shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 52:

[0116] N347E+K339R, Q172R+S452R, Q172R+T920R, Q172R+V426R, S452R+T920R, V426R+S452R, V426R+T920R.

[0117] In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 15th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 172nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 173rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 182nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 183rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 184th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 185th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 186th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 239th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 243rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 264th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 271st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 295th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 297th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 317 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 329 corresponding to the sequence shown in SEQ ID NO: 52.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 331st amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 335th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 339th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 339th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 347th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 363rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 366th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 426th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 429th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 430th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 433rd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 452nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 455th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 465th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 493 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at amino acid residue 497 corresponding to the sequence shown in SEQ ID NO: 52.In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 587th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 768th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 825th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 911th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 914th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 915th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 916th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 918th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 919th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 920th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 922nd amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. In some embodiments of the present disclosure, the Cas12 protein has a mutation at the 940th amino acid residue corresponding to the sequence as shown in SEQ ID NO: 52. Further optionally, the mutation is a mutation to residue R.

[0118] On the other hand, a technical solution provided by the present disclosure is: a guide polynucleotide comprising (i) a direct repeat sequence, wherein the direct repeat sequence has at least 50% identity with any one of SEQ ID NOs: 54-583, 704, and (ii) a guide sequence engineered to hybridize with a target nucleic acid; the direct repeat sequence is connected to the guide sequence, and the guide polynucleotide is capable of forming a complex with the Cas12 protein and guiding the complex to bind specifically to the sequence of the target nucleic acid.

[0119] In some embodiments of the present disclosure, the direct repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0120] In some embodiments of the disclosure, the direct repeat sequence has at least 60% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0121] In some embodiments of the disclosure, the direct repeat sequence has at least 65% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0122] In some embodiments of the disclosure, the direct repeat sequence has at least 70% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0123] In some embodiments of the disclosure, the direct repeat sequence has at least 75% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0124] In some embodiments of the disclosure, the direct repeat sequence has at least 80% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0125] In some embodiments of the disclosure, the direct repeat sequence has at least 85% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0126] In some embodiments of the disclosure, the direct repeat sequence has at least 90% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0127] In some embodiments of the disclosure, the direct repeat sequence has at least 95% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0128] In some embodiments of the disclosure, the direct repeat sequence has at least 96% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0129] In some embodiments of the disclosure, the direct repeat sequence has at least 97% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0130] In some embodiments of the disclosure, the direct repeat sequence has at least 98% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0131] In some embodiments of the disclosure, the direct repeat sequence has 100% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704.

[0132] In a preferred embodiment, the Cas12 protein is the Cas12 protein described in the present disclosure.

[0133] In specific embodiments of the present disclosure, the guide sequence comprises 15-60 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-50 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-40 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-35 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-30 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-25 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 18-25 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 20-25 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 18-22 nucleotides. In specific embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.

[0134] In specific embodiments of the present disclosure, the guide sequence hybridizes to the target nucleic acid, and the guide sequence is 90%-100% complementary to the target nucleic acid.

[0135] In specific embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid.

[0136] In specific embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid with no more than one nucleotide mismatch between the guide sequence and the target nucleic acid.

[0137] In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-100 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-90 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-80 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-70 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-60 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 15-50 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 15-40 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 20-40 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 20-30 nucleotides. In specific embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.

[0138] In specific embodiments of the disclosure, the guide sequence is located at the 3' end of the direct repeat sequence.

[0139] In specific embodiments of the disclosure, the guide sequence is located at the 5' end of the direct repeat sequence.

[0140] In specific embodiments of the present disclosure, the guide polynucleotide further comprises a tracrRNA.

[0141] In some embodiments of the disclosure, the tracrRNA sequence has at least 50% identity compared to any one of SEQ ID NOs: 584-695. In some embodiments of the disclosure, the tracrRNA sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to any one of SEQ ID NOs: 584-695.

[0142] In some embodiments of the present disclosure, the tracrRNA sequence is optionally selected from the sequence shown in any one of SEQ ID NOs: 584-695.

[0143] In a specific embodiment of the present disclosure, the tracrRNA can be complementary to the direct repeat sequence. In general, the complementary pairing is a complementary pairing of partial bases. In a specific embodiment of the present disclosure, the tracrRNA can interact with the direct repeat sequence.

[0144] In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence consisting of 1-10 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence consisting of 4 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a 5'-GAAA-3' sequence.

[0145] In specific embodiments of the present disclosure, the tracrRNA sequence is located at the 3' end of the direct repeat sequence.

[0146] In specific embodiments of the present disclosure, the tracrRNA sequence is located at the 5' end of the direct repeat sequence.

[0147] In specific embodiments of the present disclosure, the tracrRNA comprises 10-200 nucleotides. In specific embodiments of the present disclosure, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10- or 40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In specific embodiments of the present disclosure, the tracrRNA comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 94, 95, 96, 97, 98, 99, or 100 nucleotides.

[0148] SEQ ID NO: 1-53 shows the amino acid sequence of the Cas protein.

[0149] SEQ ID NO: 54-583 shows the direct repeat sequence (DR) corresponding to the Cas protein. When multiple DR sequences corresponding to a specific Cas protein are listed, one can be used.

[0150] SEQ ID NOs: 584-695 show the tracrRNA sequences corresponding to Cas proteins. When SEQ ID NOs: 584-695 do not list the tracrRNA sequence corresponding to a particular Cas protein, the gRNA may contain only the guide sequence and DR sequence, without the tracrRNA sequence. When SEQ ID NOs: 584-695 list the tracrRNA sequence corresponding to a particular Cas protein, the gRNA may or may not contain the tracrRNA sequence in addition to the guide sequence and DR sequence. When multiple tracrRNA sequences corresponding to a particular Cas protein are listed, one may be used.

[0151] On the other hand, a technical solution provided by the present disclosure is: a Cas12 inactivated variant, characterized in that the Cas12 inactivated variant is a nuclease activity inactivated variant of the Cas12 protein as described in the present disclosure.

[0152] In this article, according to the context, the scope of reference of the Cas12 protein may include the Cas12 inactivated variant. However, in view of the importance of the Cas12 inactivated variant (non-limiting examples such as the Cas12 inactivated variant fused with a deaminase for single base editing, fused with a transcription activation domain or a transcription repression domain for transcriptional regulation, etc.), it will be described separately herein; this does not mean that the scope of reference of the Cas12 protein must not include the Cas12 inactivated variant.

[0153] In a specific embodiment of the present disclosure, the Cas12 inactivated variant is a variant in which the nuclease activity is completely inactivated, i.e., a dead Cas12 inactivated variant (dCas12). The dCas12 can only bind to the target nucleic acid under the mediation of the guide polynucleotide, and has no or almost no function of cutting the target nucleic acid. For example, the target nucleic acid cutting efficiency of the dCas12 is ≤20%, ≤15%, ≤10%, ≤5%, ≤4%, ≤3%, ≤2% or ≤1% of the target nucleic acid cutting efficiency of the Cas12 protein before the inactivation mutation.

[0154] In a specific embodiment of the present disclosure, the Cas12 inactivated variant is a variant with partially inactivated nuclease activity. Further, the variant with partially inactivated nuclease activity is a Cas12 nickase (nCas12), which binds to the target nucleic acid under the mediation of a guide polynucleotide and then cuts one of the single strands in the double-stranded target nucleic acid without cutting the other single strand.

[0155] In a preferred embodiment of the present disclosure, the Cas12 inactivated variant is an inactivated Ruvc domain of the Cas12 protein.

[0156] In a preferred embodiment of the present disclosure, the inactivated Cas12 variant is an inactivated Ruvc-I, Ruvc-Ⅱ or Ruvc-Ⅲ domain of the Cas12 protein.

[0157] In a preferred embodiment of the present disclosure, the inactivated variant of Cas12 is obtained by introducing an inactivating mutation into the Ruvc-I, Ruvc-Ⅱ or Ruvc-Ⅲ domain of the Cas12 protein.

[0158] In a specific embodiment of the present disclosure, the inactivating mutation is selected from one or more of D651A, E891A and D1082A corresponding to the amino acid sequence shown in SEQ ID NO: 696.

[0159] In a specific embodiment of the present disclosure, the inactivating mutations are D651A, E891A and D1082A corresponding to the amino acid sequence shown in SEQ ID NO: 696.

[0160] In a specific embodiment of the present disclosure, the PAM sequence recognizable by the inactivated Cas12 variant is the same as the PAM sequence recognizable by the Cas12 protein.

[0161] In a specific embodiment of the present disclosure, the PAM sequence (5'→3') that can be recognized by the Cas12 inactivated variant is selected from any one or more of the following:

[0162] A, C, T, G,

[0163] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,

[0164] <h2 style=";text-align:left;direction:ltr">NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT ANG、GNA、GTT、NGA、TAA、GTA、GGN、GNT、NCG、ATT、CCA、CNN、AAA、AAC、ATN、GA G、CTG、ACG、NAA、TAN、NAT、CNA、GCN、GTC、NCN、CTN、CNC、ANT、NNC、CAG、NAN、 ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0165] <h2 style=";text-align:left;direction:ltr">NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA,CANC,ACGA,CCAC,CCGG,CTNG,CNGN,GGTA,NGNC,GTTT,CTAA,TNCT,CTGN,NGAC,TGTA,TANN,GCNT,GCTC,CNCG,AAAN,CCNT,GANA,CACA,CTNA,ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG,GTGG,GNGN,ACCA,NTAA,ACTN,NCTG,NCTA,TTTT,GCNG,NTAG,CAAA,GGNA,CNTN,TTAG,TCTG,NCTN,TATG,GCGT,TANT,GGGT,NACN,ACTG,CCNG,GNNT<h2 style=";text-align:left;direction:ltr">CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN<h2 style=";text-align:left;direction:ltr">GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA,AAGA,CTTT,AAAC,AGGN,ACNT,NTGT,CTTN,ATCA,NACT,NNAG,NGTN,NAAC,TGCG,GGNT,ATAN,TTGC,ANCN,CCCC,ANGA,NGCG,TCTC,CTCG,ATNA,AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,

[0166] The N is A, T, C or G.

[0167] On the other hand, a technical solution provided by the present disclosure is: a fusion protein or conjugate, which comprises the following elements: (1) the Cas12 protein as described in the present disclosure, or the inactivated variant of Cas12 as described in the present disclosure; and (2) a homologous or heterologous functional domain.

[0168] In this article, according to the context, the scope of reference of the Cas12 protein may include the Cas12 inactivated variants. However, in view of the importance of Cas12 inactivated variants (non-limiting examples such as the Cas12 inactivated variants fused with deaminase for single base editing, fused with transcription activation domain or transcription repression domain for transcriptional regulation, etc.), this article will describe them separately and emphatically; this does not mean that the scope of reference of the Cas12 protein does not necessarily include the Cas12 inactivated variants. In some embodiments of the present disclosure, a fusion protein is provided, the fusion protein comprising: (1) the Cas12 protein as described in the present disclosure, or the Cas12 inactivated variant as described in the present disclosure; and (2) homologous or heterologous functional domains.

[0169] In some embodiments of the present disclosure, a fusion protein is provided, comprising: (1) a Cas12 protein as described herein; and (2) a homologous or heterologous functional domain.

[0170] In a specific embodiment of the present disclosure, a conjugate is provided, comprising: (1) a Cas12 protein as described herein, or an inactivated variant of Cas12 as described herein; and (2) a homologous or heterologous functional domain.

[0171] In a specific embodiment of the present disclosure, a conjugate is provided, comprising: (1) the Cas12 protein as described herein; and (2) a homologous or heterologous functional domain.

[0172] In some embodiments, the functional domain has an enzymatic activity that modifies a target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristoylation activity and / or demyristoylation activity.

[0173] In a specific embodiment of the present disclosure, the inactivating mutation is selected from one or more of D651A, E891A and D1082A corresponding to the amino acid sequence shown in SEQ ID NO: 696.

[0174] In a specific embodiment of the present disclosure, the inactivating mutations are D651A, E891A and D1082A corresponding to the amino acid sequence shown in SEQ ID NO: 696.

[0175] In some embodiments, the functional domain is selected from one or more of the following: a nuclease (e.g., FokI), a methyltransferase, a demethylase, a DNA repair enzyme, a DNA damaging enzyme, a deaminase, a dismutase, an alkylase, a depurinase, an oxidase, a pyrimidine dimer-forming enzyme, an integrase, a transposase, a recombinase, a polymerase, a ligase, a helicase, a photolyase, a glycosylase, a deglycosylase, an acetyltransferase, a deacetylase, a kinase, a phosphatase, a ubiquitin ligase, a deubiquitinating enzyme, an adenylylase, a deadenylase, a SUMOylating enzyme, a deSUMOylating enzyme, a myristoylase, and / or a demyristoylase.

[0176] In specific embodiments of the present disclosure, the homologous or heterologous functional domains are optionally selected from one, two, three, four or more of the following: subcellular localization signals, DNA binding domains, protease domains, transcription activation domains, transcription repression domains, nuclease domains, deaminase domains, uracil DNA glycosylase domains (UDG), uracil DNA glycosylase inhibitory domains (UGI), methylases, demethylases, transcription release factors, histone acetylase domains, histone deacetylase domains, DNA ligases, affinity tags, reporter tags, affinity domains and reporter domains.

[0177] In some embodiments of the present disclosure, the subcellular localization signal is selected from: a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, and a chloroplast localization signal.

[0178] In specific embodiments of the present disclosure, the fusion protein or conjugate comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or more of the homologous or heterologous functional domains; the functional domains are the same or different.

[0179] In some embodiments, the fusion protein or conjugate is arbitrarily linked to 0, 1, 2, 3, 4, 5, 6, 7, 8 or more of the functional domains at the N-terminus and / or C-terminus of the Cas12 protein.

[0180] In specific embodiments of the disclosure, the fusion protein comprises 1, 2, 3, 4 or more nuclear localization signals.

[0181] In a specific embodiment of the present disclosure, the fusion protein can be used to achieve base editing, for example, in combination with a guide polynucleotide to achieve base editing. In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal and a deaminase domain.

[0182] In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and optionally one or two UGI domains. The fusion protein can be used to achieve C→T base editing of the target nucleic acid.

[0183] In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal and an adenosine deaminase domain. The fusion protein can be used to achieve A→G base editing of the target nucleic acid.

[0184] In specific embodiments of the present disclosure, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and an adenosine deaminase domain. In specific embodiments of the present disclosure, the fusion protein comprises one, two, or three nuclear localization signals and a deaminase domain. In specific embodiments of the present disclosure, the fusion protein comprises a UGI domain. In specific embodiments of the present disclosure, the fusion protein comprises one, two, or three nuclear localization signals, a deaminase domain, and one or two UGI domains.

[0185] In a specific embodiment of the present disclosure, the fusion protein can be used to achieve transcriptional activation of a specific target gene, for example, in combination with a guide polynucleotide to achieve transcriptional activation of a specific target gene. In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal and a transcriptional activation domain.

[0186] In a specific embodiment of the present disclosure, the fusion protein can be used to achieve transcriptional inhibition of a specific target gene, for example, in combination with a guide polynucleotide to achieve transcriptional inhibition of a specific target gene. In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal and a transcriptional inhibition domain.

[0187] In a specific embodiment of the present disclosure, the fusion protein can be used to achieve methylation of a specific target sequence, for example, in combination with a guide polynucleotide to achieve methylation of a specific target sequence. In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal and a DNA methylation domain.

[0188] In a specific embodiment of the present disclosure, the fusion protein can be used to achieve demethylation of a specific target sequence, for example, in combination with a guide polynucleotide to achieve demethylation of a specific target sequence. In a specific embodiment of the present disclosure, the fusion protein comprises a nuclear localization signal and a DNA demethylation domain.

[0189] In a preferred embodiment of the present disclosure, the nuclease domain comprises a polypeptide having ssDNA cleavage activity and / or a polypeptide having dsDNA cleavage activity.

[0190] In a preferred embodiment of the present disclosure, the nuclease domain comprises a polypeptide having ssDNA cleavage activity.

[0191] In a preferred embodiment of the present disclosure, the nuclease domain comprises a polypeptide having dsDNA cleavage activity.

[0192] In specific embodiments of the present disclosure, the Cas12 protein or inactivated variant is directly or indirectly linked to the homologous or heterologous functional domain.

[0193] In a preferred embodiment of the present disclosure, the direct connection is covalent connection, and the indirect connection is connection via an amino acid linker or a non-amino acid linker.

[0194] In a more preferred embodiment of the present disclosure, the homologous or heterologous functional domain is fused or conjugated at the N-terminus, C-terminus or internally relative to the Cas12 protein or inactivated variant.

[0195] In the present disclosure, the fusion protein refers to the connection between the element (1) and the element (2) via a peptide segment, or a direct connection; the conjugate refers to the connection between the element (1) and the element (2) via a non-peptide chemical bond.

[0196] In a specific embodiment of the present disclosure, the PAM sequence recognizable by the fusion protein or conjugate is the same as the PAM sequence recognizable by the Cas12 protein.

[0197] In a specific embodiment of the present disclosure, the PAM sequence (5'→3') that can be recognized by the fusion protein or conjugate is selected from any one or more of the following:

[0198] A, C, T, G,

[0199] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,

[0200] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,

[0201] <h2 style=";text-align:left;direction:ltr">NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA,CANC,ACGA,CCAC,CCGG,CTNG,CNGN,GGTA,NGNC,GTTT,CTAA,TNCT,CTGN,NGAC,TGTA,TANN,GCNT,GCTC,CNCG,AAAN,CCNT,GANA,CACA,CTNA,ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG,GTGG,GNGN,ACCA,NTAA,ACTN,NCTG,NCTA,TTTT,GCNG,NTAG,CAAA,GGNA,CNTN,TTAG,TCTG,NCTN,TATG,GCGT,TANT,GGGT,NACN,ACTG,CCNG,GNNT<h2 style=";text-align:left;direction:ltr">CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN<h2 style=";text-align:left;direction:ltr">GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA,AAGA,CTTT,AAAC,AGGN,ACNT,NTGT,CTTN,ATCA,NACT,NNAG,NGTN,NAAC,TGCG,GGNT,ATAN,TTGC,ANCN,CCCC,ANGA,NGCG,TCTC,CTCG,ATNA,AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,

[0202] The N is A, T, C or G.

[0203] In a specific embodiment of the present disclosure, the fusion protein can recognize a PAM sequence of 5'-TTN-3'.

[0204] In a specific embodiment of the present disclosure, the fusion protein can recognize a PAM sequence of 5'-TTNC-3'.

[0205] In a specific embodiment of the present disclosure, the fusion protein can recognize a PAM with the sequence 5'-WTN-3'.

[0206] In a specific embodiment of the present disclosure, the fusion protein can recognize a PAM sequence of 5'-ATN-3'.

[0207] In specific embodiments of the present disclosure, the conjugate can recognize a PAM with the sequence 5'-TTN-3'.

[0208] In a specific embodiment of the present disclosure, the conjugate can recognize a PAM with the sequence 5'-TTNC-3'.

[0209] In specific embodiments of the present disclosure, the conjugate can recognize a PAM with the sequence 5'-WTN-3'.

[0210] In a specific embodiment of the present disclosure, the conjugate can recognize a PAM with the sequence 5'-ATN-3'.

[0211] On the other hand, a technical solution provided by the present disclosure is: an isolated nucleic acid encoding the Cas12 protein as described in the present disclosure, the Cas12 inactivated variant as described in the present disclosure, or the fusion protein or conjugate as described in the present disclosure.

[0212] In some embodiments of the present disclosure, the nucleic acid encodes a Cas12 protein as described herein or a fusion protein as described herein.

[0213] In preferred embodiments of the present disclosure, the nucleic acid is codon-optimized for expression in cells.

[0214] In more preferred embodiments of the present disclosure, the nucleic acid is codon-optimized for expression in a eukaryote, a mammal such as a human or non-human mammal, a plant, an insect, a bird, a reptile, a rodent (e.g., mouse, rat), a fish, a worm / nematode, or a yeast.

[0215] On the other hand, a technical solution provided by the present disclosure is: a CRISPR-Cas12 system, the CRISPR-Cas12 system comprising:

[0216] a. a Cas12 protein as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, or an isolated nucleic acid as described herein; and

[0217] b. a guide polynucleotide, or a polynucleotide sequence encoding the guide polynucleotide;

[0218] The Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate forms a complex with the guide polynucleotide; the guide polynucleotide comprises a guide sequence, and the guide sequence is engineered to guide the sequence-specific binding of the complex to the target nucleic acid.

[0219] The isolated nucleic acid encodes the Cas12 protein as described herein, the Cas12 inactivated variant as described herein, or the fusion protein or conjugate as described herein.

[0220] In specific embodiments of the disclosure, the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence.

[0221] In specific embodiments of the present disclosure, the direct repeat sequence has at least 50% identity to any one of SEQ ID NOs: 54-583, 704.

[0222] In specific embodiments of the present disclosure, the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence. Further, in some specific embodiments, the direct repeat sequence has at least 50% identity compared to any one of SEQ ID NOs: 54-583, 704. In some specific embodiments, the direct repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% sequence identity to the sequence shown in any one of SEQ ID NOs: 54-583, 704. Furthermore, in some specific embodiments, the direct repeat sequence comprises or is a sequence shown in any one of SEQ ID NOs: 54-583, 704.

[0223] In specific embodiments of the present disclosure, the guide sequence comprises 15-60 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-50 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-40 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-35 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-30 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 15-25 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 18-25 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 20-25 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 18-22 nucleotides. In specific embodiments of the present disclosure, the guide sequence comprises 20-22 nucleotides. In specific embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides.

[0224] In specific embodiments of the present disclosure, the guide sequence hybridizes to the target nucleic acid, and the guide sequence is 90%-100% complementary to the target nucleic acid.

[0225] In specific embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid.

[0226] In specific embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid with no more than one nucleotide mismatch between the guide sequence and the target nucleic acid.

[0227] In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-100 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-90 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-80 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-70 nucleotides. In a specific embodiment of the present disclosure, the direct repeat sequence comprises 15-60 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 15-50 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 15-40 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 20-40 nucleotides. In a specific embodiment of the present disclosure, the guide sequence comprises 20-30 nucleotides. In specific embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.

[0228] In specific embodiments of the disclosure, the guide sequence is located at the 3' end of the direct repeat sequence.

[0229] In specific embodiments of the disclosure, the guide sequence is located at the 5' end of the direct repeat sequence.

[0230] In specific embodiments of the present disclosure, the guide polynucleotide further comprises a tracrRNA.

[0231] In some embodiments of the disclosure, the tracrRNA sequence is at least 50% identical to any one of SEQ ID NOs: 584-695. In some embodiments of the disclosure, the tracrRNA sequence is at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identical to any one of SEQ ID NOs: 584-695.

[0232] In a specific embodiment of the present disclosure, the tracrRNA can be complementary to the direct repeat sequence. In general, the complementary pairing is a complementary pairing of partial bases. In a specific embodiment of the present disclosure, the tracrRNA can interact with the direct repeat sequence.

[0233] In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence consisting of 1-10 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a nucleotide sequence consisting of 4 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence through a 5'-GAAA-3' sequence.

[0234] In specific embodiments of the present disclosure, the tracrRNA sequence is located at the 3' end of the direct repeat sequence.

[0235] In specific embodiments of the present disclosure, the tracrRNA sequence is located at the 5' end of the direct repeat sequence.

[0236] In specific embodiments of the present disclosure, the tracrRNA comprises 10-200 nucleotides. In specific embodiments of the present disclosure, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10- or 40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In specific embodiments of the present disclosure, the tracrRNA comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 94, 95, 96, 97, 98, 99, or 100 nucleotides.

[0237] In a preferred embodiment of the present disclosure, the guiding polynucleotide is the guiding polynucleotide as described in the present disclosure.

[0238] In specific embodiments of the present disclosure, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA.

[0239] In a preferred embodiment of the present disclosure, the DNA is eukaryotic DNA; preferably, the eukaryotic DNA is non-human mammal DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent DNA, fish DNA, worm / nematode DNA or yeast DNA.

[0240] In specific embodiments of the present disclosure, the target nucleic acid is a disease or disorder-associated gene or a signal transduction biochemical pathway-associated gene, or the target nucleic acid is a reporter gene; for example, the disease or disorder is a blood system disease or disorder, an ophthalmic disease or disorder, a nervous system disease or disorder, a respiratory system disease or disorder, a liver disease or disorder, a metabolic system disease or disorder, cancer, or an infectious disease.

[0241] In some embodiments, the target nucleic acid is a gene listed in Table 27.

[0242] In some embodiments, the target nucleic acid is a gene associated with a disease or condition, and the disease or condition is selected from: hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidosis, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, alpha-1 antitrypsin deficiency, alpha-mannosidosis, alpha-thalassemia, beta-thalassemia, Alzheimer's disease, Budd-Bieder syndrome, albicans punctata, leukocyte adhesion deficiency type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens camp Malnutrition, pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult glucan body disease, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophosphatasia, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin lymphoma, non-muscle-depleting myeloma Bladder cancer, nonalcoholic fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, Charcot-Marie-Tooth disease type 1A, Charcot-Marie-Tooth disease type 2A, pulmonary hypertension, Friedreich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type 1, coronary artery disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis dyslexia, urge incontinence, acute intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial Tay-Sachs disease, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, junctional epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy,Canavan disease, cocaine addiction, Krabbe disease, Crigler-Najjar syndrome, oral cancer, Happy Marionette syndrome, diffuse intrinsic pontine glioma, Lafora disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, anemia of chronic kidney disease, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Netherton syndrome, ornithine transcarbamylase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage diseases, meat tumors, breast cancer, Rett syndrome, triple-negative breast cancer, Sandhoff disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscinosis, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibulopathy, Stargardt's disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease I Type a, glycogen storage disease type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, head and neck squamous cell carcinoma, Wilson's disease, stable angina, Ussher syndrome, choroideremia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, novel coronavirus infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, critical limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive dystrophic epidermolysis bullosa, infantile malignant Osteoporosis, dystrophic epidermolysis bullosa, morphea, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic disease, lipoprotein lipase deficiency, severe congenital neutropenia, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, transthyretin (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.

[0243] In some embodiments, the transthyretin (ATTR) amyloidosis-associated genes include but are not limited to ATTR;

[0244] The genes related to Leber hereditary optic neuropathy include but are not limited to MT-ND4;

[0245] The AATD liver disease related genes include but are not limited to AATD;

[0246] The AATD lung disease related genes include but are not limited to AATD;

[0247] The graft-versus-host disease related genes include but are not limited to thymidine kinase gene;

[0248] The genes related to hereditary retinal dystrophy include but are not limited to RPE65;

[0249] The spinal muscular atrophy-related genes include but are not limited to SMN1;

[0250] The osteoarthritis related genes include but are not limited to TGF-β1;

[0251] The related genes of hemophilia A include but are not limited to factor VIII;

[0252] The related genes of hemophilia B include but are not limited to factor IX;

[0253] The cystic fibrosis related genes include but are not limited to CFTR;

[0254] The Parkinson's disease-related genes include but are not limited to Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB and REST;

[0255] The Usher syndrome related genes include but are not limited to USH2A;

[0256] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include but are not limited to BCL11A, HBG, HBA, and HBB;

[0257] The pulmonary hypertension related genes include but are not limited to eNOS;

[0258] The Stargardt disease-related genes include but are not limited to ABCA4;

[0259] The age-related macular degeneration-related genes include but are not limited to VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang 2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene and FGF2;

[0260] The glaucoma-related genes include but are not limited to AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4 and carbonic anhydrase CA12;

[0261] The idiopathic pulmonary fibrosis related genes include but are not limited to CTGF;

[0262] The hyperlipidemia-related genes include but are not limited to PCSK9;

[0263] The Alzheimer's disease related genes include but are not limited to NGF;

[0264] The coronary heart disease related genes include but are not limited to VEGFA and bFGF;

[0265] The related genes of chronic kidney disease anemia include but are not limited to EPO;

[0266] The related genes of congenital amaurosis include but are not limited to RPE65;

[0267] The retinitis pigmentosa related genes include but are not limited to PDE6B;

[0268] The phenylketonuria related genes include but are not limited to PAH; and / or

[0269] The epilepsy-related genes include but are not limited to GAT1.

[0270] On the other hand, a technical solution provided by the present disclosure is: a vector system, wherein the vector system comprises one or more recombinant vectors, and the recombinant vector comprises the isolated nucleic acid as described in the present disclosure, or the CRISPR-Cas12 system as described in the present disclosure.

[0271] In a specific embodiment of the present disclosure, the recombinant vector further comprises a regulatory sequence.

[0272] In a specific embodiment of the present disclosure, the vector system comprises one or more recombinant vectors, which comprise a polynucleotide sequence encoding the Cas12 protein, Cas12 inactivated variant or fusion protein or conjugate disclosed herein, and a polynucleotide sequence encoding the guide polynucleotide.

[0273] In a specific embodiment of the present disclosure, the polynucleotide sequence encoding the Cas12 protein, Cas12 inactivated variant or fusion protein or conjugate is operably linked to the regulatory sequence 1.

[0274] In specific embodiments of the present disclosure, the polynucleotide sequence encoding the guide polynucleotide is operably linked to regulatory sequence 2.

[0275] Furthermore, in a specific embodiment of the present disclosure, the regulatory sequence 1 and the regulatory sequence 2 are identical or different sequences.

[0276] In a preferred embodiment of the present disclosure, the regulatory sequence is optionally selected from: one or more of a promoter, an enhancer, an internal ribosome entry site and a transcription termination signal; the promoter is, for example, a constitutive promoter, an inducible promoter, a broad-spectrum promoter or a tissue-specific promoter, and / or the transcription termination signal is, for example, a polyadenylation signal or a poly-U sequence.

[0277] In a specific embodiment of the present disclosure, the backbone of the recombinant vector is an adeno-associated virus vector, a lentivirus vector or a virus-like particle.

[0278] In a preferred embodiment of the present disclosure:

[0279] When the backbone is an adeno-associated viral vector, the adeno-associated viral vector is a recombinant adeno-associated viral vector of serotype AAV1, AAV2, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12 or AAV13;

[0280] When the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; preferably, the isolated nucleic acid is linked to an aptamer sequence;

[0281] When the backbone is a virus-like particle, the isolated nucleic acid is linked to a gene encoding a gag protein.

[0282] On the other hand, a technical solution provided by the present disclosure is: a delivery system, comprising: (1) a delivery tool, and (2) a Cas12 protein as described in the present disclosure, a guide polynucleotide as described in the present disclosure, a Cas12 inactivated variant as described in the present disclosure, a fusion protein or conjugate as described in the present disclosure, a nucleic acid as described in the present disclosure, a CRISPR-Cas12 system as described in the present disclosure, or a vector system as described in the present disclosure.

[0283] In a preferred embodiment of the present disclosure, the delivery vehicle is a virus, lipid nanoparticle, nanoparticle, liposome, exosome, microbubble or gene gun.

[0284] In a more preferred embodiment of the present disclosure, the delivery vehicle is a lipid nanoparticle, which comprises the guide polynucleotide and the mRNA encoding the Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate.

[0285] On the other hand, a technical solution provided by the present disclosure is: a cell, comprising the Cas12 protein as described in the present disclosure, the guiding polynucleotide as described in the present disclosure, the Cas12 inactivated variant as described in the present disclosure, the fusion protein or conjugate as described in the present disclosure, the nucleic acid as described in the present disclosure, the CRISPR-Cas12 system as described in the present disclosure, or the vector system as described in the present disclosure.

[0286] In some embodiments of the disclosure, the cell is a prokaryotic cell.

[0287] In some embodiments of the disclosure, the cell is a eukaryotic cell.

[0288] In some embodiments of the disclosure, the eukaryotic cell is a mammalian cell.

[0289] On the other hand, a technical solution provided by the present disclosure is: a pharmaceutical composition, comprising the Cas12 protein as described in the present disclosure, the guiding polynucleotide as described in the present disclosure, the Cas12 inactivated variant as described in the present disclosure, the fusion protein or conjugate as described in the present disclosure, the nucleic acid as described in the present disclosure, the CRISPR-Cas12 system as described in the present disclosure, the vector system as described in the present disclosure, the delivery system as described in the present disclosure, or the cell as described in the present disclosure.

[0290] Preferably, the pharmaceutical composition comprises pharmaceutically acceptable excipients.

[0291] On the other hand, a technical solution provided by the present disclosure is: a kit comprising the Cas12 protein as described herein, the guiding polynucleotide as described herein, the Cas12 inactivated variant as described herein, the fusion protein or conjugate as described herein, the nucleic acid as described herein, the CRISPR-Cas12 system as described herein, the vector system as described herein, the delivery system as described herein, or the cell as described herein.

[0292] In a preferred embodiment of the present disclosure, the kit further comprises a cutting buffer. The cutting buffer can be any buffer known in the art suitable for Cas12 protein cutting target nucleic acid.

[0293] On the other hand, a technical solution provided by the present disclosure is: use of the Cas12 protein as disclosed herein, the guiding polynucleotide as disclosed herein, the inactivated Cas12 variant as disclosed herein, the fusion protein or conjugate as disclosed herein, the nucleic acid as disclosed herein, the CRISPR-Cas12 system as disclosed herein, the vector system as disclosed herein, the delivery system as disclosed herein, the cell as disclosed herein, the pharmaceutical composition as disclosed herein, or the kit as disclosed herein in the preparation of a reagent or drug for diagnosing, treating and / or preventing a disease or condition associated with a target nucleic acid.

[0294] In specific embodiments of the present disclosure, the disease or condition is a hematological disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease; and / or the agent or drug is used to: cut one or more target nucleic acid molecules or cause a nick in one or more target nucleic acid molecules, activate or upregulate the expression of one or more target nucleic acid molecules, activate or inhibit the transcription of one or more target nucleic acid molecules, inactivate one or more target nucleic acid molecules, visualize, label, or detect one or more target nucleic acid molecules, bind to one or more target nucleic acid molecules, transport one or more target nucleic acid molecules, and mask one or more target nucleic acid molecules.

[0295] In specific embodiments of the present disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or disorder is a disease or disorder listed in Table 27.

[0296] In specific embodiments of the present disclosure, the disease or disorder is selected from the group consisting of hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidosis, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, alpha-1-antitrypsin deficiency, alpha-mannosidosis, alpha-thalassemia, beta-thalassemia, Alzheimer's disease, Budd-Bieder syndrome, albicans punctata, leukocyte adhesion deficiency type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti's lens dystrophy, Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult glucan body disease, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophosphatasia, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin lymphoma, non-muscle invasive bladder cancer, non-alcoholic Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, Charcot-Marie-Tooth disease type 1A, Charcot-Marie-Tooth disease type 2A, pulmonary hypertension, Friedreich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type 1, coronary artery disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge urinary incontinence, acute intermittent urination Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial Tay-Sachs disease, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, junctional epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, Canavan disease, cocaine addiction, Krabbe disease,Crigler-Najjar syndrome, oral cancer, happy puppet syndrome, diffuse intrinsic pontine glioma, Lafora disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, anemia of chronic kidney disease, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Netherton syndrome, ornithine transcarbamylase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Rett syndrome, Triple-negative breast cancer, Sandhoff disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscinosis, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibulopathy, Stargardt's disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, head and neck squamous cell carcinoma, Wilson's disease, stable angina, Ussher syndrome, choroideremia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, novel coronavirus infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, critical limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive dystrophic epidermolysis bullosa, infantile malignant osteopetrosis , dystrophic epidermolysis bullosa, morphea, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic disease, lipoprotein lipase deficiency, severe congenital neutropenia, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, transthyretin (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.

[0297] In specific embodiments of the present disclosure, the transthyretin (ATTR) amyloidosis-related genes include but are not limited to ATTR;

[0298] The genes related to Leber hereditary optic neuropathy include but are not limited to MT-ND4;

[0299] The AATD liver disease related genes include but are not limited to AATD;

[0300] The AATD lung disease related genes include but are not limited to AATD;

[0301] The graft-versus-host disease related genes include but are not limited to thymidine kinase gene;

[0302] The genes related to hereditary retinal dystrophy include but are not limited to RPE65;

[0303] The spinal muscular atrophy-related genes include but are not limited to SMN1;

[0304] The osteoarthritis related genes include but are not limited to TGF-β1;

[0305] The related genes of hemophilia A include but are not limited to factor VIII;

[0306] The related genes of hemophilia B include but are not limited to factor IX;

[0307] The cystic fibrosis related genes include but are not limited to CFTR;

[0308] The Parkinson's disease-related genes include but are not limited to Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB and REST;

[0309] The Usher syndrome related genes include but are not limited to USH2A;

[0310] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include but are not limited to BCL11A, HBG, HBA, and HBB;

[0311] The pulmonary hypertension related genes include but are not limited to eNOS;

[0312] The Stargardt disease-related genes include but are not limited to ABCA4;

[0313] The age-related macular degeneration-related genes include but are not limited to VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang 2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene and FGF2;

[0314] The glaucoma-related genes include but are not limited to AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4 and carbonic anhydrase CA12;

[0315] The idiopathic pulmonary fibrosis related genes include but are not limited to CTGF;

[0316] The hyperlipidemia-related genes include but are not limited to PCSK9;

[0317] The Alzheimer's disease related genes include but are not limited to NGF;

[0318] The coronary heart disease related genes include but are not limited to VEGFA and bFGF;

[0319] The related genes of chronic kidney disease anemia include but are not limited to EPO;

[0320] The related genes of congenital amaurosis include but are not limited to RPE65;

[0321] The retinitis pigmentosa related genes include but are not limited to PDE6B;

[0322] The phenylketonuria related genes include but are not limited to PAH; and / or

[0323] The epilepsy-related genes include but are not limited to GAT1.

[0324] On the other hand, a technical solution provided by the present disclosure is: a method for detecting, binding or cutting a target nucleic acid, the method comprising contacting the target nucleic acid with the Cas12 protein as described herein, the guiding polynucleotide as described herein, the Cas12 inactivated variant as described herein, the fusion protein or conjugate as described herein, the nucleic acid as described herein, the CRISPR-Cas12 system as described herein, the vector system as described herein, the delivery system as described herein, the cell as described herein, the pharmaceutical composition as described herein or the kit as described herein.

[0325] In a preferred embodiment of the present disclosure, the method is a method for non-diagnostic and / or therapeutic purposes; and / or the fusion protein or conjugate comprises a detectable label, such as a label detectable by fluorescence, Southern blotting or FISH.

[0326] In a more preferred embodiment of the present disclosure, when the method is for cutting the target nucleic acid, the method further comprises using a cutting buffer (Cut Buffer) to carry out a cutting reaction. The cutting buffer can be any buffer suitable for Cas12 protein cutting target nucleic acid known in the art.

[0327] On the other hand, a technical solution provided by the present disclosure is: a method for changing the state of a cell, the method comprising using the Cas12 protein as described herein, the guiding polynucleotide as described herein, the Cas12 inactivated variant as described herein, the fusion protein or conjugate as described herein, the nucleic acid as described herein, the CRISPR-Cas12 system as described herein, the vector system as described herein, the delivery system as described herein, the cell as described herein, the pharmaceutical composition as described herein, or the kit as described herein to contact the cell, thereby changing the state of the cell.

[0328] In some embodiments of the present disclosure, the method results in one or more of: increased or decreased expression of a specific gene, induction of cellular senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, cell growth promotion and / or cell growth inhibition in vitro or in vivo, induction of anergy in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo.

[0329] In a more preferred embodiment of the present disclosure, the method is a method for non-diagnostic and / or therapeutic purposes.

[0330] On the other hand, a technical solution provided by the present disclosure is: a method for diagnosing, treating or preventing a disease or condition associated with a target nucleic acid, comprising administering a Cas12 protein as described herein, a guiding polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein to a sample of a subject in need or to a subject in need.

[0331] In specific embodiments of the present disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or disorder is a disease or disorder listed in Table 27.

[0332] In specific embodiments of the present disclosure, the disease or disorder is a hematological disease or disorder, an ophthalmic disease or disorder, a nervous system disease or disorder, a respiratory system disease or disorder, a liver disease or disorder, a metabolic system disease or disorder, cancer, or an infectious disease.

[0333] On the other hand, a technical solution provided by the present disclosure is: the Cas12 protein as described in the present disclosure, the guiding polynucleotide as described in the present disclosure, the Cas12 inactivated variant as described in the present disclosure, the fusion protein or conjugate as described in the present disclosure, the nucleic acid as described in the present disclosure, the CRISPR-Cas12 system as described in the present disclosure, the vector system as described in the present disclosure, the delivery system as described in the present disclosure, the cell as described in the present disclosure, the pharmaceutical composition as described in the present disclosure, or the kit as described in the present disclosure, which is used to diagnose, treat or prevent diseases or disorders associated with target nucleic acids.

[0334] In specific embodiments of the present disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or disorder is a disease or disorder listed in Table 27.

[0335] In specific embodiments of the present disclosure, the disease or disorder is a hematological disease or disorder, an ophthalmic disease or disorder, a nervous system disease or disorder, a respiratory system disease or disorder, a liver disease or disorder, a metabolic system disease or disorder, cancer, or an infectious disease.

[0336] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0337] Figure 1A shows the evolutionary relationship between the disclosed Cas proteins and known Cas12 subtype proteins (sequence alignment was performed using MAFFT, and then an evolutionary tree was constructed using FastTree).

[0338] In the evolutionary tree, some of the proteins disclosed herein form an independent, clearly separate branch (different clusters [CLUSTER]) compared with known Cas12 subtype proteins, that is, they are not mixed with known Cas12 subtypes; and the evalue of these proteins disclosed herein (Figure 1B) compared with the existing Cas12 HMM Profile model is greater than 1e-5. This part of the proteins disclosed herein includes the CLUSTER1-CLUSTER13 proteins shown in Figure 1. Taken together, this suggests that these Cas proteins may be new subgroups, for example, they may be new Cas12 subtypes.

[0339] FIG2 shows the SDS-PAGE electrophoresis diagram of the C12-279 recombinant protein.

[0340] Figure 3 shows a PAM library of Cas12 protein cleavage in vitro. The sequences in the figure are shown in SEQ ID NO: 878-881.

[0341] Figure 4 shows the motif captured after C12-279-sgRNA targeted the 7nt random sequence.

[0342] FIG5 shows the motif captured after C12-279-sgRNA-Rev targeted a 7 nt random sequence.

[0343] Figure 6 shows a fragment of a plasmid containing a 7nt random sequence in a plasmid elimination experiment. The sequences in the figure are shown as SEQ ID NOs: 720, 882-886.

[0344] FIG7 shows the motifs captured in the plasmid curative experiment of C12-279.

[0345] Figure 8 shows some indels generated after C12-279 editing of the TTR gene. The sequences in the figure are shown as SEQ ID NOs: 722, 887-916.

[0346] Figure 9 shows that the CI1062732 (SEQ ID NO: 46) protein, in which dozens of amino acid residues are deleted from the N-terminus of the C12-279 protein (SEQ ID NO: 696), is combined with a gRNA containing the DR sequence GTAATGCGTCTCCCATTGACGCC (SEQ ID NO: 529) to target a 7nt random sequence plasmid library in bacteria, and the captured "fake" PAM motif is 5'-TTNC-3'.

[0347] Figure 10 shows the secondary structure of the "fake" DR sequence analyzed by the inventors. It is speculated that the 3'-terminal C base of the captured "fake" PAM motif TTNC may be caused by an extra C at the 3' end of the DR. The sequences in the figure are shown in SEQ ID NOs: 917 and 918.

[0348] FIG11 shows the SDS-PAGE electrophoresis diagram of the C12-101-07 recombinant protein (126 KDa).

[0349] FIG12 shows the motif captured after C12-101-07-sgRNA targets a 7 nt random sequence.

[0350] FIG13 shows the motif captured after C12-101-07-sgRNA-Rev targeted a 7 nt random sequence.

[0351] FIG14 shows the motifs captured in the C12-101-09 plasmid curing experiment.

[0352] Figure 15A shows a fragment of the pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid. The sequences in the figure are shown as SEQ ID NOs: 919-921. Figure 15B shows a fragment of the library plasmid Plasmid Lib with a PAM sequence composition of NAAN. The sequences in the figure are shown as SEQ ID NOs: 922-924.

[0353] FIG16 shows the editing efficiency of different C12-279 mutants in NAAN cells, expressed as fold compared to the editing efficiency of C12-279.

[0354] FIG17 shows the editing efficiency of different C12-279 mutants in NAAN cells, expressed as fold compared to the editing efficiency of C12-279.

[0355] Figure 18 shows the editing efficiency test results of different mutant targeting reporter systems, where Wt represents wild type C12-279 (green dot); the marked number n indicates that the amino acid residue at position n is mutated to arginine R; Mut-01 is mutant C12-279-pCDH-05 (Q186R variant), Mut-02 is mutant C12-279-pCDH-28 (double point mutations of Q186R and D352R), and Mut-03 is mutant C12-279-pCDH-35 (triple point mutations of G184R, Q186R, and D352R); the green line (dashed line below) indicates that the editing efficiency is 50% higher than that of the wild type; the red line (dashed line above) indicates that the efficiency is the same as that of Mut-02.

[0356] Figure 19 shows the editing efficiency test results of different multi-point mutants in a reporter system. Mut-02-1-5-426-858-860 represents a multi-point mutant obtained by additionally introducing mutations at positions 1, 5, 426, 858, and 860 (all mutated to arginine R) based on the Mut-02 mutant. That is, this multi-point mutant contains the following mutations: amino acid residues 1, 5, 186, 352, 426, 858, and 860 are all mutated to R. The same applies to other mutants herein.

[0357] Figure 20A shows the efficiency of editing the TTR gene after combining multiple point mutants with different gRNAs. Figure 20B shows the efficiency of editing the HBG gene after combining multiple point mutants with different gRNAs. Mut-02-1-5-426-858-860 represents a multiple point mutant obtained by additionally introducing mutations at positions 1, 5, 426, 858, and 860 (all mutated to arginine R) based on the Mut-02 mutant. That is, the multiple point mutant contains the following mutations: amino acid residues at positions 1, 5, 186, 352, 426, 858, and 860 are all mutated to R. The same applies to other mutants.

[0358] Figure 21 shows the results of the dCas12-279 editing activity test. Mut-02-426-860 represents a multiple-point mutant obtained by introducing the 426R and 860R mutations into the Mut-02 mutant. Mut-02-426-860-D651A represents a multiple-point mutant obtained by introducing the 426R, 860R, and 651A mutations into the Mut-02 mutant. The same applies to other mutants herein.

[0359] Figure 22 shows the NGS sequencing results of HEK293 cells electroporated with the encoding mRNA of the mutant Mut-02-1-5-426-858-860 and the modified gRNA (C279-dmTTR01-02) to edit the TTR gene. The measured editing efficiency was as high as 92.18%. The sequences in the figure are shown in SEQ ID NOs: 925-974.

[0360] FIG. 23 shows the PAM recognized by the C12-279 mutant Mut-02-1-426-846-858-860.

[0361] FIG24 shows a schematic diagram of the proposed structure of C12-279. DETAILED DESCRIPTION

[0362] In this disclosure, unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, procedures in molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are conventional procedures widely used in the relevant fields. To facilitate a better understanding of this disclosure, definitions and explanations of relevant terms are provided below.

[0363] In this disclosure, "several" refers to greater than or equal to 2. In this disclosure, "multiple" refers to greater than or equal to 2.

[0364] In the present disclosure, "cleavage" may refer to cutting the main chain of the polynucleotide chain, depending on the context; non-limiting examples include completely breaking single-stranded DNA, breaking one of the single strands of double-stranded DNA, or breaking both single strands of double-stranded DNA.

[0365] In the present disclosure, according to the context, "modification" may refer to other forms of nucleic acid chain chemical reactions other than "cutting", including but not limited to base substitution, addition and / or deletion, and methylation and demethylation of nucleic acid chains. Non-limiting examples include, for example, single-base editing (the Cas12 disclosed herein is fused to a deaminase domain and combined with gRNA) to replace bases on the target nucleic acid chain, such as A→G, C→T, T→C or G→A nucleotide mutations, and other types of nucleotide mutations (such as A→T, C→G, T→A, G→C, etc.); also, for example, Prime editing technology (the Cas12 disclosed herein is fused to a reverse transcriptase and combined with PegRNA) to achieve base substitution, addition or deletion, or base substitution, addition or deletion by means of HDR homologous recombination (such as the Cas12 disclosed herein combined with gRNA and a donor template); Cas12 disclosed herein can also be fused with a DNA methylase or a DNA demethylase and targeted with gRNA.

[0366] In the present disclosure, depending on the context, "regulating the expression of a target nucleic acid" may refer to regulating the transcription of a target nucleic acid; a non-limiting example is the enhancement or inhibition of the transcription of a target nucleic acid by means of a transcriptional activation or repression domain fused to Cas12 through CRISPRa and CRISPRi technologies.

[0367] In the present disclosure, the letters in the amino acid sequence represent the single-letter abbreviations of amino acids well known in the art, such as those described in J. Biol. Chem, 243, p3558 (1968): Alanine: Ala-A, Arginine: Arg-R, Aspartic acid: Asp-D, Cysteine: Cys-C, Glutamine: Gln-Q, Glutamic acid: Glu-E, Histidine: His-H, Glycine: Gly-G, Asparagine: Asn-N, Tyrosine: Tyr-Y, Proline: Pro-P, Serine: Ser-S, Methionine: Met-M, Lysine: Lys-K, Valine: Val-V, Isoleucine: Ile-I, Phenylalanine: Phe-F, Leucine: Leu-L, Tryptophan: Trp-W, Threonine: Thr-T.

[0368] In the present disclosure, "amino acid difference" refers to the difference in amino acid residues at specific sites in the amino acid sequence of a protein, including substitution, addition or reduction.

[0369] In the present disclosure, the mutation of an amino acid residue at a specific site refers to the substitution, addition or reduction of the amino acid residue at that site.

[0370] It is well known to those skilled in the art that in a protein or peptide, two adjacent amino acids each remove an OH or H, dehydrate and condense to form a peptide bond, and each amino acid actually exists in the form of an amino acid residue. Therefore, in this disclosure, the terms "amino acid" and "amino acid residue" generally represent the same meaning. In addition, in order to simplify the expression, the amino acid residue before substitution is retained before the site where the amino acid residue is located in this disclosure, the letter before the site represents the original amino acid residue, and the letter after the site represents the amino acid residue after substitution. For example, S211 represents that the original amino acid residue at site 211 is S. When it is replaced by R, it can be expressed as S211R.

[0371] In this article, the symbol "+" is sometimes used to connect one amino acid mutation before and after it, indicating that these two point mutations exist simultaneously in one mutant; if multiple point mutations are connected by two or more symbols "+", it is intended to indicate that these multiple point mutations exist simultaneously.

[0372] In this disclosure, if an amino acid is substituted, it means that it is replaced by another amino acid residue that is different from the original amino acid residue. If the original amino acid residue is a positively charged amino acid and it is replaced by a positively charged amino acid, it means that it is replaced by another positively charged amino acid residue that is different from the original amino acid residue. For example, if the original amino acid residue is R and it is replaced by a positively charged amino acid, it means that it is replaced by H or K.

[0373] In this article, when referring to RNA sequences, "T" in the sequence can be used interchangeably with "U". When referring to "guide sequences", "T" in the sequence can be used interchangeably with "U". When referring to "direct repeat sequences", "T" in the sequence can be used interchangeably with "U". When referring to "direct repeat sequences", "T" in the sequence can be used interchangeably with "U".

[0374] In this article, when referring to Cas protein numbering, C12-n and Cas12-n are intended to refer to the same protein. For example, Cas12-279 and C12-279 can be used interchangeably.

[0375] Sequence identity

[0376] As used herein, the term "identity" is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. "Identity," "percent identity," and "sequence identity" are used interchangeably. When a certain position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a certain position in each of the two DNA molecules is occupied by adenine, or a certain position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions being compared × 100%. For example, if 6 out of 10 positions of two sequences match, then the two sequences have 60% sequence identity. Typically, comparison is made when two sequences are aligned to produce maximum sequence identity. Such comparison can be by using disclosed and commercially available comparison algorithms and programs, such as, but not limited to, CLUSTERalΩ, MAFFT, Probcons, T-Coffee, Probalign, BLAST, which can be reasonably selected for use by those skilled in the art. Those skilled in the art can determine suitable parameters for aligning sequences, including, for example, any algorithm required for achieving better alignment or optimal comparison over the full length of the compared sequences, and any algorithm required for achieving better alignment or optimal comparison over a portion of the compared sequences.

[0377] CRISPR-Cas12 system

[0378] As used herein, the terms "Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-CRISPR-Associated (Cas) (CRISPR-Cas) System" or "CRISPR System" are used interchangeably and have the meaning generally understood by those skilled in the art, which generally include transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of directing the activity of the Cas genes. Such transcripts or other elements may include sequences encoding Cas effector proteins and guide polynucleotides.

[0379] Zhang Feng's group discovered Cas12a in 2015, classifying it as type V within the Class I CRISPR-Cas system. Following detailed studies of subtype V (Cas12a), Zhang Feng's group also reported Cas12b (C2C1) in 2015. In 2017, Burstein et al. reported the Cas12e (CasX) nuclease. In 2019, Winston X. Yan et al. detailed the newly discovered type V Cas effector proteins, Cas12c, Cas12h, Cas12i, and Cas12g, through bioinformatics analysis.

[0380] In some embodiments, the Cas12 protein described herein refers to an amino acid sequence comprising or having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% sequence identity compared to any one of SEQ ID NOs: 1-53, 696, 728. When the CRISPR-Cas12 system comprises a fusion protein or conjugate comprising the Cas12 protein and a functional domain, the percentage of sequence identity between the Cas12 portion of the fusion protein or conjugate and the reference sequence is calculated.

[0381] In the present disclosure, the CRISPR-Cas12 system comprises a Cas12 protein having at least 50% sequence identity compared to any one of SEQ ID NOs: 1-53, 696, and 728, or a nucleic acid encoding the Cas12 protein; and a guide polynucleotide or a nucleic acid encoding the guide polynucleotide; the guide polynucleotide comprises a direct repeat sequence connected to a guide sequence, the guide sequence being engineered to hybridize with a target nucleic acid, and the guide polynucleotide is capable of forming a complex with the Cas12 protein and guiding the complex to sequence-specific binding to the target nucleic acid.

[0382] Guide polynucleotide

[0383] As used herein, the term "guide polynucleotide" is used to refer to a molecule in the CRISPR-Cas system that forms a complex with the Cas protein and guides the complex to the target sequence. Typically, the guide polynucleotide comprises a backbone sequence connected to the guide sequence, and the guide sequence can hybridize with the target sequence. The backbone sequence typically comprises a direct repeat sequence and sometimes may further comprise a tracrRNA sequence. In some embodiments, the guide polynucleotide does not comprise a tracrRNA sequence. In some embodiments, the guide polynucleotide comprises a tracrRNA sequence.

[0384] In some embodiments, the guide polynucleotide of the CRISPR-Cas12 system is a guide RNA. In some embodiments, the guide polynucleotide is a chemically modified guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least one chemically modified nucleotide.

[0385] [Corrected 03.01.2025 according to Rule 91] In some embodiments, the chemically modified nucleotides include base-modified nucleotides, phosphate-modified nucleotides and ribose-modified nucleotides.

[0386] In some preferred embodiments, the base-modified nucleotides are selected from nucleotides of non-natural bases.

[0387] In some preferred embodiments, the phosphate-modified nucleotides are selected from phosphoramidate nucleotides, phosphorothioate-based nucleotides, phosphorodithioate-based nucleotides, methylphosphonate-based nucleotides, 5'-phosphate-based nucleotides, alkylphosphate-based nucleotides and boranephosphate-based nucleotides.

[0388] [Corrected 03.01.2025 according to Rule 91] In some preferred embodiments, the ribose-modified nucleotides are selected from deoxynucleotides, 3'-terminal deoxythymidine (dT) nucleotides, 2'-O-methyl-modified nucleotides, 2'-fluoro-modified nucleotides, 2'-deoxy-modified nucleotides, 2'-amino-modified nucleotides, 2'-O-allyl-modified nucleotides, 2'-C-alkyl-modified nucleotides, 2'-hydroxyl-modified nucleotides, 2'-methoxyethyl-modified nucleotides, 2'-O-alkyl-modified nucleotides and morpholino nucleotides.

[0389] In some preferred embodiments, the guide nucleotides comprise chemically modified nucleotides at the 5' end and / or the 3' end.

[0390] In some embodiments, the guide nucleotide comprises a deoxynucleotide at the 5' end; the deoxynucleotide is 10-25 nt in length, eg, 14 or 25 nt.

[0391] In some embodiments, the guide nucleotide comprises a phosphorothioate group at the 5' end or the 3' end and is a 2'-OMe modified nucleotide.

[0392] In some embodiments, the guide polynucleotide comprises at least one guide sequence (also known as a spacer sequence) linked to at least one direct repeat (DR). In some embodiments, the guide sequence is located at the 3' end of the direct repeat. In some embodiments, the guide sequence is located at the 5' end of the direct repeat.

[0393] In some embodiments, the tracrRNA sequence is linked to the direct repeat sequence.

[0394] In some embodiments, the tracrRNA sequence is located at the 5' or 3' end of the direct repeat sequence. In some embodiments, the tracrRNA sequence is located at the 5' end of the direct repeat sequence. In some embodiments, the tracrRNA sequence is located at the 3' end of the direct repeat sequence.

[0395] In some embodiments, the nucleotide sequence of the guide polynucleotide comprises, from 5' to 3', tracrRNA, a direct repeat sequence, and a guide sequence.

[0396] In some embodiments, the nucleotide sequence of the guide polynucleotide comprises, from 5' to 3', tracrRNA, a linker sequence, a direct repeat sequence, and a guide sequence.

[0397] In some embodiments, the nucleotide sequence of the guide polynucleotide comprises, from 5' to 3', tracrRNA, a loop sequence, a direct repeat sequence, and a guide sequence.

[0398] In some embodiments, the structure of the guide polynucleotide is 5'-tracrRNA-loop-direct repeat sequence-guide sequence-3'.

[0399] In some embodiments, the tracrRNA and direct repeat sequence of the guide polynucleotide are linked by a nucleotide sequence.

[0400] In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence by a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence by a nucleotide sequence consisting of 4 nucleotides. In a specific embodiment of the present disclosure, the tracrRNA sequence is connected to the direct repeat sequence by a 5'-GAAA-3' sequence.

[0401] In some embodiments, the guide sequence has sufficient complementarity to the target nucleic acid sequence to hybridize to the target nucleic acid and guide sequence-specific binding of the CRISPR-Cas12 complex to the target nucleic acid. In some embodiments, the guide sequence has 100% complementarity to the target nucleic acid, but the guide sequence can have less than 100% complementarity to the target nucleic acid, such as at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% complementarity.

[0402] In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with no more than two nucleotide mismatches. In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with no more than one nucleotide mismatches. In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with or without mismatches.

[0403] In some specific embodiments of the present disclosure, the guide sequence is as shown in any one of SEQ ID NOs: 783-805, 775, 825-877, having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity. In some specific embodiments of the present disclosure, the guide sequence is as shown in any one of SEQ ID NOs: 783-805, 775, 825-877.

[0404] In some embodiments, the CRISPR-Cas12 system comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different guide polynucleotides. In some embodiments, the guide polynucleotides target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different target nucleic acid molecules, or target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different regions of one or more target nucleic acid molecules.

[0405] In some embodiments, the guidance polynucleotide includes a constant direct repeat sequence located upstream of the variable guide sequence. In some embodiments, a plurality of guidance polynucleotides are a part of an array (which can be a part of a vector, such as a viral vector or a plasmid). For example, the guidance array including the sequence DR-spacer-DR-spacer-DR-spacer-......-DR-spacer can include a plurality of unique unprocessed guidance polynucleotides (one for each DR-spacer or spacer-DR sequence). Once introduced into a cell or a cell-free system, the array is processed into several separate mature guidance polynucleotides by the Cas12 protein. This allows multiplexing, such as delivering a plurality of guidance polynucleotides to a cell or system to target a plurality of target nucleic acids or a plurality of regions within a single target nucleic acid.

[0406] The ability of a guide polynucleotide to guide a complex (CRISPR complex) to bind sequence-specifically to a target nucleic acid can be assessed by any suitable assay. For example, components of a CRISPR system sufficient to form a complex (CRISPR complex), including the guide polynucleotide to be tested, can be provided to a host cell with the corresponding target nucleic acid molecule, such as by transfection of a vector encoding the components of the CRISPR complex, and preferential cleavage within the target sequence can be assessed. Similarly, cleavage of a target nucleic acid sequence can be assessed in a test tube by providing a target nucleic acid, components of a CRISPR complex, including the guide polynucleotide to be tested and a control guide polynucleotide that is different from the test guide polynucleotide, and comparing the ability to bind to the target nucleic acid or the rate at which the target nucleic acid is cleaved between the test and control guide polynucleotides. The ability of a CRISPR complex to cleave a target nucleic acid or target nucleic acid can also be assessed by the assays described above.

[0407] Cas12 mutants

[0408] As described herein, when referring to "corresponding positions of the sequence set forth in SEQ ID NO: XX" or using similar text descriptions, the corresponding positions can be determined by amino acid sequence alignment. Typically, comparison is performed when two sequences are aligned to produce maximum sequence identity. Such an alignment can be performed using published and commercially available alignment algorithms and programs, such as, but not limited to, ClustalΩ, MAFFT, Probcons, T-Coffee, Probalign, and BLAST, which can be reasonably selected for use by a person skilled in the art. A person skilled in the art can determine appropriate parameters for aligning sequences, including, for example, any algorithm required to achieve optimal alignment or optimal comparison over the entire length of the compared sequences, as well as any algorithm required to achieve optimal alignment or optimal comparison over a portion of the compared sequences.

[0409] Optionally, the corresponding position is aligned online (https: / / mafft.cbrc.jp / alignment / server / index.html) by aligning the amino acids of the Cas12 protein with any of the sequences shown in SEQ ID NO: 1-53, 696, 728 using the MAFFT version 7 tool, selecting the following parameters: G-INS-i (Very slow; recommended for <200 sequences with global homology; 2 iterative cycles only), Try to align gappy regions anyway, Scoring matrix for amino acid sequences--BLOSUM62, Gap opening penalty 1.53, Offset value 0.0, Mafft-homologs--Use UniRef50 (more comprehensive and requires longer search time).

[0410] Optionally, the corresponding position is aligned online (https: / / mafft.cbrc.jp / alignment / server / index.html) by aligning the amino acids of the Cas12 protein with the sequence shown in SEQ ID NO: 696 using the MAFFT version 7 tool, selecting the following parameters: G-INS-i (Very slow; recommended for <200 sequences with global homology; 2 iterative cycles only), Try to align gappy regions anyway, Scoring matrix for amino acid sequences--BLOSUM62, Gap opening penalty 1.53, Offset value 0.0, Mafft-homologs--Use UniRef50 (more comprehensive and requires longer search time).

[0411] In some embodiments, the Cas12 protein provided herein comprises one or more mutations, such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof, compared to the Cas12 protein shown in any one of SEQ ID NOs: 1-53, 696, and 728. In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 4, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 128, 129 or 130 amino acid changes (e.g., insertions, deletions or substitutions) in the guide polynucleotide, but retain the ability to bind to a target nucleic acid molecule that is complementary to the guide sequence of the guide polynucleotide and / or retain the ability to process an RNA transcript comprising the guide sequence into a guide polynucleotide molecule.In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191,

[0412] One type of modification or mutation includes replacing amino acid residues with similar biochemical properties with amino acids, i.e., conservative substitutions. Typically, conservative substitutions have little or no effect on the activity of the resulting protein or peptide. For example, conservative substitutions are amino acid substitutions in the Cas12 protein that do not substantially affect the binding of the Cas12 protein to the target nucleic acid molecule complementary to the gRNA molecule guide sequence, and / or the process of processing the guide array RNA transcript into a gRNA molecule.

[0413] More substantial changes can be made by using less conservative substitutions, for example, by selecting residues that differ more in maintaining: (a) the structure of the polypeptide backbone in the region where the substitution occurs, for example, as a helical or sheet conformation; (b) the charge or hydrophobicity of the region that interacts with the target site; or (c) the bulk of the side chain. Substitutions that would generally be expected to produce the greatest changes in polypeptide function are (a) substitutions between a hydrophilic residue (e.g., serine or threonine) and a hydrophobic residue (e.g., leucine, isoleucine, phenylalanine, valine, or alanine); (b) substitutions between cysteine ​​or proline and any other residue; (c) substitutions between a residue with a positively charged side chain (e.g., lysine, arginine, or histidine) and a negatively charged residue (e.g., glutamic acid or aspartic acid); or (d) substitutions between a residue with a bulky side chain (e.g., phenylalanine) and a residue without a side chain (e.g., glycine).

[0414] Cas12 active fragment

[0415] In the present disclosure, the Cas12 protein may comprise only a WED-I domain, a Helical-I1 domain, a PI domain, a Helical-I2 domain, a Helical-II domain, a WED-II domain, a Ruvc-I domain, a Helical-III domain, a BH domain, a Ruvc-II domain, a Nuc domain and / or a Ruvc-III domain.

[0416] The Cas12 protein described in the present disclosure, in addition to comprising the structural domain, may also comprise the structural domains of other Cas12 proteins in the prior art, which are combined together to form the complete structure of the Cas12 protein to achieve the functions of the Cas12 protein described in the present disclosure, including but not limited to retaining the ability of the Cas12 protein to form a complex with the gRNA, retaining the ability of the Cas12 protein to form a complex with the gRNA and target the target nucleic acid, retaining the ability of the complex formed by the Cas12 protein and the gRNA to target and regulate the expression of the target nucleic acid, retaining the ability of the complex formed by the Cas12 protein and the gRNA to target and cut the single or double strand of the target nucleic acid, retaining the ability of the Cas12 protein to bind to the target nucleic acid molecule complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process the RNA transcript comprising the guide sequence into a guide polynucleotide molecule.

[0417] Cas12 inactivated variants

[0418] By inactivating the RuvC domain of Cas12 through point mutation, the Cas12 protein will lose its endonuclease activity. The resulting dCas12 (dead Cas12) can only bind to the target gene under the mediation of the guide polynucleotide, but does not have the function of cutting DNA.

[0419] The RuvC domain of Cas12 can also be partially inactivated by point mutation to form a Cas12 nickase (nCas12), which binds to the target gene under the guidance of a guide polynucleotide and cuts one of the single strands in the double-stranded nucleic acid without cutting the other single strand.

[0420] Therefore, dCas12 or nCas12 can be fused or conjugated with other domains (including but not limited to deaminase domains, transcription activation domains, transcription repression domains, methylation domains, demethylation domains, histone acetylation domains, histone deacetylation domains), and guided to the target sequence of the target nucleic acid by the guiding polynucleotide, and then the corresponding functions are performed with the help of the other domains; for example, the conversion of the target nucleic acid base C→T is achieved by deaminating the cytosine base, the conversion of the target nucleic acid base A→G is achieved by deaminating the adenine base, the transcription inhibition of the target nucleic acid is achieved by the transcription repression domain KRAB, and the transcription of the target nucleic acid is promoted by the transcription activation domain VP64.

[0421] Functional domain

[0422] In some embodiments, the Cas12 protein or Cas12 inactivated variant is covalently linked or fused to a homologous or heterologous functional domain.

[0423] In some embodiments, the functional domain has an enzymatic activity that modifies a target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristoylation activity and / or demyristoylation activity.

[0424] In some embodiments, the functional domain is selected from one or more of the following: a nuclease (e.g., FokI), a methyltransferase, a demethylase, a DNA repair enzyme, a DNA damaging enzyme, a deaminase, a dismutase, an alkylase, a depurinase, an oxidase, a pyrimidine dimer-forming enzyme, an integrase, a transposase, a recombinase, a polymerase, a ligase, a helicase, a photolyase, a glycosylase, a deglycosylase, an acetyltransferase, a deacetylase, a kinase, a phosphatase, a ubiquitin ligase, a deubiquitinating enzyme, an adenylylase, a deadenylase, a SUMOylating enzyme, a deSUMOylating enzyme, a myristoylase, and / or a demyristoylase.

[0425] In some embodiments, the functional domains are selected from one, two, three, four or more of the following: a subcellular localization signal, a DNA binding domain, a protease domain, a transcription activation domain, a transcription repression domain, a nuclease domain, a deaminase domain, a uracil DNA glycosylase domain (UDG), a uracil DNA glycosylase inhibitory domain (UGI), a methylase, a demethylase, a transcription release factor, a histone acetylase domain, a histone deacetylase domain, a DNA ligase, an epitope tag and / or a reporter domain.

[0426] In some embodiments of the present disclosure, the deaminase domain is selected from: APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, activation-induced cytidine deaminase (AID), CDA from lamprey, a mutant of adenosine deaminase engineered to act on DNA (TadA).

[0427] In some embodiments, the transcriptional activation domain is selected from the group consisting of: P65, VPR, VP16, VP64, VTR1, VTR2, VTR3, p65, MyoD1, HSF1, RTA, SET7 / 9, and histone acetyltransferase. In some embodiments, the transcriptional activation domain is optionally selected from: sequence ETFSDLWKL from p53 TAD1, sequence DDIEQWFTE from p53 TAD2, sequence SDIMDFVLK from MLL, sequence DLLDFSMMF from E2A, sequence ETLDFSLVT from Rtg3, sequence RKILNDLSS from CREB, sequence EAILAELKK from CREBaB6, sequence DDVVQYLNS from Gli3, sequence DDVYNYLFD from Gal4, sequence DLFDYDFLV from Oaf1, sequence DFFDYDLLF from Pip2, sequence EDLYSILWS from Pdr1, sequence TDLYHTLWN from Pdr3.

[0428] In some embodiments, the transcriptional repression domain is selected from the group consisting of KOX1, KAP-1, MAD, FKHR, EGR-1, ERD, SID, a tandem of SID (e.g., SID4X), TIEG, v-ERB-A, MBD2, MBD3, TRa, a histone methyltransferase, a histone deacetylase (HDAC), a nuclear hormone receptor (e.g., an estrogen receptor or a thyroid hormone receptor), a DNMT family member (e.g., DNMT1, DNMT3A, DNMT3B), the KRAB domain of MeCP2, ROM2, and AtHD2A.

[0429] In some embodiments, the transcriptional repression domain is a KRAB domain from the KOX1 protein.

[0430] In some embodiments, the nuclease domain is selected from FokI, a polypeptide having ssDNA cleavage activity, and a polypeptide having dsDNA cleavage activity.

[0431] In some embodiments, the methylase domain is selected from DNA methylases, including but not limited to DNMT1, DNMT3a, and DNMT3b.

[0432] In some embodiments, the demethylase is selected from TET1CD, TET1, ROS1, DME, DML2, and DML3.

[0433] Methylation and demethylation are recognized in the art as important means of epigenetic gene regulation.

[0434] In some embodiments, the homology or heterology functional domain is for dissolving, purifying or detecting useful sequence tags for the fusion rotein or conjugate. Suitable protein tag sequences are provided herein, and these sequences include but are not limited to biotin carboxylase carrier protein (BCCP) label, myc label, calmodulin label, FLAG label, hemagglutinin (HA) label, polyhistidine tag (also referred to as His label), maltose binding protein (MBP) label, nus label, glutathione-S-transferase (GST) label, green fluorescent protein (GFP) label, thioredoxin label, S-label, Softag (for example, Softag 1, Softag 3), strep-label, biotin ligase label, F1AsH label, V5 label and SBP label.Other suitable sequence will be apparent to those of ordinary skill in the art.

[0435] In some embodiments of the present disclosure, a single-base editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with a deaminase domain and a nuclear localization signal. In some embodiments of the present disclosure, a single-base editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with an APOBEC3A domain and an SV40 NLS.

[0436] In some embodiments of the present disclosure, a transcriptional repression epi-editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with KRAB and SV40 NLS. In some embodiments of the present disclosure, a transcriptional activation epi-editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with the VP64 domain and SV40 NLS.

[0437] Subcellular localization signal

[0438] In some embodiments, the Cas12 protein is fused to at least one homologous or heterologous subcellular localization signal. In some embodiments, the Cas12 protein is fused to at least one homologous or heterologous subcellular localization signal. Exemplary subcellular localization signals include organelle localization signals, such as nuclear localization signals (NLS), nuclear export signals (NES), or mitochondrial localization signals.

[0439] Non-limiting examples of NLSs include NLS sequences derived from the SV40 virus large T antigen, which has the amino acid sequence PKKKRKV (SEQ ID NO: 738); an NLS from a nucleoplasmin (e.g., sequence KRPAATKKAGQAKKKK, SEQ ID NO: 739); a c-myc NLS, which has the amino acid sequence PAAKRVKLD (SEQ ID NO: 740) or RQRRNELKRSP (SEQ ID NO: 741); an hRNPA1 M9 NLS, which has the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 742); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 743) from the IBB domain; the sequence VSRKRPRP (SEQ ID NO: 744) and PPKKARED (SEQ ID NO: 745) of myoma T protein. NO: 745); the sequence of human p53 PQPKKKPL (SEQ ID NO: 746); the sequence of mouse c-ablIV SALIKKKKKMAP (SEQ ID NO: 747); the sequences of influenza virus NS1 DRLRR (SEQ ID NO: 748) and PKQKKRK (SEQ ID NO: 749); the sequence of hepatitis virus delta antigen RKLKKKIKKL (SEQ ID NO: 750); the sequence of mouse Mx1 protein REKKKFLKRR (SEQ ID NO: 751); the sequence of human PARP enzyme KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 752); and the sequence of steroid hormone receptor RKCLQAGMNLEARKTKK (SEQ ID NO: 753). In some embodiments, the nuclear localization sequence is of sufficient strength to achieve detectable accumulation of the fusion protein or conjugate described herein in the nucleus of a eukaryotic cell. In general, the intensity of nuclear localization activity can be derived from the number of NLS, one or more specific NLS used or a combination of these factors. Detection of accumulation in the core can be performed by any suitable technique. For example, a detectable marker can be fused to the Cas protein so that the position in the cell can be visualized, such as by combining means (for example, a dye specific to the core, such as DAPI) for detecting the position of the core. The nucleus can also be separated from the cell, and its contents can then be analyzed by any suitable method for detecting protein, such as immunohistochemistry, western blotting or enzyme activity assays.Nuclear accumulation can also be determined indirectly, such as by measuring the effects of nucleic acid-targeting complex formation (e.g., measuring DNA or RNA cleavage or mutation at the target sequence, or measuring gene expression activity that is altered due to the effects of DNA- or RNA-targeting complex formation and / or DNA- or RNA-targeting Cas protein activity), compared to a control that is not exposed to the nucleic acid-targeting Cas protein or nucleic acid-targeting complex, or is exposed to a nucleic acid-targeting Cas protein that lacks one or more NLSs.

[0440] Vector system

[0441] Another aspect of the present disclosure relates to a vector system comprising the CRISPR-Cas12 system described herein, comprising one or more recombinant vectors comprising a polynucleotide sequence encoding the Cas12 protein and a polynucleotide sequence encoding the guide polynucleotide.

[0442] In some embodiments, the vector system comprises at least one plasmid or viral recombinant vector (e.g., retrovirus, lentivirus, adenovirus, adeno-associated virus or herpes simplex virus). In some embodiments, the polynucleotide sequence encoding the Cas12 protein and the polynucleotide sequence encoding the guide polynucleotide are located on the same recombinant vector. In some embodiments, the polynucleotide sequence encoding the Cas12 protein and the polynucleotide sequence encoding the guide polynucleotide are located on multiple recombinant vectors.

[0443] In some embodiments, the polynucleotide sequence encoding Cas12 protein and / or the polynucleotide sequence encoding guide polynucleotide are operably connected to a regulatory sequence (also referred to as a regulatory element). The regulatory element includes a promoter, an enhancer, an internal ribosome entry site (IRES) and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences). Regulatory elements include regulatory elements that make the nucleotide sequence constitutively expressed in many types of host cells, and regulatory elements (e.g., tissue-specific regulatory sequences) that make the nucleotide sequence expressed only in certain host cells. Tissue-specific promoters can be directly expressed mainly in desired tissues of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas) or specific cell types (e.g., lymphocytes). Regulatory elements can also guide expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type-specific. In some embodiments, the regulatory element is an enhancer element, such as WPRE, the CMV enhancer, the R-U5 segment in the LTR of HTLV-1, the SV40 enhancer, or the intronic sequence between exons 2 and 3 of rabbit β-globin.

[0444] In some embodiments, the recombinant vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with RSV enhancer), cytomegalovirus (CMV) promoter (optionally with CMV enhancer), SV40 promoter, dihydrofolate reductase promoter, β-actin promoter, phosphoglycerol kinase (PGK) promoter, or EF1α promoter), or a pol III promoter and a pol II promoter.

[0445] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1α, CAG, and β-actin promoters. In some embodiments, the promoter is an inducible promoter regulated by external signals or molecules (e.g., transcription factors).

[0446] In some embodiments, the promoter is a tissue-specific promoter, which can be used to drive the tissue-specific expression of Cas12 protein. Suitable muscle-specific promoters include but are not limited to CK8, MHCK7, myoglobin promoter (Mb), desmin (Desmin) promoter, muscle creatine kinase promoter (MCK) and variants thereof, and SPc5-12 synthetic promoters. Suitable immune cell-specific promoters include but are not limited to B29 promoter (B cells), CD14 promoter (monocytes), CD43 promoter (leukocytes and platelets), CD68 (macrophages) and SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include but are not limited to CD43 promoter (leukocytes and platelets), CD45 promoter (hematopoietic cells), INF-β (hematopoietic cells), WASP promoter (hematopoietic cells), SV40 / CD43 promoter (leukocytes and platelets), and SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include but are not limited to elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, the Fit-1 promoter and the ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, the GFAP promoter (astroglial cells), the SYN1 promoter (neurons), and the NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, the NphsI promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, the OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, the SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, the SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, α-MHC.

[0447] AAV vectors

[0448] Another aspect of the present disclosure relates to an adeno-associated viral (AAV) vector comprising the CRISPR-Cas12 system described herein, wherein the adeno-associated viral (AAV) vector comprises DNA encoding the Cas12 protein and guide polynucleotide described herein.

[0449] Delivery of the CRISPR-Cas system by AAV vector is described in Maeder et al., Nature Medicine 25:229-233 (2019), which has been clinically demonstrated to be safe and effective for subretinal delivery of AAV. Local delivery by subretinal injection, the natural tropism of AAV5 for photoreceptor cells, and the use of the photoreceptor-specific GRK1 promoter are all used to limit the expression of the CRISPR / Cas system to therapeutic target tissues and cell types, which is incorporated herein by reference in its entirety. In some embodiments, the AAV vector comprises an ssDNA genome comprising a coding sequence for a Cas12 protein and a guide polynucleotide flanked by ITRs.

[0450] In some embodiments, the CRISPR-Cas12 system described herein is packaged in an AAV vector, such as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, and AAVrh74. In some embodiments, the CRISPR-Cas12 system described herein is packaged in an AAV vector comprising an engineered capsid with tissue tropism, such as an engineered muscle tropism capsid. Tabebordbar et al., Cell 184:4919-4938 (2021) describes engineering of AAV capsids with tissue tropism by directed evolution, identifies a class of capsids containing RGD motifs, and systemic injection of MyoAAV can efficiently transduce primate muscles, which is incorporated herein by reference in its entirety.

[0451] lipid nanoparticles

[0452] Another aspect of the present disclosure relates to lipid nanoparticles (LNPs) comprising a CRISPR-Cas12 system described herein, wherein the LNP comprises a guide polynucleotide described herein and an mRNA encoding a Cas12 protein described herein.

[0453] Gillmore et al., N.Engl.J.Med., 385:493-502 (2021) describes the LNP delivery of the CRISPR-Cas system, lipid nanoparticles (LNPs) are composed of 4 lipids, including a proprietary ionizable lipid LP000001; DSPC; cholesterol and DMG-PEG2k, and the LNP suspension is formulated in an aqueous buffer of Tris, NaCl and sucrose, pH 7.4, which is incorporated herein by reference in its entirety. In some embodiments, in addition to RNA payload (Cas12 mRNA and guide polynucleotides), lipid nanoparticles (LNPs) also include four components: cationic or ionizable lipids, cholesterol, helper lipids, and PEG-lipids. In some embodiments, the cation or ionizable lipids include cKK-E12, C12-200, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA, FTT5, Moderna SM-102, and Intellia LP01. In some embodiments, the PEG-lipids include PEG-2000-C-DMG, PEG-2000-DMG, or ALC-0159. In some embodiments, the helper lipids include DSPC. The components of LNPs are described in Paunovska et al., Nature Reviews Genetics 23: 265-280 (2022), and FDA-approved LNPs contain variants of four basic ingredients: cation or ionizable lipids, cholesterol, helper lipids, and polyethylene glycol (PEG) lipids, which are incorporated herein by reference in their entirety.

[0454] Lentiviral vectors

[0455] Another aspect of the present disclosure relates to a lentiviral vector comprising a CRISPR-Cas12 system as described herein, wherein the lentiviral vector comprises a guide polynucleotide as described herein and an mRNA encoding a Cas12 protein as described herein. In some embodiments, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein such as VSV-G. In some embodiments, the mRNA encoding the Cas12 protein is connected to an aptamer sequence.

[0456] RNP complex

[0457] Another aspect of the present disclosure relates to a ribonucleoprotein complex comprising a CRISPR-Cas12 system as described herein, wherein the ribonucleoprotein complex is formed by a guidance polynucleotide and a Cas12 protein as described herein. In some embodiments, the ribonucleoprotein complex can be delivered to eukaryotic cells, mammalian cells, or human cells by microinjection or electroporation. In some embodiments, the ribonucleoprotein complex can be packaged in virus-like particles and delivered to a mammal or human subject in vivo.

[0458] virus-like particles

[0459] Another aspect of the present disclosure relates to a virus-like particle (VLP) comprising the CRISPR-Cas12 system described herein, wherein the virus-like particle comprises a guide polynucleotide and a Cas12 protein described herein, or a ribonucleoprotein complex consisting of the guide polynucleotide and the Cas12 protein.

[0460] Banskota et al. Cell 185(2):250-265(2022) reported the development and application of DNA-free virus-like particles (eVLPs) for efficient packaging and delivery of base editors or Cas9 ribonucleoprotein; Mangeot et al., Nature Communications 10(1):1-15(2019) used engineered mouse leukemia virus-like particles (Nanoblades) loaded with Cas9-sgRNA ribonucleoprotein to induce efficient genome editing in cell lines and primary cells (including human induced pluripotent stem cells, human hematopoietic stem cells and mouse bone marrow cells); Campbell, et al., Molecular Therapy 27:151-163 (2019) utilizes a kind of specialized extracellular vesicle called " gesicle ", effectively but transiently delivers Cas9 targeting HIV long terminal repeat sequence (LTR) in the form of ribonucleoprotein, Gesicles is produced by expressing vesicular stomatitis virus glycoprotein and packaging protein (as its cargo), so there is no need for transgenic delivery, so that Cas9 expression and Mangeot et al.Molecular Therapy, 19 (9): 1656-1666 (2011) reported that the overexpression of the spike glycoprotein of vesicular stomatitis virus (VSV-G) in human cells induced the release of fusogenic vesicles called gesicles, biochemical and functional studies have shown that glial cells bind proteins from production cells and can transport them to receptor cells, and this protein transduction method allows direct transport of cytoplasmic, nuclear or surface proteins in target cells. These documents all describe engineered VLPs, which are incorporated herein by reference in their entirety.

[0461] In some embodiments, engineered virus-like particles (VLPs) are pseudotyped with homologous or heterologous envelope proteins such as VSV-G. In some embodiments, the Cas12 protein is fused to a gag protein (e.g., MLVgag) by a cleavable linker, wherein the cutting of the linker in the target cell exposes the NLS between the linker and the Cas12 protein. In some embodiments, the fusion protein or conjugate comprises (e.g., from 5' to 3') a gag protein (e.g., MLVgag), one or more NES, a cleavable linker, one or more NLS, and Cas12, as described in Banskota et al. Cell 185 (2): 250-265 (2022).

[0462] In some embodiments, the Cas12 protein is fused to a first dimerization domain that is capable of dimerizing or heterodimerizing with a second dimerization domain fused to a membrane protein, wherein the presence of a ligand promotes the dimerization and enriches the Cas12 protein or fusion protein or conjugate into VLPs as described in Campbell, et al., Molecular Therapy 27: 151-163 (2019).

[0463] cell

[0464] Another aspect of the present disclosure relates to cells comprising CRISPR-Cas12 systems as described herein.Cells (for example, which can be used to produce cell-free systems) can be eukaryotic or prokaryotic.The examples of such cells include, but are not limited to, bacteria, archaebacteria, plants, fungi, yeasts, insects, and mammalian cells, such as lactobacilli, lactococci, bacillus (for example, bacillus subtilis), Escherichia (for example, Escherichia coli), Clostridium, Saccharomyces or Pichia (such as saccharomyces cerevisiae or Pichia pastoris), Kluyveromyces lactis, Salmonella typhimurium, Drosophila cells, Caenorhabditis elegans cells, African clawed frog cells, SF9 cells, C129 cells, 293 cells, Neurospora and immortalized mammalian cell lines (for example, Hela cells, myeloid cell lines, and lymphoid cell lines).

[0465] In some embodiments, the cell is a prokaryotic cell, such as a bacterial cell, such as escherichia coli. In some embodiments, the cell is a eukaryotic cell, such as a mammalian cell or a human cell. In some embodiments, the cell is a primary eukaryotic cell, a stem cell, a tumor / cancer cell, a circulating tumor cell (CTC), a blood cell (for example, T cell, B cell, NK cell, Tregs etc.), a hematopoietic stem cell, a specialized immune cell (such as tumor infiltrating lymphocytes or tumor suppressor lymphocytes), a stromal cell (such as cancer associated fibroblasts etc.) in a tumor microenvironment. In some embodiments, the cell is the brain or neuronal cell (for example, neuron, astrocyte, microglia, retinal ganglion cell, rod / cone cell etc.) of a central or peripheral nervous system.

[0466] Target nucleic acid or target DNA

[0467] In some embodiments of the present disclosure, the target nucleic acid is target DNA.

[0468] The CRISPR-Cas12 system described herein can be used to target one or more target nucleic acid molecules, such as target nucleic acid molecules present in a biological sample, an environmental sample (e.g., a soil, air, or water sample), and the like.

[0469] In some embodiments of the present disclosure, the target nucleic acid is a disease or disorder related gene. In some embodiments of the present disclosure, the target nucleic acid is a disease related gene. In some embodiments of the present disclosure, the disease related gene is a pathogenic gene that directly causes the disease. In some embodiments of the present disclosure, the disease related gene is an abnormal gene that directly causes the disease or a gene whose expression is abnormal. For example, an unfavorable mutation occurs in the gene, leading to the occurrence of the disease. For another example, the gene expression is too high or too low, leading to the occurrence of the disease. In some embodiments of the present disclosure, when the gene expression is too high, the occurrence of the disease is caused. In some embodiments of the present disclosure, when the gene expression is too low, the occurrence of the disease is caused. In some embodiments of the present disclosure, the gene expression is too high and is related to the occurrence of the disease. In some embodiments of the present disclosure, the gene expression is too low and is related to the occurrence of the disease.

[0470] In some embodiments of the present disclosure, the disease or disorder is a hematological disease or disorder, an ophthalmic disease or disorder, a nervous system disease or disorder, a respiratory system disease or disorder, a liver disease or disorder, a metabolic system disease or disorder, cancer, or an infectious disease.

[0471] In some embodiments of the present disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or disorder is a disease or disorder listed in Table 27. Table 27 shows the target nucleic acids and the disease or disorder corresponding to each target nucleic acid.

[0472] In some embodiments of the present disclosure, the disease or disorder is selected from the group consisting of hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidosis, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, alpha-1-antitrypsin deficiency, alpha-mannosidosis, alpha-thalassemia, beta-thalassemia, Alzheimer's disease, Budd-Bieder syndrome, albicans punctata, leukocyte adhesion deficiency type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti's lens dystrophy, Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult glucan body disease, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophosphatasia, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin lymphoma, non-muscle invasive bladder cancer, non-alcoholic Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, Charcot-Marie-Tooth disease type 1A, Charcot-Marie-Tooth disease type 2A, pulmonary hypertension, Friedreich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type 1, coronary artery disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge urinary incontinence, acute intermittent urination Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial Tay-Sachs disease, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, junctional epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, Canavan disease, cocaine addiction, Krabbe disease,Crigler-Najjar syndrome, oral cancer, happy puppet syndrome, diffuse intrinsic pontine glioma, Lafora disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, anemia of chronic kidney disease, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Netherton syndrome, ornithine transcarbamylase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Rett syndrome, Triple-negative breast cancer, Sandhoff disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscinosis, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibulopathy, Stargardt's disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, head and neck squamous cell carcinoma, Wilson's disease, stable angina, Ussher syndrome, choroideremia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, novel coronavirus infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, critical limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive dystrophic epidermolysis bullosa, infantile malignant osteopetrosis , dystrophic epidermolysis bullosa, morphea, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic disease, lipoprotein lipase deficiency, severe congenital neutropenia, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, transthyretin (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.

[0473] The genes related to transthyretin (ATTR) amyloidosis include but are not limited to ATTR;

[0474] The genes related to Leber hereditary optic neuropathy include but are not limited to MT-ND4;

[0475] The AATD liver disease related genes include but are not limited to AATD;

[0476] The AATD lung disease related genes include but are not limited to AATD;

[0477] The graft-versus-host disease related genes include but are not limited to thymidine kinase gene;

[0478] The genes related to hereditary retinal dystrophy include but are not limited to RPE65;

[0479] The spinal muscular atrophy-related genes include but are not limited to SMN1;

[0480] The osteoarthritis related genes include but are not limited to TGF-β1;

[0481] The related genes of hemophilia A include but are not limited to factor VIII;

[0482] The related genes of hemophilia B include but are not limited to factor IX;

[0483] The cystic fibrosis related genes include but are not limited to CFTR;

[0484] The Parkinson's disease-related genes include but are not limited to Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB and REST;

[0485] The Usher syndrome related genes include but are not limited to USH2A;

[0486] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include but are not limited to BCL11A, HBG, HBA, and HBB;

[0487] The pulmonary hypertension related genes include but are not limited to eNOS;

[0488] The Stargardt disease-related genes include but are not limited to ABCA4;

[0489] The age-related macular degeneration-related genes include but are not limited to VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang 2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene and FGF2;

[0490] The glaucoma-related genes include but are not limited to AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4 and carbonic anhydrase CA12;

[0491] The idiopathic pulmonary fibrosis related genes include but are not limited to CTGF;

[0492] The hyperlipidemia-related genes include but are not limited to PCSK9;

[0493] The Alzheimer's disease related genes include but are not limited to NGF;

[0494] The coronary heart disease related genes include but are not limited to VEGFA and bFGF;

[0495] The related genes of chronic kidney disease anemia include but are not limited to EPO;

[0496] The related genes of congenital amaurosis include but are not limited to RPE65;

[0497] The retinitis pigmentosa related genes include but are not limited to PDE6B;

[0498] The phenylketonuria related genes include but are not limited to PAH; and / or

[0499] The epilepsy-related genes include but are not limited to GAT1.

[0500] In some specific embodiments of the present disclosure, the sequence of the target nucleic acid is shown in any one of SEQ ID NOs: 761-782.

[0501] Non-limiting examples of such target nucleic acids also include those listed in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed on December 12, 2012 and January 2, 2013, respectively, and International Application No. PCT / US2013 / 074667, filed on December 12, 2013, all of which are incorporated herein by reference.

[0502] In some embodiments, the target nucleic acid is a reporter gene. Examples of reporter genes include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP).

[0503] Applications for treating or preventing diseases

[0504] Another aspect of the present disclosure relates to a pharmaceutical composition comprising a Cas12 protein as described in the present disclosure, a guiding polynucleotide as described in the present disclosure, a Cas12 inactivation variant as described in the present disclosure, a fusion protein or conjugate as described in the present disclosure, a nucleic acid as described in the present disclosure, a CRISPR-Cas12 system as described in the present disclosure, a vector system as described in the present disclosure, a delivery system as described in the present disclosure, or a cell as described in the present disclosure. The pharmaceutical composition may comprise, for example, an AAV vector encoding a Cas12 protein or a Cas12 inactivation variant and a guiding polynucleotide as described herein. The pharmaceutical composition may comprise, for example, lipid nanoparticles comprising a guiding polynucleotide as described herein and an mRNA encoding a Cas12 protein. The pharmaceutical composition may comprise, for example, a lentiviral vector comprising a guiding polynucleotide as described herein and an mRNA encoding a Cas12 protein. The pharmaceutical composition may comprise, for example, a virus-like particle comprising a guiding polynucleotide and a Cas12 protein as described herein or a ribonucleoprotein complex formed by the guiding polynucleotide and the Cas12 protein.

[0505] Another aspect of the present disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein in cutting or editing a target nucleic acid in a mammalian cell.

[0506] Another aspect of the present disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein for any of the following: cleaving one or more target nucleic acid molecules or nicking one or more target nucleic acid molecules, activating or upregulating the expression of one or more target nucleic acid molecules, activating or inhibiting the transcription of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling or detecting one or more target nucleic acid molecules, binding one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules.

[0507] Another aspect of the present disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein to modify one or more target nucleic acid molecules, wherein the modification of the one or more target nucleic acid molecules comprises one or more of the following: nucleic acid base substitution, nucleic acid base deletion, nucleic acid base insertion, fragmentation of the target nucleic acid, nucleic acid methylation, and nucleic acid demethylation.

[0508] Another aspect of the present disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein in the diagnosis, treatment, or prevention of a disease or disorder associated with a target nucleic acid.

[0509] Another aspect of the present disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein in the preparation of a medicament for diagnosing, treating or preventing a disease or disorder associated with a target nucleic acid.

[0510] In some embodiments of the present disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or disorder is a disease or disorder listed in Table 27. Table 27 shows the target nucleic acids and the disease or disorder corresponding to each target nucleic acid.

[0511] In some embodiments of the present disclosure, the CRISPR-Cas12 system described herein is used to target and cleave specific genes listed in Table 27, thereby preventing, diagnosing, or treating the diseases or conditions corresponding to the genes in Table 27. For example, after targeted cleavage, indels are formed through cellular repair, and the target gene is knocked out, thereby inhibiting its function.

[0512] In some embodiments of the present disclosure, specific genes listed in Table 27 are targeted and modified by the CRISPR-Cas12 system described herein, thereby preventing, diagnosing or treating the diseases or conditions corresponding to the genes in Table 27.

[0513] In some embodiments of the present disclosure, the expression of specific genes listed in Table 27 is targeted and regulated by the CRISPR-Cas12 system described herein, thereby preventing, diagnosing or treating the diseases or conditions corresponding to the genes in Table 27.

[0514] In some embodiments, the pharmaceutical composition is delivered to a human subject in vivo. The pharmaceutical composition can be delivered by any effective route. Exemplary routes of administration include, but are not limited to, intravenous infusion, intravenous injection, intraperitoneal injection, intramuscular injection, intratumoral injection, subcutaneous injection, intradermal injection, intraventricular injection, intravascular injection, intracerebellar injection, intraocular injection, subretinal injection, intravitreal injection, intracameral injection, intratympanic injection, intranasal administration, and inhalation.

[0515] Diagnostic applications

[0516] Another aspect of the present disclosure relates to an in vitro composition comprising a CRISPR-Cas12 system described herein and a labeled detector DNA that is incapable of hybridizing to a guide polynucleotide described herein.

[0517] Another aspect of the present disclosure relates to the use of the CRISPR-Cas12 system described herein for detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid.

[0518] Another aspect of the present disclosure relates to the use of the CRISPR-Cas12 system described herein for detecting a target nucleic acid in a nucleic acid sample comprising the target nucleic acid.

[0519] In some embodiments, the target nucleic acid detected is a target RNA.

[0520] In some embodiments, the target nucleic acid detected is a target DNA. In some embodiments, the method for detecting the target DNA includes a Cas12 protein fused to a fluorescent protein or other detectable marker and a guide polynucleotide comprising a guide sequence specific to the target DNA. The binding of Cas12 to the target DNA can be visualized by a microscope or other imaging methods.

[0521] In some embodiments, the method for detecting a target nucleic acid in a cell-free system results in the production of a detectable label or enzyme activity. For example, by using a Cas12 protein, a guide polynucleotide comprising a guide sequence specific for the target nucleic acid, and a detectable label, the target nucleic acid will be recognized by Cas12. The binding of Cas12 to the target nucleic acid triggers its DNase activity, which results in the cleavage of the target nucleic acid and the detectable label.

[0522] In some embodiments, the detectable label is a DNA connected to a fluorescent probe and a quencher. The complete detectable DNA is connected to the fluorescent probe and the quencher to suppress fluorescence. After the detectable DNA is cut by Cas12, the fluorescent probe is released from the quencher and shows fluorescent activity. This method can be used to determine whether the target DNA is present in a cracked cell sample, a cracked tissue sample, a blood sample, a saliva sample, an environmental sample (such as water, soil or air sample), or other cracked cells or cell-free samples. This method can also be used to detect pathogens, such as viruses or bacteria, or to diagnose disease states, such as cancer.

[0523] In some embodiments, detection of a target nucleic acid aids in diagnosing a disease and / or pathological condition, or the presence of a viral or bacterial infection.

[0524] Table 27: Target nucleic acids / target genes and corresponding diseases or conditions (when there are two or more targets in a specific cell, it means that these two or more target genes are targeted simultaneously; without limitation, the targeting can be targeted knockout, single base editing, homologous recombination, targeted enhancement of its transcription, targeted inhibition of its transcription, etc.):

[0525] Example

[0526] The present disclosure is further illustrated by way of examples below, but the present disclosure is not limited to the scope of the examples. Experimental methods in the following examples without specifying specific conditions were performed according to conventional methods and conditions, or selected according to the product specifications.

[0527] Experimental Example 1: Screening of C12-279 protein

[0528] As shown in Figures 1A and 1B, Cas12i protein C12-279 was finally obtained through screening, analysis, and design using complex bioinformatics methods.

[0529] [Corrected 03.01.2025 according to Rule 91] The amino acid sequence of the C12-279 protein (SEQ ID NO: 696) is:

[0530] The predicted structure of the C12-279 protein is shown in FIG24 .

[0531] Experimental Example 2: Preparation and purification of C12-279 protein

[0532] 1. Vector Construction

[0533] The pET28a vector plasmid was double-digested with BamHI and XhoI, and the linearized vector was recovered by agarose gel electrophoresis. The prepared pXC12-279-GFPPAM-DR5 plasmid (SEQ ID NO: 697) was used as a template, and primers ChkCas-pET28-PF1 (SEQ ID NO: 699) and ChkCas-pET28-PR1 (SEQ ID NO: 700) were used to amplify the DNA fragment containing the coding sequence of the C12-279 protein by PCR. The DNA fragment was amplified by homologous recombination (NEB, Gibson The master mix was inserted into the cloning region of the pET28a vector to construct the recombinant vector C12-279-pET28a-01 (SEQ ID NO: 698). The reaction mixture was transformed into Stbl3 competent cells, plated on LB plates containing kanamycin sulfate, and cultured overnight at 37°C. Clones were then selected and sequenced.

[0534] Positive clones with correct sequences were selected and cultured overnight. The plasmids were extracted and transformed into expression strain Rosetta (DE3). The clones were spread on LB plates containing kanamycin sulfate and cultured at 37°C overnight.

[0535] 2. Protein Expression

[0536] A single clone was picked and inoculated into 5 mL of LB culture medium containing kanamycin sulfate and cultured at 37°C overnight.

[0537] The cells were inoculated into 500 mL of LB culture medium containing kanamycin sulfate at a ratio of 1:100, cultured at 220 rpm and 37°C to an OD of 0.6, and IPTG was added to a final concentration of 0.2 mM, and induced at 16°C for 24 h.

[0538] Rinse with 15 ml PBS, collect the cells by centrifugation, add lysis buffer and ultrasonically disrupt them. Centrifuge at 10,000 g for 30 min to obtain the supernatant containing the recombinant protein. Filter the supernatant through a 0.45 μm filter membrane and then apply it to the column for purification.

[0539] 3. Protein Purification

[0540] The expressed recombinant protein has 1240aa and a structure of His tag-NLS-C12-279-SV40 NLS-nucleoplasmin NLS. The six His residues at the N-terminus were used as purification tags and purified by IMAC (Ni Sepharose 6 Fast Flow, Cytiva). TM Heparin 50 μm chromatography column, Thermo Scientific TM ) was purified to obtain C12-279 recombinant protein. The purified recombinant protein was detected by SDS-PAGE electrophoresis, and the results are shown in Figure 2.

[0541] Experimental Example 3: C12-279 protein cleaves the PAM library in vitro and captures PAM

[0542] In this experimental example, the sgRNA containing a specific guide sequence and the C12-279 recombinant protein prepared and purified as described in Experimental Example 2 were mixed, and the in vitro cleavage substrate (containing a spacer sequence and a 7nt random sequence) was cleaved (as shown in Figure 3). After incubation at 37°C, the protein was purified, a library was constructed, and NGS sequencing and analysis were performed to determine the PAM sequence of C12-279. The specific steps are as follows:

[0543] A. In vitro cleavage of substrates

[0544] The designed in vitro cleavage substrate sequence is as follows:

[0545] In the sequence, N represents any of A, T, C, and G.

[0546] The double-stranded DNA containing the above sequence was prepared by PCR amplification method and used as an in vitro cleavage substrate.

[0547] The cleavage substrate was taken to a sequencing company for PCR-free library construction and NGS sequencing. The complexity and abundance of the PAM library composed of 7nt random sequences were analyzed. The results are as follows:

[0548] The composition of the four bases, A, T, G, and C, was essentially uniform. Furthermore, the PAM library, comprised of 7-nt random sequences, contained 4^7 = 16,384 different combinations, all of which were detected. The PAM library demonstrated acceptable complexity and abundance.

[0549] B. Preparation of sgRNA

[0550] An sgRNA containing a specific guide sequence (C12-279-sgRNA) was synthesized by in vitro transcription at 37°C in a system containing T7 RNA transcriptase, four ribonucleotide triphosphates, and a DNA template with a T7 promoter. The transcript was precipitated with LiCl and purified. The sgRNA sequence is as follows:

[0551] >C12-279-sgRNA

[0552] >C12-279-sgRNA-Rev

[0553] The sgRNA backbone sequence (direct repeat sequence, DR sequence) is: 5'-gtaatgcgtctcccattgacgc-3' (SEQ ID NO: 704).

[0554] The uppercase bases are the guide sequence of the sgRNA.

[0555] C. NGS library construction and PAM analysis

[0556] 1) PAM library cutting and T4 DNA Polymerase treatment

[0557] Reaction systems containing C12-279 protein, two different sgRNAs, in vitro cleavage substrates and buffer were prepared (as shown in Table 1) and reacted at 37°C for 3 h and 75°C for 15 min.

[0558] Table 1. Reaction system for in vitro cleavage reaction

[0559] 2) T4 DNA Polymerase treatment to fill in the cleavage product

[0560] T4 DNA Polymerase (Thermo Scientific) was added to the cleaved product. The specific reaction system is shown in Table 2. After addition, the reaction was incubated at 37°C for 20 minutes and then at 85°C for 10 minutes.

[0561] Table 2. Reaction system for filling in the C12-279 cleavage product

[0562] 3) Add A to the 3' end and add a biotin-labeled adapter

[0563] a. Add 78 μL of SPRISelect Beads (Beckman COULTER) to the T4 DNA Polymerase reaction product, mix well, and let it stand at room temperature for 5 minutes. Transfer the product to a magnetic rack for adsorption for 5 minutes, and transfer the supernatant to a new 1.5 mL tube. Then add 39 μL of SPRISelect Beads (Beckman COULTER), mix well, and let it stand at room temperature for 5 minutes. Transfer the product to a magnetic rack for adsorption for 5 minutes, discard the supernatant, wash twice with 85% ethanol, let it stand at room temperature for 10 minutes, and air-dry. Elute with 50 μL of ddH2O.

[0564] b. Use the SynplSeq DNA Library Prep Kit for Illumina to construct a library. Perform 3' A addition to the product in step a according to the system in Table 3. Incubate at 37°C for 10 min, 65°C for 20 min, and then at 4°C for later use.

[0565] Table 3. C12-279 cleavage products 3' plus A

[0566] c. Add Adapter 1 (obtained by annealing the upstream primer: 5'Biosg / gttgacatgctggattgagacttcctacactctttccctacacgacgctcttccgatc*t (SEQ ID NO: 705) and the downstream primer: gatcggaagagcgtcgtgtagggaaaga gtgtaggaagtctcaatccagcatgtcaac (SEQ ID NO: 706)) according to the system in Table 4. Incubate at 20°C for 30 minutes and then at 16°C overnight. Purify the reaction product using SPRISelect Beads.

[0567] Among them, Biosg represents biotin modification, and “*” represents thiolation.

[0568] Table 4. Reaction system with Adapter 1

[0569] d. Using streptavidin-labeled magnetic beads The reaction products were purified by M-280 Streptavidin (Invitrogen).

[0570] e.Recover PCR

[0571] Design the primers in Table 5 and use them according to the system in Table 6 and the reaction procedure in Table 7. Recover PCR reaction was performed using Hot Start High-Fidelty 2x Master Mix (NEB).

[0572] Table 5. Recover PCR primers

[0573] Table 6. Recover PCR reaction system

[0574] Table 7. Recover PCR reaction program

[0575] f. Move the Recover PCR product to a magnetic rack and allow it to adsorb for 5 minutes. Transfer the supernatant to a new 1.5ml centrifuge tube, take 3μL of the Recovery PCR product, and dilute it with 148.5μL of ddH2O.

[0576] g.Index PCR

[0577] Select the primers in Table 8 and perform Index PCR according to the system in Table 9 and the reaction program in Table 10.

[0578] Table 8. Index PCR primers

[0579] Table 9. Index PCR reaction system

[0580] Table 10. Index PCR reaction program

[0581] h. Index PCR products were purified using 0.7x SPRISelect Beads, eluted with 38 μl of ddH2O, and subjected to NGS sequencing after concentration determination using Qubit.

[0582] i. Analysis of NGS results: NGS sequencing was performed and WebLogo software was used to analyze the results according to the reference (A compact Cas9 ortholog from Staphylococcus Auricularis (SauriCas9) expands the DNA targeting scope. PLoS biology, 2020, 18(3), e3000686.), and the captured motifs shown in Figures 4 and 5 were obtained. Figures 4 and 5 both demonstrate that C12-279 can recognize a PAM with the sequence 5'-TTN-3'.

[0583] Experimental Example 4: Verification of PAM based on the in vivo editing activity of C12-279 protein

[0584] In this experiment, a plasmid library containing a 7nt random sequence was first constructed, followed by a bacterial expression plasmid containing the C12-279 protein coding sequence. After the expression plasmid was transformed into competent bacteria to prepare them, the 7nt random sequence plasmid library was electroporated. If plasmids in the 7nt random sequence library were recognized and targeted by C12-279, they would be removed from the library, and the corresponding bacteria would not grow. This is shown in Figure 6.

[0585] The specific steps are as follows:

[0586] 1. Construction of 7nt random sequence plasmid library

[0587] The pLVX-EF1a-BSD vector plasmid (SEQ ID NO: 711) was double-digested with EcoRV and XhoI, and the linearized vector was recovered by agarose gel electrophoresis. The prepared pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid (SEQ ID NO: 712) was used as a template, and primers Puro-PF1 (SEQ ID NO: 714) and Puro-PR1 (SEQ ID NO: 715) were used to amplify the DNA fragment containing the coding sequence of the Puro resistance gene by PCR. The DNA fragment was amplified by homologous recombination (NEB, Gibson The master mix was inserted into the digested pLVX-EF1a-BSD vector to construct the recombinant vector pLVX-7NN-Puro library plasmid (SEQ ID NO: 713) containing the 7NN random sequence. The reaction mixture was transformed into Stbl3 competent cells, plated on LB plates containing ampicillin, and incubated overnight at 37°C. All colonies were scraped and used for plasmid extraction.

[0588] 2. Construction of bacterial expression plasmid P15A-C12-279 for C12-279 protein and preparation of competent cells containing the plasmid

[0589] a. Construction of bacterial expression plasmid for C12-279 protein

[0590] The P15A-Cas-NC03 vector plasmid (SEQ ID NO: 716) was digested with SalI and the linearized vector was recovered by agarose gel electrophoresis. The prepared pXC12-279-GFPPAM-DR5 plasmid (SEQ ID NO: 697) was used as a template, primers ChkNLS-PF1 (SEQ ID NO: 718) and the synthesized C12-279-sgRNA fragment (SEQ ID NO: 717) were used to amplify the DNA fragment encoding the C12-279 protein and the fusion fragment expressing the sgRNA. The fragments were then cloned and cloned using homologous recombination (NEB, Gibson The recombinant vector P15A-C12-279 (SEQ ID NO: 719) was constructed by inserting the master mix into the digested P15A-Cas-NC03 vector. The reaction solution was transformed into Stbl3 competent cells, coated on LB plates containing chloramphenicol, and cultured overnight at 37°C. Single clones were then selected for sequencing verification.

[0591] b. Preparation of competent cells containing bacterial expression plasmid P15A-C12-279

[0592] The correct P15A-C12-279 plasmid verified by sequencing was transformed into DH5a competent cells (Weidi Biotechnology, CAT#: DL1001), and a single clone was picked and inoculated into LB medium containing chloramphenicol and cultured at 37°C overnight.

[0593] Prepare competent cells for electroporation as follows:

[0594] The culture solution was inoculated into 100 mL of fresh LB medium containing chloramphenicol at a ratio of 1:100 and cultured at 37°C and 220 rpm for amplification.

[0595] Cultivate until OD600 = 0.5, transfer the bacterial solution to a 50 mL centrifuge tube, and pre-cool on ice for 30 min;

[0596] Centrifuge at 4000 rpm, 4°C for 10 min, collect the cells, and resuspend in an equal volume of pre-cooled sterile water;

[0597] Repeat the above steps;

[0598] Resuspend the cells in 1 / 10 volume of pre-cooled sterile water containing 10% glycerol, divide the cells into 50 μL / tubes, and store at -80°C to obtain P15A-C12-279 competent cells.

[0599] c. Plasmid elimination to identify the C12-279 protein PAM sequence

[0600] 100 ng of the pLVX-7NN-Puro library plasmid was electroporated into the prepared P15A-C12-279 competent cells and DH5a electroporated cells, respectively, and labeled as Lib1 (electroporated into P15A-C12-279 competent cells) and Lib2 (electroporated into DH5a competent cells).

[0601] After electroporation, add 10 mL of LB medium and culture at 37°C, 220 rpm for 2 h.

[0602] After thawing, the bacterial suspension was centrifuged at 4000 rpm for 2 min to collect the cells. The cells were resuspended in 400 μl LB and then coated on LB plates. The DH5a electroporated cells were coated on LB plates containing ampicillin, while the P15A-C12-279 competent cells were coated on LB plates containing chloramphenicol and ampicillin. The cells were cultured at 37°C overnight.

[0603] The bacterial cells were scraped from the culture plates and the plasmid DNA was extracted by alkaline lysis method;

[0604] 100 ng of each of the two extracted plasmid DNA samples was used as a PCR template and PCR amplification was performed using primers SiteSeq-PF1 (SEQ ID NO: 720) and SiteSeqPuro-PR (SEQ ID NO: 721). The obtained fragments were used to construct amplicon libraries using an NGS library preparation kit (SynplSeq DNA Library Prep Kit for Illumina), and the completed libraries were sequenced by NGS.

[0605] The differences in NGS sequencing between transformed Lib1 and Lib2 cells were compared and analyzed, as shown in Figure 7. This demonstrates that C12-279 can recognize the PAM sequence of 5'-TTN-3'.

[0606] Experimental Example 5: Cleavage activity of C12-279 protein on target nucleic acid in 293T cells

[0607] In this experimental example, a sgRNA targeting the TTR gene in HEK293T cells was first designed and constructed into the pXC12-279-GFPPAM-DR5 plasmid (SEQ ID NO: 697) to obtain the TTR gene-targeting plasmid pXC12-279-TTR01. After transfection into HEK293T cells, the indel ratio was verified by NGS, and the cleavage activity of the C12-279 protein in 293T cells was verified. The specific steps are as follows:

[0608] 1. sgRNA Plasmid Construction

[0609] Based on the TTR gene sequence information in the HEK293T cell line, an sgRNA target site catgagcatgcagaggtgagtat (SEQ ID NO: 722) was designed.

[0610] Table 11. sgRNA annealing primers

[0611] Annealing primers were designed according to the sgRNA target information, and plasmid pXC12-279-GFPPAM-DR5 (SEQ ID NO: 697) was cloned using BsmBI (Thermo Scientific TM ,ER0451) and Acc65I (Thermo Scientific TM ,ER0901) enzyme digestion, primers (Table 11) annealed and ligated with the vector to obtain the pXC12-279-TTR01 expression clone.

[0612] 2. TTR gene editing efficiency detection

[0613] Plating: 293T cell lines were plated when the confluency reached 70-80%, and the number of cells seeded in a 24-well plate was 5*10^5 cells / well.

[0614] Transfection: Perform transfection 12-14 hours after plating. Add 100 μL Opti-MEM, 1.5 μL PEI (Yisheng Bio, Polyethylenimine Linear (PEI) MW25000), and 500 ng pXC12-279-TTR01 plasmid to each well of a 24-well plate, mix well, and add the culture medium to 293T cells after standing at room temperature for 20 minutes for cell transfection. Replace with fresh culture medium after overnight transfection and continue culturing.

[0615] DNA extraction, PCR amplification, and NGS library construction: After 72 h of culture, cells were washed with PBS and then 100 μL of cell lysis buffer (Viagen, Lysis Reagent (Cell) was used to generate a lysate containing genomic DNA. The region near the target sequence was amplified from the genomic DNA. The PCR product was subjected to NGS library construction and sequencing, and the sequencing results were analyzed. The indel rate was higher than 14%, as shown in Figure 8.

[0616] Experimental Example 6: PAM Identification of CI1062732

[0617] Compared with the C12-279 protein (SEQ ID NO: 696), the protein CI1062732 (SEQ ID NO: 46) lacks dozens of amino acid residues at the N-terminus.

[0618] The inventors previously used a method essentially identical to that used in Experimental Example 4 to identify the PAM sequence recognized by CI1062732. A gRNA containing the "false" DR sequence GTAATGCGTCTCCCATTGACGCC (SEQ ID NO: 529) was combined with CI1062732 to target a 7 nt random sequence plasmid library in bacteria. Sequencing analysis (as shown in Figure 9) revealed that the "false" PAM motif captured by CI1062732 was 5'-TTNC-3'.

[0619] Subsequently, the inventors analyzed the secondary structure of the "fake" DR sequence and hypothesized that the C base at the 3' end of the captured "fake" PAM motif TTNC might be due to an extra C at the 3' end of the DR ( Figure 10 ). Therefore, in subsequent CI1062732 and C12-279 tests, the DR sequence used was GTAATGCGTCTCCCATTGACGC (SEQ ID NO: 704).

[0620] Experimental Example 7: Screening, design, preparation, and testing of C12-101-07 protein

[0621] (1) Through complex bioinformatics methods, screening and design were carried out, and finally the C12-101-07 protein was obtained.

[0622] The amino acid sequence of the C12-101-07 protein is:

[0623] The recombinant vector C12-101-07-pET28a-01 (SEQ ID NO: 725) was constructed using essentially the same methods as in Experimental Example 2. The C12-101-07 recombinant protein, with an amino acid sequence of 1122 aa and a structure of His tag-NLS-C12-101-07-SV40 NLS-nucleoplasmin NLS, was expressed and purified. The purified recombinant protein was analyzed by SDS-PAGE, as shown in Figure 11.

[0624] (2) The C12-101-07 protein in vitro cleavage PAM library was prepared using a method substantially the same as that in Experimental Example 3 to capture PAM.

[0625] In vitro transcription was used to synthesize sgRNA containing a specific guide sequence, the sequence of which is as follows:

[0626] >C12-101-07-sgRNA

[0627] >C12-101-07-sgRNA-Rev

[0628] The sgRNA backbone sequence is: 5'-atcgcaacatctcagaaacccgtcctaagttgacgg-3' (SEQ ID NO: 534).

[0629] The C12-101-07 protein was combined with the forward and reverse gRNAs for editing.

[0630] The PAM motifs shown in Figures 12 and 13 were captured during the experiment.

[0631] It was demonstrated that C12-101-07 could recognize the PAM sequence of 5'-TTN-3'.

[0632] Experimental Example 8: In vivo editing activity of C12-101-09 protein in bacteria, verifying PAM

[0633] A similar protein C12-101-09 was designed based on C12-101-07, and its sequence is:

[0634] >C12-101-09

[0635] The recombinant vector plasmid P15A-C12-101-09 (SEQ ID NO: 729) was constructed using conventional methods and then tested using methods substantially the same as those in Experimental Example 4.

[0636] The motif recognized by C12-101-09 was captured, as shown in FIG14 .

[0637] It was demonstrated that C12-101-09 could recognize the PAM sequence of 5'-TTN-3'.

[0638] Experimental Example 9: Construction of cell lines containing different PAM reporter systems

[0639] Stable cell lines were prepared using lentiviral infection. We constructed lentiviral expression plasmids containing different PAM sequences, packaged them, and infected 293T cells to construct cell lines containing GFP reporter systems with different PAM sequences for subsequent mutant screening.

[0640] (1) Construction of lentiviral expression plasmids for GFP reporter systems composed of different PAM sequences

[0641] In addition to the prepared plasmid pCDH-CMV-EGFP-reporter3-EF1-Puro (SEQ ID NO: 712), a plasmid library with a PAM composition of NAAN was simultaneously constructed. The specific construction scheme is as follows:

[0642] A mixed library of EGFP fragments containing the detection system was gene synthesized, and the synthesized fragments were digested with XbaI+NotI enzymes. The fragments were then ligated with the XbaI+NotI enzyme-digested vector of the plasmid pCDH-CMV-EGFP-reporter3-EF1-Puro using T4 DNA ligase, transformed into Stbl3, and cultured on ampicillin-containing plates at 37°C overnight.

[0643] For the fragment mixed library, multiple clones need to be picked for sequencing and identification to obtain plasmids with 16 different sequence compositions (i.e., PAM sequences are AAAA, AAAT, AAAG, AAAC, TAAA, TAAT, TAAG, TAAC, GAAA, GAAT, GAAG, GAAC, CAAA, CAAT, CAAG, CAAC), and then mixed in equal mass to obtain the plasmid library Plasmid Lib (Puro-NAAN-eGFP-Lib, SEQ ID NO: 730) with a PAM sequence composition of NAAN for subsequent lentiviral packaging and stable cell line construction.

[0644] The unedited reporter system has a non-multiple of 3 base insertion from the start codon to the normal reading frame of GFP (pCDH-CMV-EGFP-reporter3-EF1-Puro has a 32bp base insertion, and the plasmid library Plasmid Lib has a 29bp base insertion, as shown in Figure 15), which causes the normal reading frame of GFP to be interrupted and GFP is not expressed; the sgRNA target is set inside the GFP expression frame, and indel is generated by editing of Cas12. There is a chance of restoring the normal reading frame of GFP, so that GFP can be expressed normally, and the higher the editing efficiency, the higher the probability that the generated indel will restore the correct reading frame of GFP. The number of cells that can normally express GFP is detected by flow cytometry to characterize the editing efficiency of Cas12 protein.

[0645] (2) Lentiviral packaging of GFP reporter system plasmids with different PAM sequences

[0646] The constructed pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid and the library plasmid Plasmid Lib were mixed with the viral packaging helper plasmids pMD2.G (Miaoling Biotechnology) and psPAX2 (Miaoling Biotechnology) at a molar ratio of 1:1:1, respectively, and then transfected into 293T cells using PEI. 48 hours after transfection, the culture supernatant was collected and filtered through a 0.45μm filter to obtain the crude viruses pCDH-CMV-EGFP-reporter3-EF1-Puro and Plasmid Lib.

[0647] (3) 293T cells were infected with pCDH-CMV-EGFP-reporter3-EF1-Puro and Plasmid Lib crude virus to construct detection cell lines

[0648] 1 / 4 volume of pCDH-CMV-EGFP-reporter3-EF1-Puro and Plasmid Lib crude virus were added to the culture medium to infect 293T cells. After 48 hours of infection, the medium was changed and 2 μg / ml of Puromycin was added for screening.

[0649] For 293T cells infected with pCDH-CMV-EGFP-reporter3-EF1-Puro, the screened cells were screened for monoclonal clones using limiting dilution, and the screened monoclonal clones were used as the cell line for detection (called Reporter3 cell line);

[0650] The cell pool obtained after drug screening of 293T cells infected with Plasmid Lib is the cell line used for testing (called NAAN cell line).

[0651] Experimental Example 10: Design of C12-279 protein mutants and testing of editing efficiency

[0652] a. Determination of mutation site

[0653] The 3D structure of the C12-279 protein was predicted and simulated through bioinformatics analysis and AI methods. The possible DNA binding, recognition and cleavage sites of C12-279 were analyzed in combination with the 3D structure. Mutant clones were constructed targeting these sites using molecular cloning point mutation methods.

[0654] First, the mutants shown in Table 12 were designed.

[0655] Table 12. Editing efficiency of different mutation sites of C12-279 in NAAN cell lines

[0656] b. Construction of mutant clones

[0657] After determining the specific mutation site, primers are used to introduce the mutated bases and construct expression clones containing different mutation sites. The following uses the D352R mutation clone construction as an example to illustrate.

[0658] Table 13. Primer sequences for the construction of C12-279 mutant clone C12-279-GFPPAM-01

[0659] For the mutation site D352R, primers were designed (as shown in Table 13), and the mutation site was introduced by primers C279-D352R-PF1 and C279-D352R-PR1. PCR amplification was performed using pXC12-279-GFPPAM-DR5 (SEQ NO: 1) as a template ChkCas12-PF1 + C279-D352R-PR1 (Yijin Bio, PC019, UltraHiPF TM DNA Polymerase Kit was used to obtain fragment C12-279-D352R-F1, and fragment C12-279-D352R-F2 was obtained by PCR amplification of ChkCas12-PR1+C279-D352R-PF1. Plasmid pXC12-279-GFPPAM-DR5 (SEQ ID NO: 697) was digested with HindIII+KpnI and recovered on gel (Guangzhou Meiji Biotechnology Co., Ltd., D2110, HiPure Gel Pure Micro Kit). The 5646 bp vector fragment was then recombined in vitro with fragments C12-279-D352R-F1 and C12-279-D352R-F2 (NEB, E2611L, Gibson Master Mix) was used to transform Escherichia coli by heat shock to obtain the mutant clone plasmid C12-279-GFPPAM-01.

[0660] c. Detection of editing efficiency of mutants in the NAAN reporter system

[0661] Plating: NAAN cell lines were plated when the confluency reached 70-80%, and the number of cells seeded in a 24-well plate was 5*10^5 cells / well.

[0662] Transfection: Transfection was performed 12-14 hours after plating. 1.5ul PEI (Yisheng Bio, 40815ES03, Polyethylenimine Linear (PEI) MW25000) + 500ng mutant clone plasmid was added to 100μl Opti-MEM per well of a 24-well plate, mixed, and added to the NAAN cell line after standing at room temperature for 20 minutes for cell transfection. After overnight transfection, fresh culture medium was replaced and cultured. After 72 hours of culture, flow cytometry was used to detect the editing efficiency of different mutant clones based on the proportion of GFP-positive cells.

[0663] The results are shown in Table 16 and Figure 16. The editing efficiency of different C12-279 mutants in NAAN cells is expressed as a multiple of the editing efficiency of C12-279. Multiple mutants improved the editing efficiency.

[0664] d. Based on the results of the first round of mutagenesis, we combined some mutation sites that significantly improved editing efficiency, attempting to combine multiple mutation sites to further improve editing efficiency. The designed mutants are shown in Table 14.

[0665] Table 14. Editing efficiency of different mutation site combinations of C12-279 in NAAN cell lines

[0666] The vector was constructed using essentially the same method as described above, and the editing efficiency was tested in the NAAN cell line. The results are shown in FIG17 .

[0667] e. Construction of targeted fixed PAM combination mutant plasmids

[0668] Since the PAM of the NAAN cell line is a mixed library, it may theoretically have a certain impact on the actual editing efficiency. To more intuitively and efficiently demonstrate the effect of mutations on editing efficiency, based on the sequence between the start codon of the Reporter3 cell line and the GFP reading frame, a target with a PAM sequence of TTG and a target sequence of CTCACCTCGCGACGCAATGATG (SEQ ID NO: 735) was selected for subsequent editing efficiency testing.

[0669] Construction of targeted Reporter3 cell line mutation clones

[0670] Based on C12-279-GFPPAM-DR5, the sgRNA was changed to target CTCACCTCGCGACGCAATGATG (SEQ ID NO: 735). Primers pCDH-PF1: GTACCGAAAAACATCATTGCGTCGCGAGGTGAGGCGTC (SEQ ID NO: 754) and pCDH-PR1: CATTGACGCCTCACCTCGCGACGCAATGATGTTTTTCG (SEQ ID NO: 755) were synthesized and annealed. Plasmid C12-279-GFPPAM-DR5 was digested with Acc65I (Thermo Scientific) and BsmBI (Thermo Scientific) to recover the vector and ligate with the annealed product to obtain the mutant clone C12-279-pCDH targeting the Reporter3 cell line. The method for constructing mutant clones targeting the Reporter3 cell line is the same as the protocol described in this experimental example, except that the vector plasmid needs to be changed from C12-279-GFPPAM-DR5 to C12-279-pCDH, and the primer ChkCas12-PR1 needs to be replaced with ChkCas12-PR2: GCGACGCAATGATGTTTTTCGGTACC (SEQ ID NO: 736). The rest of the methods and steps are exactly the same.

[0671] Some mutants were selected for editing efficiency testing in the Reporter3 cell line. The testing method was the same as the editing efficiency testing method of mutants in the NAAN reporter system, except that the cell line was changed from the NAAN cell line to the Reporter3 cell line.

[0672] Based on the results of the first round of mutations, the entire 3D structure was modified and labeled, and the first round of data was placed in a new model for predictive analysis. Finally, the possible second round of mutation sites were analyzed and predicted, and mutations and detection were performed. By analogy, through multiple rounds of mutation, selection, and accumulation, the optimal mutation combination for C12-279 was determined. Table 15 shows the C12-279 mutants and their editing efficiency in the Reporter3 cell line. Among them, the absolute value (average) of the editing efficiency of the C12-279-pCDH group, i.e., the C12-279 protein, was 8.55%.

[0673] Table 15. Editing efficiency of C12-279 mutants in Reporter3 cell line

[0674] (Expressed as multiple of editing efficiency compared to C12-279)

[0675] Experimental Example 11: Design of C12-101-07 protein mutants and detection of editing efficiency

[0676] Mutants shown in Table 16 were designed for the C12-101-07 protein.

[0677] The mutant vector plasmid was constructed using a method basically the same as that in Experimental Example 10, and the editing efficiency was tested.

[0678] a. Detection of editing efficiency of single-point mutations based on the NAAN reporter system

[0679] For the selected mutation site, C12-101-07-GFPPAM (SEQ ID NO: 737) was used as a cloning template plasmid and a control plasmid during the transfection test. Molecular cloning point mutation was used to construct mutant clones for editing efficiency detection. The results are shown in Table 16. The editing efficiency of different C12-101-07 mutants in the NAAN cell line is expressed as a multiple of the editing efficiency of C12-101-07. The absolute value (average value) of the editing efficiency of the C12-101-07-GFPPAM group, i.e., the C12-101-07 protein, is 0.23%.

[0680] Table 16. Different mutants of C12-101-07

[0681] b. Combining the mutation sites that significantly improved editing efficiency in the first round of mutagenesis, we attempted two point mutations. The mutants shown in Table 17 were designed.

[0682] Mutant clones were constructed using molecular cloning point mutagenesis methods, and the editing efficiency was determined based on the NAAN cell line.

[0683] The results are shown in Table 17. The editing efficiency of different C12-101-07 mutants in the NAAN cell line is expressed as a multiple compared to the editing efficiency of C12-101-07.

[0684] Table 17. C12-101-07 mutants

[0685] c. Construction of mutant clones of targeted Reporter3 cell lines and detection of editing efficiency

[0686] Using a method basically the same as that of Experimental Example 10 above, the target sequence on the original C12-101-07-GFPPAM vector was replaced, and the target targeting the NAAN cell line was replaced with the target targeting the Reporter3 cell line to obtain the control plasmid 101-07-sgRNA02. Then, the mutant vector was constructed, and the editing efficiency was tested in the Reporter3 cell line.

[0687] The results are shown in Table 18. The editing efficiency of different C12-101-07 mutants in the Reporter3 cell line is expressed as a multiple of the editing efficiency of C12-101-07. The absolute value (average) of the editing efficiency of the 101-07-sgRNA02 control group, i.e., the C12-101-07 protein, was 10.32%.

[0688] Table 18. Editing efficiency of different mutants of C12-101-07 in Reporter 3 cell line

[0689] (Expressed as a multiple of the editing efficiency compared to C12-101-07)

[0690] Experimental Example 12: Construction of different mutants by site-directed mutagenesis of C12-279 protein

[0691] The amino acid sequence of the C12-279 protein is SEQ ID NO: 696. Two mutation primers F / R are designed at the mutation site, and the required mutation sequence is introduced by the primers. Combined with the universal primers at both ends of the vector, PCR amplification is performed to obtain two mutation fragments F1 and F2. The mutation fragments F1 and F2 are homologously recombined with the enzyme-linearized vector to obtain a mutant plasmid. In this experimental example, the two sites of amino acid 426 and amino acid 860 are used as examples to construct single-point mutations (D426R, L860R) and multi-point mutation mutants (D426R&L860R). The specific steps are as follows:

[0692] The primers for site-directed mutagenesis at positions 426 and 860 are shown in Table 19.

[0693] Table 19. Primers

[0694] Construction of single point mutation clones at sites 426 and 860

[0695] ChkCas12-PF1+279_426_R was amplified by PCR using C12-279-pCDH plasmid as template (Yijin Bio, PC019, UltraHiPF TMDNA Polymerase Kit was used to obtain the mutant fragment D426R-F1, and the mutant fragment D426R-F2 was obtained by PCR amplification of ChkCas12-PR2+279_426_F. The plasmid C12-279-pCDH was digested with HindIII+KpnI and the 5647 bp vector fragment was recovered by gel (Guangzhou Meiji Biotechnology Co., Ltd., D2110, HiPure Gel Pure Micro Kit). The fragment was then recombined in vitro with the fragments D426R-F1 and D426R-F2 (NEB, E2611L, Gibson Master Mix) and heat-shocked to transform Escherichia coli to obtain the 426-site mutant plasmid C12-279-pCDH-426;

[0696] The same steps and methods were used to replace primers 279_426_F with 279_860_F and 279_426_R with 279_860_R. The same amplification and recombination were performed, and Escherichia coli was heat-shock transformed to obtain the 860 site mutant plasmid C12-279-pCDH-860. The construction method of Cas protein mutant plasmids at other different sites was consistent with the above-mentioned method for sites 426 and 860, except that mutant primers for each different site needed to be designed.

[0697] Verification of editing activity of Cas protein mutants

[0698] The test was conducted based on the Reporter3 cell line constructed in the aforementioned experimental example. The sequence between the start codon and the GFP reading frame is 32 bases, which is not an integer multiple of 3, resulting in abnormal GFP reading frame and GFP not being expressed. The Reporter3 cell line was edited using a Cas protein mutant that recognizes a target with a PAM of TTG and a target sequence of CTCACCTCGCGACGCAATGATG (SEQ ID NO: 735), generating indels and restoring the normal GFP reading frame. The proportion of cells in which GFP expression was restored was detected by flow cytometry to characterize the editing efficiency of different mutants. The specific steps are as follows:

[0699] Cell culture and plating: Cell lines were plated when they reached 70-80% confluency. The number of cells seeded in a 24-well plate was 5*10^5 cells / well.

[0700] Transfection: Transfection was performed 12-14 hours after plating. 1.5ul PEI (Yisheng Bio) + 500ng mutant plasmid was added to 100μl Opti-MEM per well of a 24-well plate, mixed, and added to the Reporter3 cell line for cell transfection after standing at room temperature for 20 minutes. Fresh culture medium was replaced after overnight transfection and continued to be cultured. After 72 hours of culture, flow cytometry was used to detect the editing efficiency of different mutant clones according to the proportion of GFP-positive cells, and the average value of multiple batches of data was taken. The specific results are shown in Table 20 and Figure 18. The average absolute value of the editing efficiency of the wild-type C12-279 group was 9.0%.

[0701] Table 20. Editing efficiency of C12-279 mutants in Reporter3 cell line

[0702] (Expressed as multiple of editing efficiency compared to C12-279)

[0703] In the table, nR notation (n is an integer) indicates that the n-th amino acid residue mutation is R.

[0704] Construction of clones with different mutation site combinations of Cas variant proteins and verification of editing activity:

[0705] Based on the editing efficiency results of different mutation sites mentioned above, different mutation sites were selected for multiple mutation combinations. The plasmid construction of 426 and 860 double-site Cas mutants was used as an example to introduce it.

[0706] Using C12-279-pCDH plasmid as a template, ChkCas12-PF1+279_426_R was PCR amplified to obtain the mutant fragment D426R-F1, 279_426_F+279_860_R was PCR amplified to obtain the mutant fragment D426R-L860R-F1, and ChkCas12-PR2+279_860_F was PCR amplified to obtain the mutant fragment L860R-F2. Plasmid C12-279-pCDH was digested with HindIII+KpnI, and the 5647 bp vector fragment was recovered from gel. It was recombined in vitro with fragments D426R-F1, D426R-L860R-F1, and L860R-F2, and heat-shocked transformed into Escherichia coli to obtain the variant plasmid C12-279-pCDH-426-860. The construction method of other plasmids with multiple mutations was consistent with that of C12-279-pCDH-426-860.

[0707] Among them, the amino acid sequence of the Mut-02-1-426-846-858-860 mutant is:

[0708] The editing activity of multiple point mutants was verified using the same method as described above in this experimental example.

[0709] The results showed that the editing efficiency of most multi-point mutants was improved compared to the wild type, as shown in Table 21 and Figure 19. The absolute value of the editing efficiency of the wild type C12-279 group was 8.7%.

[0710] Table 21. Editing efficiency of C12-279 mutants in Reporter3 cell line

[0711] (Expressed as multiple of editing efficiency compared to C12-279)

[0712] In the table, the integer n indicates that the amino acid residue at position n has mutated to R. Mut-01 is the Q186R variant, Mut-02 is a double-point mutant of Q186R and D352R, and Mut-03 is a triple-point mutant of G184R, Q186R, and D352R. 5+426+860 indicates that positions 5, 426, and 860 have simultaneously mutated to R. Mut-02+860 represents a multi-point mutant obtained by introducing an additional mutation at position 860 (to arginine R) into the Mut-02 mutant. The same applies to other multi-point mutants (mutated to arginine R).

[0713] Experimental Example 13: Detection of editing efficiency of different mutants against different targets of TTR and HBG

[0714] In this experimental example, the TTR gene (GeneBank: NG_009490.1) and the HBG gene (GeneBank: NC_000011.10) were first targeted, the PAM sequence was selected as the target of TTN, and sgRNAs targeting different positions were designed and constructed (as shown in Table 22) for combination with different mutants to detect the editing efficiency.

[0715] Table 22. sgRNAs targeting TTR and HBG

[0716] sgRNA plasmids were constructed according to the sgRNA sequences in the above table: the vector plasmid SpCas9-gRNA-pUC57Kan was linearized using BbsI (Thermofisher) and XhoI (Thermofisher), primers were synthesized for different sgRNA sequences, annealed and ligated into the linearized vector, and Escherichia coli was transformed to obtain the final sgRNA expression vector plasmid.

[0717] Different mutant plasmids were combined with different sgRNA plasmids, and HEK293T cells were transfected with PEI. After 48 h, the cells were collected and lysed using DirectPCR Lysis Reagent (Cell) (VIAGEN: 302-C). Different primers were selected according to different targets for PCR amplification, followed by Sanger sequencing, and the editing efficiency was analyzed using TIDE.

[0718] Specific amplification and sequencing primers are shown in Table 23. The experiment was repeated three times, and the editing efficiency results are shown in Figures 20A and 20B. Each sgRNA group in each batch of experiments produced effective editing, and the editing efficiency at many target sites was better than that of wild-type C12-297.

[0719] Table 23. Primers for amplification and sequencing of different targets

[0720] Experimental Example 14: Preparation and Activity Detection of Inactivated C12-279 Mutant

[0721] Based on the multi-point mutants designed in the previous experimental example, any one or more of the D651A, E891A, and D1082A mutations were further introduced. The same method as in the previous experimental example 10 was used to construct a mutant clone plasmid, which was edited in the Reporter3 cell line. The proportion of cells that restored GFP expression after editing was analyzed by flow cytometry to determine the mutant editing efficiency. The specific results are shown in Figure 21. As long as each mutant in the figure contains at least one of the D651A, E891A, and D1082A mutations, the editing efficiency is reduced to below 1.3%. DeadCas12 (dCas12) can be obtained by introducing any of D651A, E891A, and D1082A.

[0722] Among them, the amino acid sequence of the dCas mutant Mut-02-1-426-846-858-860-D651A-E891A-D1082A is:

[0723] Experimental Example 15: mRNA+gRNA delivery

[0724] The TTR gene was edited using the encoding mRNA of the Mut-02-3-426-860 mutant (obtained by in vitro transcription) together with the modified gRNA (as shown in Table 24).

[0725] Table 24. Modified gRNA

[0726] In the table, “r” indicates a natural base, “d” indicates a deoxy modification, “m” indicates a methylation modification, and “*” indicates a phosphorothioate modification.

[0727] HEK293 cells were transformed with mRNA and gRNA by electroporation. After 48 hours, cells were harvested and lysed using DirectPCR Lysis Reagent (Cell) (VIAGEN: 302-C). Sanger sequencing was performed, and editing efficiency was analyzed using TIDE. The results are shown in Table 24.

[0728] After adding DNA nucleotide sequence at the 3' end, the editing activity is significantly improved.

[0729] In addition, primers TTR-NGS-PF2 (GCGTAACTTAATCCAGACTTTCACACCTT, SEQ ID NO: 823) and TTR-NGS-PR2 (GGTCATTCATCACCTTCCTTAGGACA, SEQ ID NO: 824) were used for PCR amplification and library construction, followed by NGS sequencing to test editing efficiency. The mutant Mut-02-3-426-860 combined with C279-dmTTR01-01 achieved an electroporation editing efficiency of 79.64%.

[0730] Subsequently, C279-dmTTR01-02 was tested in combination with different mutants. The specific NGS sequencing results are shown in Table 25 and Figure 22.

[0731] Table 25. Editing efficiency of modified gRNA (C279-dmTTR01-02) in combination with different mutants (NGS assay)

[0732] Experimental Example 16: Verification of Editing Efficiency of Different Endogenous Target Genes

[0733] Different gRNA targets were designed according to different disease-related targets (Table 26), and the mutant Mut-02-1-426-846-858-860 protein was edited in combination with the gRNA in Table 26. The same method as in Example 13 was used, including plasmid construction, transfection, and editing efficiency testing. The results are shown in Table 26.

[0734] Table 26. Editing efficiency of targeting different endogenous genes

[0735] Experimental Example 17: PAM Recognition of C12-279 Mutant Mut-02-1-426-846-858-860

[0736] [Corrected 03.01.2025 according to Rule 91] Using the same method as in the previous example, in vivo editing experiments were performed in bacteria, and finally, NGS sequencing confirmed that the PAM recognized by the mutant Mut-02-1-426-846-858-860 was: 5'-WTN-3' (W is A or T), as shown in Figure 23.

[0737] Although specific embodiments of the present disclosure have been described above, those skilled in the art will appreciate that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present disclosure. Therefore, the scope of protection of the present disclosure is defined by the appended claims.

Claims

1. A Cas12 protein, characterized in that (a) the Cas12 protein is CLUSTER1 protein, CLUSTER2 protein, CLUSTER3 protein, CLUSTER4 protein, CLUSTER5 protein, CLUSTER6 protein, CLUSTER7 protein, CLUSTER8 protein, CLUSTER9 protein, CLUSTER10 protein, CLUSTER11 protein, CLUSTER12 protein or CLUSTER13 protein; or, (b) the amino acid sequence of the Cas12 protein comprises or has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-53, 696 and 728; Optionally, the Cas12 protein retains the function of a protein having an amino acid sequence as shown in any one of SEQ ID NOs: 1-53, 696 and 728; Optionally, the PAM sequence (5'→3') recognizable by the Cas12 protein is selected from any one or more of the following: A, C, T, G, TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA, <h2 style=";text-align:left;direction:ltr">NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT ANG、GNA、GTT、NGA、TAA、GTA、GGN、GNT、NCG、ATT、CCA、CNN、AAA、AAC、ATN、GA G、CTG、ACG、NAA、TAN、NAT、CNA、GCN、GTC、NCN、CTN、CNC、ANT、NNC、CAG、NAN、 ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG <h2 style=";text-align:left;direction:ltr">NGAA, GANC, GCNC, NTNT, TGGG, AAGG, AAGN, NTNN, TCGT, CNTG, NTGG, CCGN, ATAT, TGCA, NGGT, TGNT, NNTG, NCCG, ACAT, GNTG, CGCG, GACN, NTCG, TCNG, CTGC, TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, C GGC、CGAN、CNCA、ATNT、NNNG、NGCT、CTGG、GGAN、NTNC、ATTC、AATG、CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT 、NAGG、NANC、CTTA、GTCT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、GACC、AACG、TTAA、TC CG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GNAN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、 CNNT、AGGC、GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CGAA、CNAC、GCCT、TAGG、ANGC、TNA A, GANT, NCNA, NCCT, AGAN, GTAA, TTTN, ATGA, TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAAN, CCNT, GANA, CACA, CTNA, ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT <h2 style=";text-align:left;direction:ltr">TTTC、GNCG、TACA、GTAC、GAGC、ACNN、ATGG、AANT、ATCC、ACCG、AGNC、TGTT、NCAT、ATTA、GNTT、GAGN、TNAC、GCCG、NTNG、GTGG、GNGN、ACCA、NTAA、ACTN、NCTG、 NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNT, CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN TNGG、CNNN、AAGT、ATTN、GGNN、CAGC、CGTN、GCCC、GCTT、CNAT、NANA、CCNN、GN GA、TNGN、GCAG、CGNG、CCTT、NGAG、NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、 ACGG, TGAC, TCCN, ANNN, TCGN, TAAN, CAGG, TTAN, NGAN, NTGC, CCNC, TNTN, ATGN, GTGN, GCAT, NNGN, NNCC, CCNA, CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN, GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN, GAAC, NCCA, TTCC, NCNN, GNNN, ANGT, NTNA, CCCT, GNAA, TTNG, GTNN, GGNG, TCTA, NCAC, GANG, TTCG, CCTC, CNGG, ANNA, TCAN, ATCG, NTGA, CGTA, TTAC, GCTN, GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT<h2 style=";text-align:left;direction:ltr">AGAT、GATG、GCCN、TGNG、GCGC、CCGA、GNCN、NTTG、NNAT、TNCG、NANG、GGTG、NC CC、GNCC、CAAT、CGCN、CNGA、NTTC、TTCT、NGGA、AGTC、CNNC、NACG、AGTN、NANN、 ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN, GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG, ACGN, GGCA, AAAAG, TTGT, NGNA, NAAN, TATN CGGG, CATA, ATGC, ACGC, ACCN, ATTT, TCNA, TNGC, NACA, NACC, CTCN, GGCC, TANG, AGAA, TNGA, TAGC, CAGN, GGCN, ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT,TCGG,GCAC,GTAG,NTCA,NATT,ANTA,CCCN,ACTA,AAAA,GAAN,TATT,NNAC,TGAT,GGGN,CCAA,GNGG,CCAN,GTCC,NNCT,AGNG,CNTT,CNCT,GANN,GGTT, AGCT、CATG、NTAC、TNCN、NNTN、TGGA、GATT、AGCA、TAAG、GCGA、ACTT、ANGN、NT GN、AACN、AACT、TCAA、NTAT、TCGA、NCTC、NNGG、ANGG、NNTT、GTNT、CTNN、CGGN、 TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC, ATNN, NNNT, AGAC, AACC, ANNC, ANNG, ACAN GTTC、TATA、GNTC、NCGT、NGNT、CGTC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TA TC、AGNA、CTNC、TTCA、ANCA、ACCC、AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、<h2 style=";text-align:left;direction:ltr">AANN, CNAN, NNCN, ANAA, ANAN, CTTG, NGNN, AGAG, TANA, TCNC, GCAA, NGNG, NAGC, NATA, ATCN, CGTT, CNGC, GATN, NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATA N, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATT, NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC; Optionally, the Cas12 protein is at positions 1-10, 12-16, 19, 21-30, 32-44, 46-51, 53-62, 64-86, 88-99, 101-106, 108-112, 114-149, 151-167, 169-180, 182-192, 194-200, 202-225, 227-240, 242-245, 247-253, 255-263, 265-276, 278-279, 281-303, 305-306, 308-310, 313, 314-315 5-325, 327-337, 339-348, 350-429, 431-433, 435-437, 439-444, 446-464, 467, 469-494, 496-497, 499-504, 506-529, 531-550, 552-553, 555-587, 589-590, 592-599, 601-616, 618-628, 630-676, 678-681, 683-686, 688-689, 691-713, 715-717, 71 9-725, 727-734, 736-749, 751-762, 764-769, 771-776, 779-787, 789-792, 794-795, 797-798, 800-802, 804-815, 817-819, 821-842, 844-860, 862-868, 870, 872-877, 879-906, 909-914, 916-928, 930-961, 963-977, 979-983, 985-1018, 1021, 1023 -1033, 1035-1050, 1052-1053, 1055-1105, 1107-1113, 1115-1117, 1120, 1122-1125, 1128-1135 amino acid residues at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11 or at least 12 have mutations; optionally, the mutations are mutations to any other natural amino acid residues; further optionally, the mutations are mutations to residues R, H, K or A; Optionally, the Cas12 protein has any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more amino acid mutations as shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 696: I33R, G184R, S185R, Q186R, G194R, N195R, G196R, G197R, N245R, G256R, L260R, Y278R, S285R, Y316R, H350R, D3 52R, A355R, A356R, C385R, P386R, H387R, G390R, K391R, N392R, D429R, Q461R, Q462R, Q469R, E485R, S491R, K521R, P525R, L611R, K629R, K631R, N633R, D841R, N898R, K987R, A988R, G989R, Q990R, T991R, D1010R, E1013R, A1136R, K1138R, T1139R; Optionally, the Cas12 protein has any one, any two, any three, any four, any five, or more amino acid mutations as shown in SEQ ID NO: 696 at the corresponding position of the sequence shown in SEQ ID NO: 696: D352R+Q186R, D352R+L260R, D352R+A355R, A355R+L260R, P386R+C385R, E485R+Q462R, D352R+Q186R, A 355R+L260R, G184R+R186Q, D352R+Q186R+I33R, D352R+Q186R+G184R, D352R+Q18 6R+S185R, D352R+Q186R+G256R, D352R+Q186R+Y278R, D352R+Q186R+S285R, D352 R+Q186R+Y316R, D352R+Q186R+H350R, D352R+Q186R+A356R, D352R+Q186R+Q469R , D352R+Q186R+S491R, D352R+Q186R+K521R, D352R+Q186R+P525R, D352R+Q186R+ K629R, D352R+Q186R+N633R, D352R+Q186R+D841R, D352R+Q186R+N898R, D352R+Q 186R+K987R, D352R+Q186R+T991R, D352R+Q186R+D1010R, D352R+Q186R+E1013R; Optionally, the Cas12 protein can form a complex with the guide polynucleotide; further, the complex can specifically bind to the target nucleic acid; further, the complex can cut the target nucleic acid, modify the target nucleic acid and / or regulate the expression of the target nucleic acid; Optionally, the Cas12 protein can form a complex with a guide polynucleotide, wherein the guide polynucleotide comprises a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence, and the backbone sequence can interact with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; Optionally, the backbone sequence does not include the backbone sequence of tracrRNA; Optionally, the PAM sequence recognizable by the Cas12 protein is 5'-TTN-3' and / or 5'-TTNC-3'; Optionally, the Cas12 protein is a variant with inactivated nuclease activity; optionally, the Cas12 protein is a dead Cas12 inactivated variant or a nickase Cas12 inactivated variant; optionally, the Ruvc domain of the Cas12 protein is inactivated.

2. A guiding polynucleotide, characterized in that It comprises (i) a direct repeat sequence, the direct repeat sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to any one of SEQ ID NOs: 54-583, and (ii) a guide sequence engineered to hybridize with a target nucleic acid; the direct repeat sequence is connected to the guide sequence, and the guide polynucleotide is capable of forming a complex with the Cas12 protein and guiding the complex to sequence-specific binding to the target nucleic acid through the guide sequence; Preferably, the Cas12 protein is the Cas12 protein as claimed in claim 1; Optionally, the guide sequence comprises 15-60 nucleotides; Optionally, the guide sequence hybridizes to the target nucleic acid, and the guide sequence has no more than one nucleotide mismatch with the target nucleic acid; Optionally, the nucleotide sequence of the guide sequence is shown in any one of SEQ ID NOs: 54-583, 704; Optionally, the guide polynucleotide further comprises a tracrRNA; Optionally, the tracrRNA sequence is linked to the direct repeat sequence; Optionally, the tracrRNA comprises 10-200 nucleotides; Optionally, the guide sequence is located at the 3' end of the direct repeat sequence; Optionally, the guide sequence is located at the 5' end of the direct repeat sequence; Optionally, the tracrRNA sequence is located at the 5' or 3' end of the direct repeat sequence; Optionally, the tracrRNA sequence has at least 50% sequence identity with any one of SEQ ID NOs: 584-695.

3. A Cas12 inactivated variant, characterized in that The Cas12 inactivated variant is a nuclease activity inactivated variant of the Cas12 protein as claimed in claim 1; the nuclease activity inactivation means that a double-stranded gap cannot be generated on the target nucleic acid; Optionally, the Cas12 inactivated variant is a dead Cas12 inactivated variant or a nickase Cas12 inactivated variant; Optionally, the Cas12 inactivated variant is a variant obtained by inactivating the Ruvc domain of the Cas12 protein.

4. A fusion protein or conjugate, characterized in that The fusion protein or conjugate comprises the following elements: (1) the Cas12 protein as described in claim 1, or the inactivated variant of Cas12 as described in claim 3; and (2) homologous or heterologous functional domains; Optionally, the functional domain has an enzymatic activity that modifies a target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylylation activity, deadenylating activity, SUMOylating activity, deSUMOylating activity, myristoylation activity and / or demyristoylation activity; Optionally, the functional domain is selected from one or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase inhibitory domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, affinity tag, reporter tag and affinity domain and reporter domain; Optionally, the nuclease domain comprises a polypeptide having ssDNA cleavage activity and / or a polypeptide having dsDNA cleavage activity; Optionally, the Cas12 protein or the inactivated variant of Cas12 is directly or indirectly connected to the homologous or heterologous functional domain; preferably, the direct connection is covalent connection, and the indirect connection is connected through an amino acid linker or a non-amino acid linker; Optionally, the homologous or heterologous functional domain is fused or conjugated at the N-terminus, C-terminus or internally relative to the Cas12 protein or inactivated variant.

5. An isolated nucleic acid, characterized in that The nucleic acid encodes the Cas12 protein according to claim 1, the Cas12 inactivated variant according to claim 3, or the fusion protein or conjugate according to claim 4; Optionally, the nucleic acid is codon optimized for expression in a cell; Optionally, the nucleic acid is codon optimized for expression in prokaryotes; Optionally, the nucleic acid is codon optimized for expression in eukaryotic cells; Optionally, the nucleic acid is codon optimized for expression in a eukaryote, preferably a mammal such as a human or non-human mammal, a plant, an insect, a bird, a reptile, a rodent (e.g., a mouse, a rat), a fish, a worm / nematode, or a yeast.

6. A CRISPR-Cas12 system, characterized in that The CRISPR-Cas12 system comprises: a. The Cas12 protein as described in claim 1, the Cas12 inactivated variant as described in claim 3, the fusion protein or conjugate as described in claim 4, or the nucleic acid as described in claim 5; and b. The guide polynucleotide of claim 2, or a polynucleotide sequence encoding the guide polynucleotide; The Cas12 functional domain or the fusion protein or conjugate forms a complex with the guide polynucleotide; the guide polynucleotide comprises a guide sequence, and the guide sequence is engineered to guide the sequence-specific binding of the complex to the target nucleic acid; Optionally, the guide polynucleotide further comprises a direct repeat sequence linked to the guide sequence, preferably, the direct repeat sequence has at least 50% sequence identity compared to any one of SEQ ID NOs: 54-583 and 704; Optionally, the guide sequence comprises 15-35 nucleotides, and / or the guide sequence hybridizes with the target nucleic acid, the guide sequence and the target nucleic acid are 90%-100% complementary, preferably with no more than one nucleotide mismatch; Optionally, the guide sequence comprises 15-60 nucleotides; Optionally, the guide sequence hybridizes to the target nucleic acid; Optionally, the guide sequence mismatches with the target nucleic acid by no more than one nucleotide; Optionally, the guide polynucleotide further comprises a tracrRNA; Optionally, the tracrRNA sequence is linked to the direct repeat sequence; Optionally, the tracrRNA comprises 10-200 nucleotides; Optionally, the guide sequence is located at the 3' end of the direct repeat sequence; Optionally, the guide sequence is located at the 5' end of the direct repeat sequence; Optionally, the tracrRNA sequence is located at the 5' or 3' end of the direct repeat sequence; Optionally, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA; Optionally, the DNA is eukaryotic DNA; preferably, the eukaryotic DNA is non-human mammal DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent DNA, fish DNA, worm / nematode DNA or yeast DNA; Optionally, the target nucleic acid is a disease or disorder-related gene or a signal transduction biochemical pathway-related gene, or the target nucleic acid is a reporter gene; optionally, the target nucleic acid is a gene listed in Table 27.

7. A vector system, characterized in that The vector system comprises one or more recombinant vectors, wherein the recombinant vector comprises the isolated nucleic acid of claim 5, or the CRISPR-Cas12 system of claim 6; Optionally, the recombinant vector further comprises a regulatory sequence; Optionally, the polynucleotide sequence encoding the Cas12 protein, Cas12 inactivated variant or fusion protein or conjugate is operably linked to a regulatory sequence, and / or the polynucleotide sequence encoding the guide polynucleotide is operably linked to a regulatory sequence; preferably, the regulatory sequence is selected from: one or more of a promoter, an enhancer, an internal ribosome entry site and a transcription termination signal, such as a constitutive promoter, an inducible promoter, a broad-spectrum promoter or a tissue-specific promoter, and / or the transcription termination signal is such as a polyadenylation signal or a poly-U sequence; Optionally, the backbone of the recombinant vector is an adeno-associated virus vector, a lentivirus vector or a virus-like particle; preferably, when the backbone When it is an adeno-associated virus vector, the adeno-associated virus vector is a recombinant adeno-associated virus vector of serotype AAV1, AAV2, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12 or AAV13, and when the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; Optionally, the isolated nucleic acid is ligated to an aptamer sequence; When the backbone of the recombinant vector is a virus-like particle, the isolated nucleic acid is connected to the gene encoding the gag protein.

8. A delivery system, characterized in that The delivery system comprises: (1) a delivery vehicle, and (2) the Cas12 protein of claim 1, the guide polynucleotide of claim 2, the inactivated variant of Cas12 of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, or the vector system of claim 7; Optionally, the delivery vehicle is a virus, a lipid nanoparticle, a nanoparticle, a liposome, an exosome, a microbubble or a gene gun; Optionally, the delivery vehicle is a lipid nanoparticle comprising the guide polynucleotide and mRNA encoding the Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate.

9. A cell, characterized in that The cell comprises the Cas12 protein of claim 1, the guide polynucleotide of claim 2, the Cas12 inactivated variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, or the delivery system of claim 8; Optionally, the cell is a prokaryotic cell; Optionally, the cell is a eukaryotic cell; Optionally, the eukaryotic cell is a mammalian cell.

10. A pharmaceutical composition, characterized in that The pharmaceutical composition comprises the Cas12 protein of claim 1, the guiding polynucleotide of claim 2, the inactivated variant of Cas12 of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, the delivery system of claim 8 or the cell of claim 9; Preferably, the pharmaceutical composition comprises a pharmaceutically acceptable excipient.

11. A kit, characterized in that: The kit comprises the Cas12 protein of claim 1, the guiding polynucleotide of claim 2, the inactivated variant of Cas12 of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, the delivery system of claim 8, the cell of claim 9 or the pharmaceutical composition of claim 10.

12. Use of the Cas12 protein of claim 1, the guiding polynucleotide of claim 2, the inactivated variant of Cas12 of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, the delivery system of claim 8, the cell of claim 9, the pharmaceutical composition of claim 10 or the kit of claim 11 in the preparation of an agent or drug for diagnosing, treating and / or preventing a disease or disorder associated with a target nucleic acid; Optionally, the disease or condition is a blood disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer or an infectious disease; and / or, the agent or drug is used to: cut one or more target nucleic acid molecules or cause a nick in one or more target nucleic acid molecules, activate or upregulate the expression of one or more target nucleic acid molecules, activate or inhibit the transcription of one or more target nucleic acid molecules, inactivate one or more target nucleic acid molecules, visualize, label or detect one or more target nucleic acid molecules, bind one or more target nucleic acid molecules, transport one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules; Optionally, the disease or condition is selected from the diseases or conditions listed in Table 27; Optionally, the disease or disorder is selected from: hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidosis, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosidosis, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bieder syndrome, white punctate retinal degeneration, leukocyte adhesion deficiency type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti crystal dystrophy, Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult glucan body disease, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophosphatasia, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin lymphoma, non-muscle-invasive bladder cancer, non-alcoholic Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scars, obesity, Charcot-Marie-Tooth disease type 1A, Charcot-Marie-Tooth disease type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial Tay-Sachs disease, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, junctional epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, Canavan disease, cocaine addiction, Krabbe disease,Crigler-Najjar syndrome, oral cancer, happy puppet syndrome, diffuse intrinsic pontine glioma, Lafora disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, anemia of chronic kidney disease, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Netherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, Glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Rett syndrome, triple-negative breast cancer, Sandhoff disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscinosis, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube Cancer, bilateral vestibular disease, Stargardt's disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, head and neck squamous cell carcinoma, Wilson's disease, stable angina, Usher syndrome, choroideremia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, new coronavirus infection, pleural mesothelioma, acne vulgaris sores, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive dystrophic epidermolysis bullosa, malignant osteosclerosis of infancy, dystrophic epidermolysis bullosa, morphea, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic disease, lipoprotein lipase deficiency, severe congenital neutropenia, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, transthyretin (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.

13. A method for detecting, binding or cleaving a target nucleic acid, characterized in that: The method comprises contacting a target nucleic acid with a Cas12 protein as described in claim 1, a guide polynucleotide as described in claim 2, an inactivated variant of Cas12 as described in claim 3, a fusion protein or conjugate as described in claim 4, a nucleic acid as described in claim 5, a CRISPR-Cas12 system as described in claim 6, a vector system as described in claim 7, a delivery system as described in claim 8, a cell as described in claim 9, a pharmaceutical composition as described in claim 10, or a kit as described in claim 11; Optionally, the method is a method for non-diagnostic and / or therapeutic purposes; and / or the fusion protein or conjugate comprises a detectable label, such as a label detectable by fluorescence, Southern blot or FISH.

14. A method for changing a cell state, characterized in that: The method comprises contacting a cell with a Cas12 protein as claimed in claim 1, a guide polynucleotide as claimed in claim 2, a Cas12 inactivated variant as claimed in claim 3, a fusion protein or conjugate as claimed in claim 4, a nucleic acid as claimed in claim 5, a CRISPR-Cas12 system as claimed in claim 6, a vector system as claimed in claim 7, a delivery system as claimed in claim 8, a cell as claimed in claim 9, a pharmaceutical composition as claimed in claim 10, or a kit as claimed in claim 11, thereby changing the cell state; Optionally, the method results in one or more of: increase or decrease in expression of a specific gene, induction of cellular senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, cell growth promotion and / or cell growth inhibition in vitro or in vivo, induction of anergy in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo; Optionally, the method is a method for non-diagnostic and / or therapeutic purposes.

15. A method for diagnosing, treating or preventing a disease or condition associated with a target nucleic acid, characterized in that: Administering the Cas12 protein of claim 1 , the guide polynucleotide of claim 2 , the inactivated Cas12 variant of claim 3 , the fusion protein or conjugate of claim 4 , the nucleic acid of claim 5 , the CRISPR-Cas12 system of claim 6 , the vector system of claim 7 , the delivery system of claim 8 , the cell of claim 9 , the pharmaceutical composition of claim 10 , or the kit of claim 11 to a sample of a subject in need or to a subject in need; Optionally, the disease or condition is a blood disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer or an infectious disease; Optionally, the disease or condition is selected from the diseases or conditions listed in Table 27; Optionally, the disease or disorder is selected from: hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidosis, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosidosis, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bieder syndrome, white punctate retinal degeneration, leukocyte adhesion deficiency type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti crystal dystrophy, pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult glucan body disease, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophosphatasia, epilepsy, multiple Myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin lymphoma, non-muscle-invasive bladder cancer, non-alcoholic fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scars, obesity, Charcot-Marie-Tooth Disease Type 1A, Charcot-Marie-Tooth Disease Type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, muscle amyotrophic lateral sclerosis, urge incontinence, acute intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial Tay-Sachs disease, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, junctional epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, Canavan disease, cocaine addiction, Krabbe disease, Crigler-Najjar syndrome, oral cancer, happy puppet syndrome, diffuse intrinsic pontine glioma, Lafora disease, rheumatoid arthritis, sickle Cytopathy, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, anemia of chronic kidney disease, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Netherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Rett syndrome, triple-negative breast cancer, Sandhoff disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscinosis, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome , esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibulopathy, Stargardt's disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, head and neck squamous cell carcinoma, Wilson's disease, stable angina pectoris, Usher syndrome, choroideremia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, new coronavirus infection, pleural mesothelioma, acne vulgaris,Severe combined immunodeficiency, critical limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive dystrophic epidermolysis bullosa, malignant osteosclerosis of infancy, dystrophic epidermolysis bullosa, morphea, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy Muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic disease, lipoprotein lipase deficiency, severe congenital neutropenia, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, transthyretin (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.

16. The Cas12 protein of claim 1, the guide polynucleotide of claim 2, the inactivated variant of Cas12 of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, the delivery system of claim 8, the cell of claim 9, the pharmaceutical composition of claim 10 or the kit of claim 11, for use in diagnosing, treating or preventing a disease or disorder associated with a target nucleic acid; Optionally, the disease or condition is a blood disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer or an infectious disease; Optionally, the disease or condition is selected from the diseases or conditions listed in Table 27; Optionally, the disease or condition is selected from: hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, Blood disease, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidosis, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, alpha-1-antitrypsin deficiency, alpha-mannosidosis, alpha-thalassemia, beta-thalassemia, Alzheimer's disease, Budd-Bieder syndrome, white punctate retinal degeneration, leukocyte adhesion deficiency type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti crystal dystrophy, pyruvate kinase deficiency, erectile dysfunction Functional impairment, autosomal recessive congenital ichthyosis, adult glucan body disease, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophosphatasia, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin lymphoma, non-muscle-invasive bladder cancer, non-alcoholic fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, Charcot-Marie-Tooth disease type 1A, Charcot-Marie-Tooth disease type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge urinary incontinence, acute intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress 1 type, spinal muscular atrophy, familial Tay-Sachs disease, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, junctional epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, Canavan disease, cocaine addiction, Krabbe disease, Crigler-Najjar syndrome, oral cancer, happy puppet syndrome, diffuse intrinsic pontine glioma, Lafora disease, rheumatoid arthritis,sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, anemia with chronic kidney disease, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Netherton syndrome, ornithine carbamyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Rett syndrome, triple-negative breast cancer, Sandhoff disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscinosis storage disease, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Stargardt's disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease type IIb, atopic dermatitis, hearing loss, hearing impairment, head Neck cancer, head and neck squamous cell carcinoma, Wilson's disease, stable angina, Usher syndrome, choroideremia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, new coronavirus infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive dystrophic epidermolysis bullosa, infantile malignant osteosclerosis, dystrophic bullous epidermolysis, morphea, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic disease, lipoprotein lipase deficiency, severe congenital neutropenia, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, transthyretin (ATTR) amyloidosis, AATD liver disease and, AATD lung disease.

Citation Information

Patent Citations

  • Novel cas12b enzymes and systems

    CN113286884A

  • Novel Cas effector proteins, gene editing systems and their applications

    CN114934031A

  • Cas enzyme and application thereof

    CN116716277A

  • Cas protein and application thereof

    CN117683749A

  • Novel crispr DNA targeting enzymes and systems

    US20200063126A1