Cas12 proteins and uses thereof
Patent Information
- Application Number
- CN202480023951.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-01
- Filing Date
- 2024-09-19
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-09-19
Smart Images

Figure BSB0000212908380000931 
Figure BSB0000212908380000941 
Figure BSB0000212908380000951
Abstract
Description
[0001] This application claims priority to Chinese Patent Application No. 2023112143306, filed on September 19, 2023, and Chinese Patent Application No. 2024103885922, filed on April 1, 2024. The full text of the aforementioned Chinese patent applications is incorporated herein by reference. Technical Field
[0002] This disclosure pertains to the field of CRISPR gene editing, specifically to a Cas12 protein and its applications. Background Technology
[0003] The CRISPR-Cas system is an adaptive immune defense developed by bacteria and archaea over a long period of evolution, used to combat invading viruses and foreign DNA. Clustered, regularly spaced short palindromic repeats (CRISPR) and the CRISPR-associated protein system (CRISPR-Cas system) can directly alter gene sequences within cells, providing a rapid and effective method.
[0004] Many researchers in this field are working to find new Cas12 proteins and CRISPR-Cas12 gene editing systems. Summary of the Invention
[0005] This disclosure provides information about the Cas12 protein and its applications.
[0006] In one aspect, the technical solution provided in this disclosure is: a Cas12 protein, wherein the Cas12 protein is CLUSTER1 protein, CLUSTER2 protein, CLUSTER3 protein, CLUSTER4 protein, CLUSTER5 protein, CLUSTER6 protein, CLUSTER7 protein, CLUSTER8 protein, CLUSTER9 protein, CLUSTER10 protein, CLUSTER11 protein, CLUSTER12 protein or CLUSTER13 protein.
[0007] In another aspect, this disclosure provides a Cas12 protein whose amino acid sequence comprises or is an amino acid sequence having at least 50% sequence identity with any one of SEQ ID NO: 1-53, 696, 728.
[0008] In the specific embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with any of SEQ ID NO: 1-53, 696, 728.
[0009] In the specific embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of SEQ ID NO: 1-53, 696, 728.
[0010] In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 85% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 90% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 95% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 97% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 98% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99.5% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99.7% identity with any one of SEQ ID NO: 1-53, 696, and 728. In a specific embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 99.8% identity with any one of SEQ ID NO: 1-53, 696, and 728. In the specific embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is 100% identical to any one of SEQ ID NO: 1-53, 696, 728.
[0011] In the specific embodiments disclosed herein, the Cas12 protein retains the function of the protein shown in any of the sequences SEQ ID NO: 1-53, 696, 728.
[0012] In a specific embodiment disclosed herein, the Cas12 protein can form a complex with a guide polynucleotide. In a specific embodiment disclosed herein, the Cas12 protein can specifically bind to the target nucleic acid with the guide polynucleotide.
[0013] In a specific embodiment disclosed herein, the Cas12 protein can form a complex with a guide polynucleotide, and the complex can specifically bind to the target nucleic acid. In a specific embodiment disclosed herein, the Cas12 protein can form a complex with a guide polynucleotide, and the complex can specifically bind to the target DNA.
[0014] In specific embodiments disclosed herein, the Cas12 protein can specifically bind to and cleave target nucleic acids with a guide polynucleotide. In specific embodiments disclosed herein, the Cas12 protein can specifically bind to and cleave target DNA with a guide polynucleotide. In specific embodiments disclosed herein, the Cas12 protein can form a complex with the guide polynucleotide, and this complex can specifically bind to and cleave target nucleic acids. In specific embodiments disclosed herein, the Cas12 protein can form a complex with the guide polynucleotide, and this complex can specifically bind to and cleave target DNA.
[0015] In this disclosure, the function of retaining the protein shown in any of the sequences SEQ ID NO: 1-53, 696, 728 refers to retaining the ability to form a complex with the guide polynucleotide, retaining the ability to bind to a target nucleic acid complementary to the guide sequence of the guide polynucleotide, retaining the ability to target and cleave the target nucleic acid with the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.
[0016] In the specific embodiments disclosed herein, the function of retaining the protein as shown in any of the sequences SEQ ID NO: 1-53, 696, 728 is to retain and guide the formation of complexes with polynucleotides.
[0017] In the specific implementation of this disclosure, the function of retaining the protein as shown in any of the sequences SEQ ID NO: 1-53, 696, 728 is to retain the ability to bind to target nucleic acids complementary to the guide sequence of the guide polynucleotide.
[0018] In the specific implementation of this disclosure, the function of retaining the protein as shown in any of the sequences SEQ ID NO: 1-53, 696, 728 is to retain and guide the polynucleotide to target and cleave the target nucleic acid.
[0019] In the specific embodiments disclosed herein, the function of retaining the protein as shown in any of the sequences SEQ ID NO: 1-53, 696, 728 is to retain the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.
[0020] In a preferred embodiment disclosed herein, the amino acid sequence of the Cas12 protein comprises or is any of the amino acid sequences shown in SEQ ID NO: 1-53, 696, 728.
[0021] In the specific embodiments disclosed herein, the PAM sequence (5′→3′) recognizable by the Cas12 protein is selected from any one or more of the following:
[0022] A, C, T, G
[0023] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0024] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,
[0025] NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,
[0026] N is A, T, C, or G.
[0027] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-T-3′.
[0028] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-G-3′.
[0029] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-A-3′.
[0030] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-C-3′.
[0031] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TA-3′.
[0032] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TC-3′.
[0033] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TG-3′.
[0034] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TT-3′.
[0035] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TN-3′.
[0036] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TTN-3′.
[0037] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TTT-3′.
[0038] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TTG-3′.
[0039] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TTC-3′.
[0040] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-TTA-3′.
[0041] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-WTN-3′.
[0042] In the specific implementation of this disclosure, the Cas12 protein can recognize PAM with the sequence 5′-ATN-3′.
[0043] N can be any one of A, T, C, and G. W is either A or T.
[0044] In some embodiments disclosed herein, the Cas12 protein is a Cas12 inactivating variant. In some embodiments disclosed herein, the Cas12 protein is a nuclease-inactivating variant. In some embodiments disclosed herein, the Cas12 protein is a dead Cas12 inactivating variant or a nickase Cas12 inactivating variant. Optionally, the RuvC domain of the Cas12 protein is inactivated.
[0045] In some embodiments of this disclosure, the Cas12 protein is selected from the active fragments constituting any of the Cas12 proteins described in this disclosure.
[0046] In some embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with any of SEQ ID NO: 46, 696, 52, 728.
[0047] Optionally, the Cas12 protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid.
[0048] Optionally, the Cas12 protein may form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the Cas12 protein; even further, the backbone sequence comprises or is a unidirectional repeat sequence; still further, the unidirectional repeat sequence comprises or is a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NO: 704, 529, 534.
[0049] Optionally, the backbone sequence does not include the tracrRNA sequence.
[0050] Optionally, the Cas12 protein can recognize PAM sequences of 5'-TTN-3' and / or 5'-TTNC-3'. N can be any of A, T, C, and G.
[0051] In some embodiments disclosed herein, a Cas12 protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 46.
[0052] Optionally, the Cas12 protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0053] Optionally, the Cas12 protein may form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the Cas12 protein; even further, the backbone sequence comprises or is a repetitive sequence; still further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in SEQ ID NO: 704.
[0054] In some embodiments disclosed herein, a Cas12 protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 696.
[0055] Optionally, the Cas12 protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid.
[0056] Optionally, the Cas12 protein may form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the Cas12 protein; even further, the backbone sequence comprises or is a repetitive sequence; still further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in SEQ ID NO: 704.
[0057] In some embodiments disclosed herein, a Cas12 protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 52.
[0058] Optionally, the Cas12 protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid.
[0059] Optionally, the Cas12 protein may form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the Cas12 protein; even further, the backbone sequence comprises or is a repetitive sequence; still further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in SEQ ID NO: 534.
[0060] In some embodiments disclosed herein, a Cas12 protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 728.
[0061] Optionally, the Cas12 protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid.
[0062] Optionally, the Cas12 protein may form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the Cas12 protein; even further, the backbone sequence comprises or is a repetitive sequence; still further, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in SEQ ID NO: 534.
[0063] In some embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 46.
[0064] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide contains a guide sequence that is inversely complementary to the target nucleic acid and a repetitive sequence.
[0065] Optionally, the repetitive sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 704.
[0066] Optionally, the complex binds to the target nucleic acid under the guidance of the guide sequence.
[0067] Optionally, the Cas12 protein can recognize PAM with the sequence 5'-TTN-3';
[0068] N can be any one of A, T, C, and G.
[0069] In some embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 696.
[0070] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide contains a guide sequence that is inversely complementary to the target nucleic acid and a repetitive sequence.
[0071] Optionally, the repetitive sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 704.
[0072] Optionally, the complex binds to the target nucleic acid under the guidance of the guiding sequence;
[0073] Optionally, the Cas12 protein can recognize PAM with the sequence 5'-TTN-3'.
[0074] In some embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 52.
[0075] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide contains a guide sequence that is inversely complementary to the target nucleic acid and a repetitive sequence.
[0076] Optionally, the repetitive sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 534.
[0077] Optionally, the complex binds to the target nucleic acid under the guidance of the guiding sequence;
[0078] Optionally, the Cas12 protein can recognize PAM with the sequence 5'-TTN-3'.
[0079] In some embodiments disclosed herein, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with SEQ ID NO: 728.
[0080] The Cas12 protein can form a complex with a guide polynucleotide; the guide polynucleotide contains a guide sequence that is inversely complementary to the target nucleic acid and a repetitive sequence.
[0081] Optionally, the repetitive sequence comprises or is a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 534.
[0082] Optionally, the complex binds to the target nucleic acid under the guidance of the guide sequence.
[0083] Optionally, the Cas12 protein can recognize PAM with the sequence 5'-TTN-3';
[0084] N can be any one of A, T, C, and G.
[0085] Optionally, the reverse complementarity can be partial or complete complementarity. In some embodiments, the guide sequence hybridizes with the target nucleic acid.
[0086] In some embodiments, the Cas12 protein is a mutant of the Cas protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728.
[0087] In some embodiments, the Cas12 protein is an inactivated variant of the Cas protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728.
[0088] In some embodiments, the Cas12 protein provided herein contains one, two, or more mutations compared to the Cas12 protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728, such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof. In some examples, the Cas12 protein, compared to the Cas12 protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728, contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 3 4, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.In some instances, the Cas12 protein, compared to the Cas12 protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728, comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 7 3, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide.
[0089] In some embodiments disclosed herein, the Cas12 protein has a mutation at any amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. Optionally, the mutation is a mutation to any other natural amino acid residue. Further optionally, the mutation is a mutation to residue R, H, K, or A. In some embodiments, the mutation is a mutation to residue R. In some embodiments, the mutation is a mutation to residue A. In some embodiments, the mutation is a mutation to residue H. In some embodiments, the mutation is a mutation to residue K.
[0090] In some embodiments disclosed herein, the Cas12 protein has mutations at amino acid residues corresponding to positions 1-41, 42-195, 196-290, 291-358, 359-479, 480-636, 637-689, 690-846, 847-884, 885-959, 960-1080, or 1081-1139 of the sequence shown in SEQ ID NO: 696.
[0091] In some embodiments disclosed herein, the Cas12 protein contains a mutation in the RuvC domain corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein contains a mutation at amino acid residues 637-689, 885-959, or 1081-1139 corresponding to the sequence shown in SEQ ID NO: 696.
[0092] In some embodiments disclosed herein, the Cas12 protein contains a mutation in the helical domain corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein contains mutations at amino acid residues 42-195, 291-479, or 690-846 corresponding to amino acid residues in the sequence shown in SEQ ID NO: 696.
[0093] In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 1-41 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 42-195 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 196-290 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 291-358 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 359-479 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 480-636 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 637-689 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 690-846 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 847-884 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 885-959 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 960-1080 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residues 1081-1139 corresponding to the sequence shown in SEQ ID NO: 696.
[0094] In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 1-41 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 42-195 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 196-290 of the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 291-358 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 359-479 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 480-636 of the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 637-689 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 690-846 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 847-884 of the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 885-959 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 960-1080 of the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has at least 50%, 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in SEQ ID NO: 696 at positions corresponding to amino acid residues 1081-1139 of the sequence shown in SEQ ID NO: 696.
[0095] In some embodiments disclosed herein, the Cas12 protein corresponds to, for example, SEQ ID. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 19, 21, 22, 23, of the sequence shown in NO: 696 24, 25, 26, 27, 28, 29, 30, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 46, 47, 48, 49, 50, 51, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 88, 8 9, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 101, 102, 103, 104, 105, 106, 108, 109, 110, 111, 112, 114, 115, 116, 117, 118, 119, 120, 121, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 151, 152, 153, 154, 155, 156, 1 57, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 169, 170, 171, 172, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 194, 195, 196, 197, 198, 199, 200, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 22 2, 223, 224, 225, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 242, 243, 244, 245, 247, 248, 249, 250, 251, 252, 253, 255, 256, 257, 258, 259, 260, 261, 262, 263, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 281, 282, 283, 284, 285, 286, 287, 288, 289,290、291、292、293、294、295、296、297、298、299、300、301、302、303、305、306、308、309、310、313、315、316、317、318、319、320、321、322、323、324、325、327、328、329、330、331、332、333、334、335、336、337、339、340、341、342、343、344、345、346、347、348、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、431、432、433、435、436、437、439、440、441、442、443、444、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、467、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、496、497、499、500、501、502、503、504、506、507、508、509、510、511、512、513、514、515、516、517、518、519、520、521、522、523、524、525、526、527、528、529、531、532、533、534、535、536、537、538、539、540、541、542、543、544、545、546、547、548、549、550、552、553、555、556、557、558、559、560、561、562、563、564、565、566、567、568、569、570、571、572、573、574、575、576、577、578、579、580、581、582、583、584、585、586、587、589、590、592、593、594、595、596、597、598、599、601、602、603、604、605、606、607、608、609、610、611、612、613、614、615、616、618、619、620、621、622、623、624、625、626、627、628、630、631、632、633、634、635、636、637、638、639、640、641、642、643、644、645、646、647、648、649、650、651、652、653、654、655、656、657、658、659、660、661、662、663、664、665、666、667、668、669、670、671、672、673、674、675、676、678、679、680、681、683、684、685、686、688、689、691、692、693、694、695、696、697、698、699、700、701、702、703、704、705、706、707、708、709、710、711、712、713、715、716、717、719、720、721、722、723、724、725、727、728、729、730、731、732、733、734、736、737、738、739、740、741、742、743、744、745、746、747、748、749、751、752、753、754、755、756、758、759、760、761、762、764、765、766、767、768、769、771、772、773、774、775、776、779、780、781、782、783、784、785、786、787、789、790、791、792、794、795、797、798、800、801、802、804、805、806、807、808、809、810、811、812、813、814、815、817、818、819、821、822、823、824、825、826、827、828、829、830、831、832、833、834、835、836、837、838、839、840、841、842、844、845、846、847、848、849、850、851、852、853、854、855、856、857、858、859、860、862、863、864、865、866、867、868、870、872、873、874、875、876、877、879、880、881、882、883、884、885、886、887、888、890、891、892、893、894、895、896、897、898、899、900、901、902、903、904、905、906、909、910、911、912、913、914、916、917、918、919、920、921、922、923、924、925、926、927、928、930、931、932、933、934、935、936、937、938、939、940、941、942、943、944、945、946、947、948、949、950、951、952、953、954、955、956、957、958、959、960、961、963、964、965、966、967、968、969、970、971、972、973、974、975、976、977、979、980、981、982、983、985、986、987、988、989、990、991、992、993、994、995、996、997、998、999、1000、1001、1002、1003、1004、1005、1006、1007、1008、1009、1010、1011、1012、1013、1014、1015、1016、1017、1018、1021、1023、1024、1025、1026、1027、1028、1029、1030、1031、1032、1033、1035、1036、1037、1038、1039、1040、1041、1042、1043、1044、1045、1046、1047、1048、1049、1050、1052、1053、1055、1056、1057、1058、1059、1060、1061、1062、1063、1064、1065、1066、1067、1068、1069、1070、1071、1072、1073、1074、1075、1076、1077、1078、1079、1080、1081、1082、1083、1084、1085、1086, 1087, 1088, 1089, 1090, 1091, 1092, 1093, 1094, 1095, 1096, 1097, 1098, 1100, 1101, 1102, 1103, 1104, 1105, 1107, 1108, 1109, 1110, 1111, 1112, 1113, 1115, 1116, 11 At least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, or at least twelve of the amino acid residues at positions 17, 1120, 1122, 1123, 1124, 1125, 1128, 1129, 1130, 1131, 1132, 1133, 1134, and 1135 have mutations. Optionally, the mutation is a mutation to any other natural amino acid residue. Further optionally, the mutation is a mutation to residue R, H, K, or A. In some embodiments, the mutation is a mutation to residue R. In some embodiments, the mutation is a mutation to residue A. In some embodiments, the mutation is a mutation to residue H. In some embodiments, the mutation is a mutation to residue K.
[0096] In some embodiments disclosed herein, the Cas12 protein is located at positions 1, 2, 3, 4, 5, 7, 10, 24, 30, 48, 51, 55, 58, 59, 66, 108, 118, 138, 141, 175, 178, 185, 186, 257, 333, 352, 356, 375, 376, 378, 379, 383, 397, 400, 416, 426, corresponding to the sequence shown in SEQ ID NO: 696. 443, 449, 456, 459, 462, 469, 484, 485, 509, 561, 597, 607, 609, 623, 638, 639, 640, 697, 722, 731, 733, 755, 758, 771, 773, 779, 781, 784, 785, 786, 789, 792, 794, 798 822, 823, 825, 826, 829, 830, 833, 834, 836, 842, 845, 846, 847, 850, 851, 853, 855, 856, 858, 859, 860, 866, 884, 892, 893, 900, 904, 926, 956, 985, 988, 989, 992, 99 3. At least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, or at least twelve of the amino acid residues at positions 3, 996, 1016, 1033, 1045, 1050, 1073, 1074, 1095, 1100, 1124, 1129, and 1132 have mutations. Optionally, the mutation is a mutation to residue R, H, or K. More preferably, the mutation is a mutation to residue R.
[0097] In some embodiments disclosed herein, the Cas12 protein is located at positions 12, 29, 35, 36, 40, 53, 57, 60, 64, 71, 72, 73, 75, 94, 95, 96, 97, 99, 137, 148, 149, 153, 164, 167, 171, 172, 174, 177, 190, 192, 194, 199, 204, 207, 208, 211, 215, 228, 232, 236, 238, 244, 248, 253, 256, 258, 261, 262, corresponding to the sequence shown in SEQ ID NO: 696. 275, 282, 286, 292, 298, 300, 302, 320, 324, 328, 332, 336, 339, 366, 373, 374, 384, 389, 393, 395, 415, 432, 436, 440, 453, 458, 460, 471, 472, 474, 506, 508, 510, 519, 523, 526, 528, 531, 534, 544, 550, 570, 571, 573, 592, 594, 596 613, 615, 616, 619, 622, 624, 626, 648, 650, 651, 652, 658, 661, 663, 684, 686, 689, 693, 704, 744, 783, 787, 800, 849, 854, 876, 888, 890, 891, 894, 912, 940, 941, 942, 945, 946, 948, 961, 971, 979, 987, 998, 1000, 1002, 1006, 10 At least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, or at least twelve of the amino acid residues at positions 13, 1015, 1035, 1042, 1048, 1049, 1052, 1053, 1057, 1058, 1063, 1079, 1081, 1082, 1083, 1084, 1085, 1086, 1087, 1088, 1089, 1090, and 1091 have mutations. Optionally, the mutation is a mutation to residue A.
[0098] In some embodiments disclosed herein, the Cas12 protein has mutations at any of the following amino acid residues corresponding to any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more positions corresponding to the sequence shown in SEQ ID NO: 696:
[0099] 133, G184, S185, Q186, G194, N195, G196, G197, N245, G256, L260, Y278, S2 85, Y316, H350, D352, A355, A356, C385, P386, H387, G390, K391, N392, D42 9. Q461, Q462, Q469, E485, S491, K521, P525, L611, K629, K631, N633, D841 , N898, K987, A988, G989, Q990, T991, D1010, E1013, A1136, K1138, T1139.
[0100] In some embodiments disclosed herein, the Cas12 protein has undergone a mutation of any one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, or more amino acids at the position corresponding to the sequence shown in SEQ ID NO: 696, as described in SEQ ID NO: 696:
[0101] I33R, G184R, S185R, Q186R, G194R, N195R, G196R, G197R, N245R, G256R, L260R, Y278R, S285R, Y316R, H350R, D352R, A3 55R, A356R, C385R, P386R, H387R, G390R, K391R, N392R, D429R, Q461R, Q462R, Q469R, E485R, S491R, K521R, P525R, L61 1R, K629R, K631R, N633R, D841R, N898R, K987R, A988R, G989R, Q990R, T991R, D1010R, E1013R, A1136R, K1138R, T1139R.
[0102] In some embodiments disclosed herein, the Cas12 protein has undergone any one, two, three, four, five, or more of the following amino acid mutation combinations (separated by commas) at the corresponding positions of the sequence shown in SEQ ID NO: 696: 186+352+1+426+846+858+860, 186+352+1+426+846+860, 186+352+1+5+426+858+860, 186+352+1+3+426+858+860, 186+352+3+426+860, 186+352+1+333+426+858+860, 186+352+1+426+485+858+860, 186+ 352+1+5+426+860, 186+352+3+426+858+860, 186+352+5+426+858+860, 186+352+333+426+858+860, 186+352+333+426+860, 186+352+426+846+860, 186+352+1+3+426+860, 5+426+860, 186+352+426+846+858+860, 186+352+1+426+485+860, 426+858+860, 7+426+ 858, 186+352+426+860, 186+352+426+485+858+860, 186+352+426+485+860, 426+846+858+860, 7+426+846, 186+352+860, 184+186+352+376+1132, 5+333+426, 5+426+858, 186+352+1+333+426+860, 184+186+352+3+107+426, 186+352+5+426, 333+376+426, 186+352+5, 2+5+846+858, 186+352+3+426, 5+846+858, 186+352+426+846, 186+352+7, 186+352+3+376+426, 186+352+376+426+860+865, 186+352+376+426, 3+846+860, 846+858+988, 3+426+858, 3+846+858, 3+858+860, 186+352+858, 184+186+352+860, 333+426, 184+186+352+3+376, 3+426+860, 846+860+988,3+860, 184+186+352+3+639, 186+352+333+426, 846+585+860, 186+352+426, 333+426+846, 333+426+485, 5+846+860, 3+846+988, 184+186+352+376, 3+333+426, 186+352+426+485, 333+426+858, 333+426+1132, 428+485, 186+352+333+352+376+426, 858+860, 186+352+376+426+485+ 860, 184+186+352+426, 186+352+639, 5+858+988, 3+858+988, 5+858, 3+858, 184+186+352+846, 184+186+352+639, 858+988, 184+186+352+426+1132, 186+352+1132, 184+186+352+5, 184+186+352+858, 858+860+1132, 3+5, 426+649, 186+352+426+485+860, 186+352+333, 184+186+35 2. 186+376, 846+860, 858+988+1132, 846+858, 333+376, 376+426, 184+186+352+3, 3+846+1132, 5+846+1132, 186+352+426+1132, 376+426+485+660, 426, 5+846, 846+860+1132, 333+376+485, 184+186+352+639+1132, 352+426, 333+485, 184+186+352+333, 846+858+1132, 333+426+86 0, 186+352+988, 5+860, 846+988, 186+352, 3+846, 846+1132, 184+186+352+1132, 186+485, 988+1132, 184+186+352+485, 376+485, 5+1132, 3+7, 186+352+485, 184+186+352+7, 184+186+352+333+336, 3+1132, 426+858+988, 186+352+376, 186+352+3 and 186+352+333+336+352+376+426. Further, optionally, the mutation is a mutation to residue R or A.
[0103] In some embodiments disclosed herein, the Cas12 protein has any one of the following amino acid mutation combinations (separated by commas) at the position corresponding to the sequence shown in SEQ ID NO: 696: 186+352+1+426+846+858+860, 186+352+1+426+846+860, 186+352+1+5+426+858+860, 186+352+1+3+426+858+860, 186+35 2+3+426+860, 186+352+1+333+426+858+860, 186+352+1+426+485+858+860, 186+352+1+5+426+860, 186+352+3+426+858+860, 186+352+5+426+858+860, 186+352+333+42 6+858+860, 186+352+333+426+860, 186+352+426+846+860, 186+352+5+426+860, 186+352+1+3+426+860, 5+426+860, 186+352+426+846+858+860, 186+352+1+426+485+86 0, 426+858+860, 7+426+858, 186+352+426+860, 186+352+426+485+858+860, 186+352+426+485+860, 426+846+858+860, 7+426+846, 186+352+860; Further optionally, the mutation is a mutation to residue R or A.
[0104] In some embodiments disclosed herein, the Cas12 protein has undergone any one, two, three, four, five, or more of the following amino acid mutations at the corresponding position of the sequence shown in SEQ ID NO: 696:
[0105] D352R+Q186R, D352R+L260R, D352R+A355R, A355R+L260R, P386R+C385R, E485R+Q462R, D352R+Q186R, A355R+L260R, G184R+R186Q, D352R+Q186R+I33R, D352R+Q186R+G184R, D352R+Q186R+S185R, D352R+Q186R+G256R, D352R+Q186R+Y278R, D352R+Q186R+S285R, D352R+Q186R+Y316R, D352R+Q186R+H350R, D352R+Q186R+A356R, D352R+Q186R+Q469R, D352R+Q186R+S491R, D352R+Q186R+K521R, D352R+Q186R+P525R, D352R+Q186R+K629R, D352R+Q186R+N633R, D352R+Q186R+D841R, D352R+Q186R+N898R, D352R+Q186R+K987R, D352R+Q186R+T991R, D352R+Q186R+D1010R, D352R+Q186R+E1013R.
[0106] In some embodiments disclosed herein, the Cas12 protein has a mutation at the first amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at the second amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at the third amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at the fourth amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at the fifth amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at the seventh amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at the tenth amino acid residue corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 24 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 30 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 48 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 51 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 55 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 58 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 59 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 66 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 108 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 118 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 138 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 141 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 175 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 178 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 185 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 186 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 257 corresponding to the sequence shown in SEQ ID NQ: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 333 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 352 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 356 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 375 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 376 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 378 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 379 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 383 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 397 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 400 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 416 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 426 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 443 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 449 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 456 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 459 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 462 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 469 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 484 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 485 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 509 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 561 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 597 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 607 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 609 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 623 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 638 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 639 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 640 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 697 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 722 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 731 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 733 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 755 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 758 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 771 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 773 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 779 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 781 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 784 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 785 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 786 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 789 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 792 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 794 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 798 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 822 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 823 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 825 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 826 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 829 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 830 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 833 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 834 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 836 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 842 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 845 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 846 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 847 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 850 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 851 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 853 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 855 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 856 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 858 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 859 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 860 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 866 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 884 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 892 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 893 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 900, corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 904, corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 926 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 956 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 985 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 988 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 989 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 992 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 993 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 996 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1016 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1033 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1045 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1050 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1073 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1074 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1095 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1100 corresponding to the sequence shown in SEQ ID NO: 696.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1124 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1129 corresponding to the sequence shown in SEQ ID NO: 696. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1132 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is a mutation to residue R.
[0107] In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 651 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is a mutation to residue A.
[0108] In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 891 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is a mutation to residue A.
[0109] In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 1082 corresponding to the sequence shown in SEQ ID NO: 696. Further optionally, the mutation is a mutation to residue A.
[0110] In some embodiments disclosed herein, the Cas12 protein has mutations at any of the following amino acid residues corresponding to any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more positions corresponding to the sequence shown in SEQ ID NO: 52:
[0111] V15, Q172, A173, G182, E183, G184, K185, K186, G239, V243, D264, E271, Y295, L297, N317, T329, E331, I335, K339, K339, N347, E363, H366, V426, K429, S430, L433, S452, S455, S465, E493, P497, K587, E768, T825, A911, G914, K915, I916, K918, T919, T920, A922, E940.
[0112] In some embodiments disclosed herein, the Cas12 protein has undergone a mutation of any one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, or sixteen or more amino acids at the position corresponding to the sequence shown in SEQ ID NO: 52, as shown in SEQ ID NO: 696:
[0113] V15R, Q172R, A173W, G182R, E183R, G184R, K185R, K186R, G239R, V243R, D264R, E271R, Y295R, L297R, N317R, T329R, E331R, I335R, K339R, K339R, N347E, E363R, H366R, V426R, K429R, S430R, L433R, S452R, S455R, S465R, E493R, P497R, K587R, E768R, T825R, A911R, G914R, K915R, 1916R, K918R, T919R, T920R, A922R, E940R.
[0114] In some embodiments disclosed herein, the Cas12 protein has undergone any one, two, three, four, five, or more of the following amino acid mutations at the corresponding position to the sequence shown in SEQ ID NO: 52, as shown in SEQ ID NO: 696:
[0115] N347E+K339R, Q172R+S452R, Q172R+T920R, Q172R+V426R, S452R+T920R, V426R+S452R, V426R+T920R.
[0116] In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 15 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 172 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 173 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 182 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 183 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 184 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 185 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 186 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 239 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 243 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 264 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 271 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 295 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 297 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 317 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 329 corresponding to the sequence shown in SEQ ID NO: 52.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 331 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 335 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 339 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 339 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 347 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 363 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 366 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 426 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 429 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 430 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 433 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 452 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 455 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 465 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 493 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 497 corresponding to the sequence shown in SEQ ID NO: 52.In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 587 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 768 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 825 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 911 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 914 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 915 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 916 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 918 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 919 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 920 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 922 corresponding to the sequence shown in SEQ ID NO: 52. In some embodiments disclosed herein, the Cas12 protein has a mutation at amino acid residue 940 corresponding to the sequence shown in SEQ ID NO: 52. Further optionally, the mutation is a mutation to residue R.
[0117] On the other hand, one technical solution provided in this disclosure is: a guide polynucleotide comprising (i) a homologous repeat sequence having at least 50% identity with any one of SEQ ID NO: 54-583, 704, and (ii) a guide sequence engineered to hybridize with a target nucleic acid; the homologous repeat sequence is linked to the guide sequence, and the guide polynucleotide is capable of forming a complex with the Cas12 protein and guiding the complex to bind to the target nucleic acid in a sequence-specific manner.
[0118] In some embodiments disclosed herein, the unidirectional repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0119] In some embodiments disclosed herein, the same-direction repeat sequence has at least 60% sequence identity compared to any one of SEQ ID NO: 54-583 and 704.
[0120] In some embodiments disclosed herein, the same-direction repeat sequence has at least 65% sequence identity compared to any one of SEQ ID NO: 54-583 and 704.
[0121] In some embodiments disclosed herein, the same-direction repeat sequence has at least 70% sequence identity compared to any one of SEQ ID NO: 54-583 and 704.
[0122] In some embodiments disclosed herein, the same-direction repeat sequence has at least 75% sequence identity compared to any one of SEQ ID NO: 54-583 and 704.
[0123] In some embodiments disclosed herein, the same-direction repeat sequence has at least 80% sequence identity compared to any one of SEQ ID NO: 54-583 and 704.
[0124] In some embodiments disclosed herein, the same-direction repeat sequence has at least 85% sequence identity compared to any one of SEQ ID NO: 54-583 and 704.
[0125] In some embodiments disclosed herein, the same-direction repeat sequence has at least 90% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0126] In some embodiments disclosed herein, the same-direction repeat sequence has at least 95% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0127] In some embodiments disclosed herein, the same-direction repeat sequence has at least 96% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0128] In some embodiments disclosed herein, the same-direction repeat sequence has at least 97% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0129] In some embodiments disclosed herein, the same-direction repeat sequence has at least 98% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0130] In some embodiments disclosed herein, the same-direction repeat sequence has 100% sequence identity with any one of SEQ ID NO: 54-583 and 704.
[0131] In a preferred embodiment, the Cas12 protein is the Cas12 protein described in this disclosure.
[0132] In specific embodiments disclosed herein, the guide sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-35 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-30 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-22 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0133] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the guide sequence and the target nucleic acid are 90%-100% complementary.
[0134] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid.
[0135] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the mismatch between the guide sequence and the target nucleic acid does not exceed one nucleotide.
[0136] In specific embodiments disclosed herein, the directed repeat sequence comprises 15-100 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-90 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-80 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-70 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-30 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0137] In the specific implementation of this disclosure, the guiding sequence is located at the 3′ end of the same-direction repeating sequence.
[0138] In the specific implementation of this disclosure, the guiding sequence is located at the 5′ end of the same-direction repeating sequence.
[0139] In the specific implementation disclosed herein, the guiding polynucleotide further comprises tracrRNA.
[0140] In some embodiments disclosed herein, the tracrRNA sequence has at least 50% identity with any one of SEQ ID NO: 584-695. In some embodiments disclosed herein, the tracrRNA sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NO: 584-695.
[0141] In some embodiments disclosed herein, the tracrRNA sequence may be selected from any one of the sequences shown in SEQ ID NO: 584-695.
[0142] In specific embodiments disclosed herein, the tracrRNA may pair complementaryly with the unidirectional repeat sequence. Typically, this complementary pairing is partial base pairing. In specific embodiments disclosed herein, the tracrRNA may interact with the unidirectional repeat sequence.
[0143] In specific embodiments disclosed herein, the tracrRNA sequence is linked to the same-direction repeat sequence. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence consisting of 1-10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence consisting of 4 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a 5'-GAAA-3' sequence.
[0144] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 3′ end of the same-direction repeat sequence.
[0145] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 5′ end of the same-direction repeat sequence.
[0146] In the specific embodiments disclosed herein, the tracrRNA comprises 10-200 nucleotides. In the specific embodiments disclosed herein, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, or 10-30 nucleotides. 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50 or 30-50 nucleotides. In the specific implementation scheme disclosed herein, the tracrRNA contains 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, and 5 5, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0147] SEQ ID NO: 1-53 shows the amino acid sequence of the Cas protein.
[0148] SEQ ID NO: 54-583 shows the direct repeat (DR) sequences corresponding to Cas proteins. When multiple DR sequences are listed for a particular Cas protein, one can be selected.
[0149] SEQ ID NO: 584-695 shows the tracrRNA sequences corresponding to Cas proteins. When SEQ ID NO: 584-695 does not list the tracrRNA sequence corresponding to a specific Cas protein, the gRNA may contain only the guide sequence and DR sequence, and not the tracrRNA sequence. When SEQ ID NO: 584-695 lists the tracrRNA sequence corresponding to a specific Cas protein, the gRNA may or may not contain the tracrRNA sequence in addition to the guide sequence and DR sequence. When multiple tracrRNA sequences corresponding to a specific Cas protein are listed, one may be selected.
[0150] On the other hand, one technical solution provided in this disclosure is: a Cas12 inactivation variant, characterized in that the Cas12 inactivation variant is a nuclease activity inactivation variant of the Cas12 protein as described in this disclosure.
[0151] In this document, depending on the context, the term "Cas12 protein" may refer to the inactivated Cas12 variants. However, given the importance of the inactivated Cas12 variants (non-limiting examples include their fusion with deaminases for single-base editing, their fusion with transcriptional activation or repression domains for transcriptional regulation, etc.), they will be described separately and with emphasis; this does not mean that the term "Cas12 protein" necessarily excludes the inactivated Cas12 variants.
[0152] In the specific embodiments disclosed herein, the Cas12 inactivating variant is a variant with completely inactivated nuclease activity, namely, the dead Cas12 inactivating variant (dCas12). The dCas12 can only bind to the target nucleic acid under the mediation of a guide polynucleotide, and has little or no ability to cleave the target nucleic acid. For example, the target nucleic acid cleavage efficiency of the dCas12 is ≤20%, ≤15%, ≤10%, ≤5%, ≤4%, ≤3%, ≤2%, or ≤1% of the target nucleic acid cleavage efficiency of the Cas12 protein before inactivation mutation.
[0153] In the specific embodiments disclosed herein, the Cas12 inactivating variant is a variant with partially inactivated nuclease activity. Further, the partially inactivated nuclease variant is a Cas12 nickase (nCas12), which binds to the target nucleic acid under the mediation of a guide polynucleotide and then cleaves one single strand of the double-stranded target nucleic acid without cleaving the other single strand.
[0154] In a preferred embodiment disclosed herein, the Cas12 inactivating variant is the inactivation of the Ruvc domain of the Cas12 protein.
[0155] In a preferred embodiment disclosed herein, the Cas12 inactivating variant is the inactivation of the Ruvc-I, Ruvc-II, or Ruvc-III domains of the Cas12 protein.
[0156] In a preferred embodiment disclosed herein, the Cas12 inactivating variant is obtained by introducing an inactivating mutation into the Ruvc-I, Ruvc-II, or Ruvc-III domain of the Cas12 protein.
[0157] In the specific embodiments disclosed herein, the inactivation mutation is selected from one or more of D651A, E891A, and D1082A corresponding to the amino acid sequence shown in SEQ ID NO: 696.
[0158] In the specific implementation of this disclosure, the inactivation mutations are D651A, E891A and D1082A corresponding to the amino acid sequences shown in SEQ ID NO: 696.
[0159] In the specific embodiments disclosed herein, the PAM sequence recognizable by the Cas12 inactivated variant is the same as the PAM sequence recognizable by the Cas12 protein.
[0160] In the specific embodiments disclosed herein, the PAM sequence (5′→3′) identifiable by the Cas12 inactivation variant is selected from any one or more of the following:
[0161] A, C, T, G
[0162] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0163] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT ANG、GNA、GTT、NGA、TAA、GTA、GGN、GNT、NCG、ATT、CCA、CNN、AAA、AAC、ATN、GA G、CTG、ACG、NAA、TAN、NAT、CNA、GCN、GTC、NCN、CTN、CNC、ANT、NNC、CAG、NAN、 ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG
[0164] NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,
[0165] N is A, T, C, or G.
[0166] On the other hand, one technical solution provided by this disclosure is: a fusion protein or conjugate comprising the following elements: (1) the Cas12 protein as described in this disclosure, or the Cas12 inactivated variant as described in this disclosure; and (2) homologous or heterologous functional domains.
[0167] In this document, depending on the context, the term "Cas12 protein" may refer to the inactivated Cas12 variant. However, given the importance of the inactivated Cas12 variant (non-limiting examples include the fusion of the inactivated Cas12 variant with a deaminase for single-base editing, or with a transcriptional activation domain or transcriptional repression domain for transcriptional regulation, etc.), it will be described separately and with emphasis; this does not mean that the term "Cas12 protein" necessarily excludes the inactivated Cas12 variant. In some embodiments of this disclosure, a fusion protein is provided comprising: (1) the Cas12 protein as described herein, or the inactivated Cas12 variant as described herein; and (2) homologous or heterologous functional domains.
[0168] In some embodiments disclosed herein, a fusion protein is provided, the fusion protein comprising: (1) the Cas12 protein as described herein; and (2) homologous or heterologous functional domains.
[0169] In a specific embodiment of this disclosure, a conjugate is provided, the conjugate comprising: (1) the Cas12 protein as described in this disclosure, or the Cas12 inactivating variant as described in this disclosure; and (2) a homologous or heterologous functional domain.
[0170] In a specific embodiment of this disclosure, a conjugate is provided, the conjugate comprising: (1) the Cas12 protein as described in this disclosure; and (2) a homologous or heterologous functional domain.
[0171] In some embodiments, the functional domain has enzymatic activity that modifies the target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristylation activity, and / or demyristylation activity.
[0172] In the specific embodiments disclosed herein, the inactivation mutation is selected from one or more of D651A, E891A, and D1082A corresponding to the amino acid sequence shown in SEQ ID NO: 696.
[0173] In the specific implementation of this disclosure, the inactivation mutations are D651A, E891A and D1082A corresponding to the amino acid sequences shown in SEQ ID NO: 696.
[0174] In some embodiments, the functional domain is optionally selected from one or more of the following: nucleases (e.g., FokI), methyltransferases, demethylases, DNA repair enzymes, DNA damage enzymes, deaminases, superoxide dismutases, alkylating enzymes, depurinases, oxidases, pyrimidine dimer forming enzymes, integrases, transposases, recombinases, polymerases, ligases, helicases, photolyases, glycosylation enzymes, deglycosylation enzymes, acetyltransferases, deacetylases, kinases, phosphatases, ubiquitin ligases, deubiquitinating enzymes, adenylate acylases, deadenylate acylases, SUMOylating enzymes, deSUMOylating enzymes, myristylases, and / or demyristylases.
[0175] In the specific implementation scheme disclosed herein, the homologous or heterologous functional domains are selected from one, two, three, four or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase repression domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, affinity tag, reporter tag, affinity domain and reporter domain.
[0176] In some embodiments disclosed herein, the subcellular localization signal may be selected from: nuclear localization signal, nuclear output signal, mitochondrial localization signal, or chloroplast localization signal.
[0177] In the specific embodiments disclosed herein, the fusion protein or conjugate comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or more of the homologous or heterologous functional domains; the functional domains may be the same or different.
[0178] In some embodiments, the fusion protein or conjugate is arbitrarily linked to 0, 1, 2, 3, 4, 5, 6, 7, 8 or more of the functional domains at the N-terminus and / or C-terminus of the Cas12 protein.
[0179] In the specific implementation scheme disclosed herein, the fusion protein contains one, two, three, four or more nuclear localization signals.
[0180] In specific embodiments disclosed herein, the fusion protein can be used to achieve base editing, for example, in conjunction with a guide polynucleotide to achieve base editing. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a deaminase domain.
[0181] In the specific embodiments disclosed herein, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and optionally one or two UGI domains. The fusion protein can be used to achieve C→T base editing of target nucleic acids.
[0182] In the specific implementation disclosed herein, the fusion protein includes a nuclear localization signal and an adenosine deaminase domain. The fusion protein can be used to achieve A→G base editing of target nucleic acids.
[0183] In the specific embodiments disclosed herein, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and an adenosine deaminase domain. In the specific embodiments disclosed herein, the fusion protein comprises one, two, or three nuclear localization signals and a deaminase domain. In the specific embodiments disclosed herein, the fusion protein comprises a UGI domain. In the specific embodiments disclosed herein, the fusion protein comprises one, two, or three nuclear localization signals, a deaminase domain, and one or two UGI domains.
[0184] In specific embodiments disclosed herein, the fusion protein can be used to achieve transcriptional activation of a specific target gene, for example, in combination with a guide polynucleotide to achieve transcriptional activation of a specific target gene. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a transcription activation domain.
[0185] In specific embodiments disclosed herein, the fusion protein can be used to achieve transcriptional repression of specific target genes, for example, in combination with a guide polynucleotide to achieve transcriptional repression of specific target genes. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a transcriptional repression domain.
[0186] In specific embodiments disclosed herein, the fusion protein can be used to achieve methylation of a specific target sequence, for example, in conjunction with a guide polynucleotide to achieve methylation of a specific target sequence. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a DNA methylation domain.
[0187] In specific embodiments disclosed herein, the fusion protein can be used to achieve demethylation of specific target sequences, for example, in conjunction with a guide polynucleotide to achieve demethylation of specific target sequences. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a DNA demethylation domain.
[0188] In a preferred embodiment disclosed herein, the nuclease domain includes a polypeptide having ssDNA cleavage activity and / or a polypeptide having dsDNA cleavage activity.
[0189] In a preferred embodiment disclosed herein, the nuclease domain includes a polypeptide having ssDNA cleavage activity.
[0190] In a preferred embodiment disclosed herein, the nuclease domain comprises a polypeptide having dsDNA cleavage activity.
[0191] In the specific embodiments disclosed herein, the Cas12 protein or its inactivated variant is directly or indirectly connected to the homologous or heterologous functional domain.
[0192] In the preferred embodiment disclosed herein, the direct connection is a covalent connection, and the indirect connection is a connection via an amino acid linker or a non-amino acid linker.
[0193] In a preferred embodiment disclosed herein, the homologous or heterologous functional domains are fused or conjugated at the N-terminus, C-terminus, or internally relative to the Cas12 protein or inactivated variant.
[0194] In this disclosure, the fusion protein refers to the connection between the element (1) and the element (2) via peptides or direct connection; the conjugate refers to the connection between the element (1) and the element (2) via non-peptide chemical bonds.
[0195] In the specific embodiments disclosed herein, the PAM sequence recognizable by the fusion protein or conjugate is the same as the PAM sequence recognizable by the Cas12 protein.
[0196] In the specific embodiments disclosed herein, the PAM sequence (5′→3′) recognizable by the fusion protein or conjugate is selected from any one or more of the following:
[0197] A, C, T, G
[0198] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0199] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,
[0200] NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,
[0201] N is A, T, C, or G.
[0202] In the specific implementation of this disclosure, the fusion protein can recognize PAM with the sequence 5′-TTN-3′.
[0203] In the specific implementation of this disclosure, the fusion protein can recognize PAM with the sequence 5′-TTNC-3′.
[0204] In the specific implementation of this disclosure, the fusion protein can recognize PAM with the sequence 5′-WTN-3′.
[0205] In the specific implementation of this disclosure, the fusion protein can recognize PAM with the sequence 5′-ATN-3′.
[0206] In the specific implementation of this disclosure, the conjugate can recognize PAM with the sequence 5′-TTN-3′.
[0207] In the specific implementation of this disclosure, the conjugate can recognize PAM with the sequence 5′-TTNC-3′.
[0208] In the specific implementation of this disclosure, the conjugate can recognize PAM with the sequence 5′-WTN-3′.
[0209] In the specific implementation of this disclosure, the conjugate can recognize PAM with the sequence 5′-ATN-3′.
[0210] On the other hand, one technical solution provided in this disclosure is: an isolated nucleic acid that encodes the Cas12 protein as described in this disclosure, the Cas12 inactivated variant as described in this disclosure, or the fusion protein or conjugate as described in this disclosure.
[0211] In some embodiments disclosed herein, the nucleic acid encodes the Cas12 protein or the fusion protein as described herein.
[0212] In a preferred embodiment disclosed herein, the nucleic acid is codon-optimized for expression in cells.
[0213] In a preferred embodiment of this disclosure, the nucleic acid is codon-optimized for expression in eukaryotes, mammals such as humans or non-human mammals, plants, insects, birds, reptiles, rodents (e.g., mice, rats), fish, worms / nematodes, or yeast.
[0214] On the other hand, one technical solution provided in this disclosure is: a CRISPR-Cas12 system, the CRISPR-Cas12 system comprising:
[0215] a. The Cas12 protein as described in this disclosure, the Cas12 inactivating variants as described in this disclosure, the fusion proteins or conjugates as described in this disclosure, or the isolated nucleic acids as described in this disclosure; and
[0216] b. A guiding polynucleotide, or a polynucleotide sequence encoding the guiding polynucleotide;
[0217] The Cas12 protein, the inactivated variant of Cas12, or the fusion protein or conjugate forms a complex with the guiding polynucleotide; the guiding polynucleotide contains a guiding sequence that is engineered to guide the complex to bind to the target nucleic acid in a sequence-specific manner.
[0218] The isolated nucleic acid encodes the Cas12 protein as described in this disclosure, the Cas12 inactivated variant as described in this disclosure, or the fusion protein or conjugate as described in this disclosure.
[0219] In a specific embodiment disclosed herein, the guiding polynucleotide comprises a homologous repeat sequence linked to the guiding sequence.
[0220] In the specific implementation of this disclosure, the same-direction repeat sequence has at least 50% identity with any one of SEQ ID NO: 54-583 and 704.
[0221] In specific embodiments disclosed herein, the guiding polynucleotide comprises a homologous repeat sequence linked to the guiding sequence. Further, in some embodiments, the homologous repeat sequence has at least 50% identity with any one of SEQ ID NO: 54-583, 704. In some embodiments, the homologous repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in any one of SEQ ID NO: 54-583, 704. Furthermore, in some specific embodiments, the same-direction repeating sequence includes or is the sequence shown in any one of SEQ ID NO: 54-583, 704.
[0222] In specific embodiments disclosed herein, the guide sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-35 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-30 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-22 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-22 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides.
[0223] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the guide sequence and the target nucleic acid are 90%-100% complementary.
[0224] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid.
[0225] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the mismatch between the guide sequence and the target nucleic acid does not exceed one nucleotide.
[0226] In specific embodiments disclosed herein, the directed repeat sequence comprises 15-100 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-90 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-80 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-70 nucleotides. In specific embodiments disclosed herein, the directed repeat sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-30 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0227] In the specific implementation of this disclosure, the guiding sequence is located at the 3′ end of the same-direction repeating sequence.
[0228] In the specific implementation of this disclosure, the guiding sequence is located at the 5′ end of the same-direction repeating sequence.
[0229] In the specific implementation disclosed herein, the guiding polynucleotide further comprises tracrRNA.
[0230] In some embodiments disclosed herein, the tracrRNA sequence has at least 50% identity with any one of SEQ ID NO: 584-695. In some embodiments disclosed herein, the tracrRNA sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any one of SEQ ID NO: 584-695.
[0231] In specific embodiments disclosed herein, the tracrRNA may pair complementaryly with the unidirectional repeat sequence. Typically, this complementary pairing is partial base pairing. In specific embodiments disclosed herein, the tracrRNA may interact with the unidirectional repeat sequence.
[0232] In specific embodiments disclosed herein, the tracrRNA sequence is linked to the same-direction repeat sequence. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence consisting of 1-10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a nucleotide sequence consisting of 4 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the same-direction repeat sequence are linked via a 5'-GAAA-3' sequence.
[0233] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 3′ end of the same-direction repeat sequence.
[0234] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 5′ end of the same-direction repeat sequence.
[0235] In the specific embodiments disclosed herein, the tracrRNA comprises 10-200 nucleotides. In the specific embodiments disclosed herein, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10- 40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In the specific implementation scheme disclosed herein, the tracrRNA contains 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, and 5 5, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0236] In a preferred embodiment of this disclosure, the guiding polynucleotide is the guiding polynucleotide as described in this disclosure.
[0237] In the specific implementation scheme disclosed herein, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA.
[0238] In the preferred embodiment disclosed herein, the DNA is eukaryotic DNA; preferably, the eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent DNA, fish DNA, worm / nematode DNA, or yeast DNA.
[0239] In the specific implementation scheme disclosed herein, the target nucleic acid is a disease or symptom-related gene or a gene related to signal transduction biochemical pathways, or the target nucleic acid is a reporter gene; for example, the disease or symptom is a hematological disease or symptom, an ophthalmic disease or symptom, a nervous system disease or symptom, a respiratory disease or symptom, a liver disease or symptom, a metabolic system disease or symptom, cancer, or an infectious disease.
[0240] In some implementations, the target nucleic acid is a gene listed in Table 27.
[0241] In some embodiments, the target nucleic acid is a disease or symptom-related gene, wherein the disease or symptom is selected from: hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes mellitus, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Bader-Bieder syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens camp Malnutrition, pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle invasive tumors. Infiltrative bladder cancer, non-alcoholic fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis (ALS). Acute urinary incontinence, acute intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy.Canavan disease, cocaine addiction, Krabby's disease, Krieger-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, meat Tumors, breast cancer, Ritter syndrome, triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease I Type A glycogen storage disease, type IIb glycogen storage disease, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal dysplasia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignancy Osteosclerosis, nutritional epidermolysis bullosa, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0242] In some embodiments, the genes associated with the thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0243] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0244] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0245] The genes related to AATD lung disease include, but are not limited to, AATD;
[0246] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0247] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0248] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0249] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0250] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0251] The genes related to hemophilia B include, but are not limited to, factor IX;
[0252] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0253] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0254] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0255] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0256] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0257] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0258] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0259] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0260] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0261] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0262] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0263] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0264] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0265] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0266] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B;
[0267] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0268] The epilepsy-related genes include, but are not limited to, GAT1.
[0269] On the other hand, one technical solution provided in this disclosure is: a vector system comprising one or more recombinant vectors, wherein the recombinant vectors comprise isolated nucleic acids as described in this disclosure, or a CRISPR-Cas12 system as described in this disclosure.
[0270] In the specific implementation scheme disclosed herein, the recombinant vector further includes a regulatory sequence.
[0271] In specific embodiments disclosed herein, the vector system comprises one or more recombinant vectors, the recombinant vectors comprising a multinucleotide sequence encoding the Cas12 protein, Cas12 inactivation variant, fusion protein, or conjugate disclosed herein, and a multinucleotide sequence encoding the guide multinucleotide.
[0272] In a specific embodiment disclosed herein, the polynucleotide sequence encoding the Cas12 protein, Cas12 inactivation variant, fusion protein, or conjugate is operatively linked to regulatory sequence 1.
[0273] In a specific implementation of this disclosure, the polynucleotide sequence encoding the guide polynucleotide is operatively linked to regulatory sequence 2.
[0274] Furthermore, in the specific implementation scheme disclosed herein, the regulatory sequence 1 and the regulatory sequence 2 may be the same or different sequences.
[0275] In a preferred embodiment disclosed herein, the regulatory sequence is selected from one or more of a promoter, enhancer, internal ribosome entry site, and transcription termination signal; the promoter is, for example, a constitutive promoter, an inducible promoter, a broad-spectrum promoter, or a tissue-specific promoter; and / or, the transcription termination signal is, for example, a polyadenylation signal or a polyU sequence.
[0276] In the specific implementation scheme disclosed herein, the backbone of the recombinant vector is an adeno-associated virus vector, a lentiviral vector, or a virus-like particle.
[0277] In the preferred embodiment disclosed herein:
[0278] When the backbone is an adeno-associated virus vector, the adeno-associated virus vector is a recombinant adeno-associated virus vector of serotype AAV1, AAV2, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12 or AAV13.
[0279] When the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; preferably, the isolated nucleic acid is linked to an aptamer sequence.
[0280] When the backbone is a virus-like particle, the isolated nucleic acid is linked to a gene encoding a gag protein.
[0281] On the other hand, one technical solution provided by this disclosure is: a delivery system comprising: (1) a delivery tool, and (2) a Cas12 protein as described in this disclosure, a guide polynucleotide as described in this disclosure, a Cas12 inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, a CRISPR-Cas12 system as described in this disclosure, or a vector system as described in this disclosure.
[0282] In a preferred embodiment disclosed herein, the delivery tool is a virus, lipid nanoparticles, nanoparticles, liposomes, exosomes, microvesicles, or a gene gun.
[0283] In a preferred embodiment disclosed herein, the delivery tool is a lipid nanoparticle comprising the guiding polynucleotide and mRNA encoding the Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate.
[0284] On the other hand, one technical solution provided in this disclosure is: a cell comprising the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, or the vector system as described in this disclosure.
[0285] In some embodiments disclosed herein, the cells are prokaryotic cells.
[0286] In some embodiments disclosed herein, the cells are eukaryotic cells.
[0287] In some embodiments disclosed herein, the eukaryotic cells are mammalian cells.
[0288] On the other hand, one technical solution provided in this disclosure is: a pharmaceutical composition comprising the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, or the cell as described in this disclosure.
[0289] Preferably, the pharmaceutical composition comprises pharmaceutically acceptable excipients.
[0290] On the other hand, one technical solution provided in this disclosure is: a kit comprising the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, or the cell as described in this disclosure.
[0291] In a preferred embodiment disclosed herein, the kit further comprises a cut buffer. The cut buffer may be any buffer known in the art suitable for Cas12 protein to cleave target nucleic acids.
[0292] On the other hand, one technical solution provided in this disclosure is the use of the Cas12 protein, the guide polynucleotide, the Cas12 inactivating variant, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas12 system, the vector system, the delivery system, the cell, the pharmaceutical composition, or the kit as described in this disclosure in the preparation of reagents or drugs for the diagnosis, treatment, and / or prevention of diseases or conditions related to the target nucleic acid.
[0293] In the specific implementation scheme disclosed herein, the disease or condition is a hematological disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory disease or condition, a liver disease or condition, a metabolic disease or condition, cancer, or an infectious disease; and / or, the reagent or drug is used to: cleave one or more target nucleic acid molecules or create nicks in one or more target nucleic acid molecules, activate or upregulate the expression of one or more target nucleic acid molecules, activate or inhibit the transcription of one or more target nucleic acid molecules, inactivate one or more target nucleic acid molecules, visualize, label, or detect one or more target nucleic acid molecules, bind one or more target nucleic acid molecules, transport one or more target nucleic acid molecules, and mask one or more target nucleic acid molecules.
[0294] In the specific implementation of this disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or symptom is the disease or symptom listed in Table 27.
[0295] In the specific implementation scheme disclosed herein, the diseases or conditions mentioned are selected from: hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes mellitus, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC 13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Bader-Bieder syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens dystrophy, pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, pure Zygote familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle-invasive bladder cancer, non-alcoholic fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scars, obesity. Peroneal muscular atrophy type 1A, Peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent porphyria, acute lymphoblastic leukemia. Leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, canavan disease, cocaine addiction.Krabby's disease, Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Rhett Butler's disease. Syndrome, triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen Storage disease type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, and infantile malignant osteosclerosis. Diseases, nutritional epidermolysis bullosa, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0296] In the specific implementation scheme disclosed herein, the genes related to thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0297] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0298] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0299] The genes related to AATD lung disease include, but are not limited to, AATD;
[0300] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0301] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0302] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0303] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0304] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0305] The genes related to hemophilia B include, but are not limited to, factor IX;
[0306] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0307] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0308] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0309] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0310] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0311] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0312] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAs, STING, mTOR, GCN2, Nrf2, Ang 2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0313] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0314] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0315] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0316] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0317] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0318] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0319] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0320] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B;
[0321] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0322] The epilepsy-related genes include, but are not limited to, GAT1.
[0323] On the other hand, one technical solution provided in this disclosure is: a method for detecting, binding to, or cleaving target nucleic acids, the method comprising contacting the target nucleic acid with a Cas12 protein as described in this disclosure, a guide polynucleotide as described in this disclosure, a Cas12 inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, a CRISPR-Cas12 system as described in this disclosure, a vector system as described in this disclosure, a delivery system as described in this disclosure, a cell as described in this disclosure, a pharmaceutical composition as described in this disclosure, or a kit as described in this disclosure.
[0324] In a preferred embodiment disclosed herein, the method is a method for non-diagnostic and / or therapeutic purposes; and / or the fusion protein or conjugate contains a detectable marker, such as a marker detectable by fluorescence, DNA blotting, or FISH.
[0325] In a preferred embodiment disclosed herein, when the method involves cleaving the target nucleic acid, the method further includes performing the cleavage reaction using a cut buffer. The cut buffer can be any buffer known in the art suitable for Cas12 protein cleavage of the target nucleic acid.
[0326] On the other hand, one technical solution provided in this disclosure is: a method for altering cell state, the method comprising contacting cells with a Cas12 protein as described in this disclosure, a guide polynucleotide as described in this disclosure, a Cas12 inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, a CRISPR-Cas12 system as described in this disclosure, a vector system as described in this disclosure, a delivery system as described in this disclosure, cells as described in this disclosure, a pharmaceutical composition as described in this disclosure, or a kit as described in this disclosure, thereby altering the cell state.
[0327] In some embodiments disclosed herein, the method results in one or more of the following: increased or decreased expression of a specific gene, induction of cell senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, promotion and / or inhibition of cell growth in vitro or in vivo, induction of non-responsiveness in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo.
[0328] In a preferred embodiment disclosed herein, the method is a method for non-diagnostic and / or therapeutic purposes.
[0329] On the other hand, one technical solution provided in this disclosure is: a method for diagnosing, treating, or preventing diseases or conditions related to target nucleic acids, comprising administering to a sample of a subject in need or to a subject in need the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, the cell as described in this disclosure, the pharmaceutical composition as described in this disclosure, or the kit as described in this disclosure.
[0330] In the specific implementation of this disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or symptom is the disease or symptom listed in Table 27.
[0331] In the specific implementation scheme disclosed herein, the disease or condition is a blood system disease or condition, an eye disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease.
[0332] On the other hand, one technical solution provided in this disclosure is: the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, the cell as described in this disclosure, the pharmaceutical composition as described in this disclosure, or the kit as described in this disclosure, for the diagnosis, treatment, or prevention of diseases or conditions related to the target nucleic acid.
[0333] In the specific implementation of this disclosure, the target nucleic acid is selected from the genes listed in Table 27, and the disease or symptom is the disease or symptom listed in Table 27.
[0334] In the specific implementation scheme disclosed herein, the disease or condition is a blood system disease or condition, an eye disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease.
[0335] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain the preferred examples disclosed herein. Attached Figure Description
[0336] Figure 1A shows the evolutionary relationship between the disclosed Cas protein and known Cas12 isoform proteins (sequence alignment was performed using MAFFT, and then the evolutionary tree was constructed using FastTree).
[0337] In the phylogenetic tree, the proteins disclosed herein form a distinct and clearly separate branch (a different cluster) compared to known Cas12 isotypes, meaning they are not mixed with known Cas12 isotypes; and these proteins disclosed herein ( Figure 1 B) The e-value compared to existing Cas12 HMM Profile models is greater than 1e-5. This portion of the protein disclosed includes... Figure 1 The CLUSTER1-CLUSTER13 proteins are shown in the image. Taken together, this suggests that these Cas proteins may be a new subgroup, such as a novel Cas12 isotype.
[0338] Figure 2 The SDS-PAGE electrophoresis image of the C12-279 recombinant protein is shown.
[0339] Figure 3 A PAM library containing Cas12 protein cleaved in vitro is shown. The sequences in the figure are shown in SEQ ID NO: 878-881.
[0340] Figure 4 The motif captured by C12-279-sgRNA after targeting a 7nt random sequence is shown.
[0341] Figure 5 The motif captured by C12-279-sgRNA-Rev after targeting a 7nt random sequence is shown.
[0342] Figure 6 The figure shows a fragment of a plasmid containing a 7nt random sequence during plasmid elimination experiments. The sequences in the figure are shown in SEQ ID NO: 720, 882-886.
[0343] Figure 7 The motif captured during the plasmid elimination experiment of C12-279 is shown.
[0344] Figure 8 The figure shows partial indel results generated after editing the TTR gene with C12-279. The sequences in the figure are shown in SEQ ID NO: 722, 887-916.
[0345] Figure 9The protein CI1062732 (SEQ ID NO: 46), which is a protein with dozens of amino acid residues deleted from the N-terminus of C12-279 protein (SEQ ID NO: 696), was shown to be combined with gRNA containing the DR sequence GTAATGCGTCTCCCATTGACGCC (SEQ ID NO: 529) to target a 7nt random sequence plasmid library in bacteria, and the captured "pseudo" PAM motif was 5'-TTNC-3'.
[0346] Figure 10 shows the secondary structure of the "pseudo" DR sequence analyzed by the inventors, who speculated that the 3' C base of the captured "pseudo" PAM motif TTNC might be due to an extra C at the 3' end of the DR. The sequences in the figure are shown in SEQ ID NO: 917 and 918.
[0347] Figure 11 shows the SDS-PAGE electrophoresis image of the C12-101-07 recombinant protein (126 kDa).
[0348] Figure 12 shows the motif captured by C12-101-07-sgRNA after targeting a 7nt random sequence.
[0349] Figure 13 shows the motif captured after C12-101-07-sgRNA-Rev targets a 7nt random sequence.
[0350] Figure 14 shows the motifs captured in the C12-101-09 plasmid elimination experiment.
[0351] Figure 15A shows a fragment of the pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid. The sequence in the figure is shown in SEQ ID NO: 919-921. Figure 15B shows a fragment of the library plasmid Plasmid Lib with a PAM sequence composition of NAAN. The sequence in the figure is shown in SEQ ID NO: 922-924.
[0352] Figure 16 shows the editing efficiency in NAAN cells of different C12-279 mutants, expressed as a multiple of the editing efficiency compared to C12-279.
[0353] Figure 17 shows the editing efficiency in NAAN cells of different C12-279 mutants, expressed as a multiple of the editing efficiency compared to C12-279.
[0354] Figure 18 shows the editing efficiency test results of different mutant targeted reporter systems. Wt represents wild-type C12-279 (green dot); the marked number n indicates that the amino acid residue at position n is mutated to arginine R; Mut-01 is mutant C12-279-pCDH-05 (Q186R variant), Mut-02 is mutant C12-279-pCDH-28 (Q186R and D352R double point mutation), and Mut-03 is mutant C12-279-pCDH-35 (G184R, Q186R and D352R triple point mutation); the green line (lower dashed line) indicates a 50% improvement in editing efficiency compared to wild-type; the red line (upper dashed line) indicates the same efficiency as Mut-02.
[0355] Figure 19 shows the editing efficiency test results of different multi-point mutants in the reporting system. Mut-02-1-5-426-858-860 represents a multi-point mutant obtained by introducing additional mutations at positions 1, 5, 426, 858, and 860 (all mutated to arginine R) on the basis of the Mut-02 mutant. That is, this multi-point mutant contains the following mutations: all amino acid residues at positions 1, 5, 186, 352, 426, 858, and 860 are mutated to R. Other mutants in this article are deduced similarly.
[0356] Figure 20A shows the efficiency of editing the TTR gene with multiple point mutants combined with different gRNAs. Figure 20B shows the efficiency of editing the HBG gene with multiple point mutants combined with different gRNAs. Mut-02-1-5-426-858-860 represents a multiple point mutant obtained by introducing additional mutations at positions 1, 5, 426, 858, and 860 (all mutated to arginine R) on the basis of the Mut-02 mutant. That is, this multiple point mutant contains the following mutations: all amino acid residues at positions 1, 5, 186, 352, 426, 858, and 860 are mutated to R. Other mutants follow the same principle.
[0357] Figure 21 shows the results of the dCas12-279 editing activity test. Mut-02-426-860 represents a multi-point mutant obtained by introducing additional 426R and 860R mutations into the Mut-02 mutant. Mut-02-426-860-D651A represents a multi-point mutant obtained by introducing additional 426R, 860R, and 651A mutations into the Mut-02 mutant. Other mutants in this paper are analogous.
[0358] Figure 22 shows the NGS sequencing results obtained after electrotransfecting HEK293 cells with the encoding mRNA of mutant Mut-02-1-5-426-858-860 in combination with modified gRNA (C279-dmTTR01-02) and editing the TTR gene, and the measured editing efficiency is as high as 92.18%. The sequences in the figure are shown in SEQ ID NO: 925-974.
[0359] Figure 23 shows the PAM recognized by C12-279 mutant Mut-02-1-426-846-858-860.
[0360] Figure 24 shows a schematic diagram of the putative structure of C12-279. DETAILED DESCRIPTION
[0361] In the present disclosure, unless otherwise specified, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Moreover, the operating procedures such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics and recombinant DNA used herein are all conventional procedures widely used in the corresponding fields. Meanwhile, for a better understanding of the present disclosure, definitions and explanations of related terms are provided below.
[0362] In the present disclosure, "several" refers to greater than or equal to 2. In the present disclosure, "plurality" refers to greater than or equal to 2.
[0363] In the present disclosure, depending on the context, "cleavage" may refer to cleavage of the backbone of a polynucleotide chain; non-limiting examples include completely breaking a single-stranded DNA, breaking one single strand of a double-stranded DNA, or breaking both two single strands of a double-stranded DNA.
[0364] In this disclosure, depending on the context, "modification" can refer to any form of nucleic acid chain chemical reaction other than "cutting"; including but not limited to base substitution, addition and / or deletion, as well as methylation and demethylation of nucleic acid chains. Non-limiting examples include base substitution on the target nucleic acid chain through single-base editing (fusion of the Cas12 disclosed herein with a deaminase domain, in conjunction with gRNA), such as A→G, C→T, T→C, or G→A nucleotide mutations, and other types of nucleotide mutations (e.g., A→T, C→G, T→A, G→C, etc.); base substitution, addition, or deletion can also be achieved through Prime editing technology (fusion of the Cas12 disclosed herein with a reverse transcriptase, in conjunction with PegRNA), or through HDR homologous recombination (e.g., the Cas12 disclosed herein in conjunction with gRNA and a donor template); the Cas12 disclosed herein can also be fused with DNA methyltransferase or DNA demethylase, in conjunction with gRNA for targeting.
[0365] In this disclosure, depending on the context, "regulating the expression of target nucleic acids" may refer to the regulation of the transcription of target nucleic acids; non-limiting examples include enhancing or inhibiting the transcription of target nucleic acids by means of CRISPRa, CRISPRi technologies, using Cas12 fusion transcriptional activation or repression domains.
[0366] In this disclosure, the letters in the amino acid sequence represent single-letter abbreviations of amino acids known in the art, such as those described in J. Biol. Chem, 243, p3558 (1968): Alanine: Ala-A, Arginine: Arg-R, Aspartic acid: Asp-D, Cysteine: Cys-C, Glutamine: Gln-Q, Glutamic acid: Glu-E, Histidine: His-H, Glycine: Gly-G, Asparagine: Asn-N, Tyrosine: Tyr-Y, Proline: Pro-P, Serine: Ser-S, Methionine: Met-M, Lysine: Lys-K, Valine: Val-V, Isoleucine: Ile-I, Phenylalanine: Phe-F, Leucine: Leu-L, Tryptophan: Trp-W, Threonine: Thr-T.
[0367] In this disclosure, "amino acid difference" refers to differences in amino acid residues at specific sites on the amino acid sequence of a protein, including substitutions, additions, or deletions.
[0368] In this disclosure, a mutation of an amino acid residue at a specific site refers to the substitution, addition, or reduction of an amino acid residue at that site.
[0369] As is known to those skilled in the art, in proteins or peptides, two adjacent amino acids each lose an OH or H atom, undergoing dehydration condensation to form a peptide bond. Each amino acid actually exists in the form of an amino acid residue. Therefore, in this disclosure, the terms "amino acid" and "amino acid residue" generally refer to the same thing. Furthermore, for the sake of simplicity, this disclosure retains the original amino acid residue before the position of the amino acid residue; the letter before the position indicates the original amino acid residue, and the letter after the position indicates the substituted amino acid residue. For example, S211 represents that the original amino acid residue at position 211 is S, and when it is replaced by R, it can be represented as S211R.
[0370] This article sometimes uses the symbol "+" to connect one amino acid mutation before and after it, indicating that two point mutations exist simultaneously in a mutant; if multiple point mutations are connected by two or more "+" symbols, it is intended to indicate that these multiple point mutations exist simultaneously.
[0371] In this disclosure, if an amino acid is substituted, it means that it is substituted by another amino acid residue that is different from the original amino acid residue. If the original amino acid was originally a positively charged amino acid, and it is substituted to a positively charged amino acid, it means that it is substituted by another positively charged amino acid residue that is different from the original amino acid residue. For example, if the original amino acid residue is R, and it is substituted to a positively charged amino acid, it means that it is substituted to H or K.
[0372] In this article, when referring to RNA sequences, the letter "T" can be used interchangeably with "U". When referring to "guide sequences", the letter "T" can be used interchangeably with "U". When referring to "directed repeat sequences", the letter "T" can be used interchangeably with "U".
[0373] In this article, when referring to Cas protein numbering, C12-n and Cas12-n are intended to refer to the same protein. For example, Cas12-279 and C12-279 are used interchangeably.
[0374] Sequence identity
[0375] As used herein, the term "identity" refers to the sequence matching between two polypeptides or two nucleic acids. "Identity," "percent identity," and "sequence identity" are used interchangeably. Two compared sequences are identical at that position when a position is occupied by the same base or amino acid monomer subunit (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared, multiplied by 100%. For example, if six out of ten positions in two sequences match, then the two sequences have 60% sequence identity. Typically, two sequences are compared to produce the maximum sequence identity. Such alignments can be performed using publicly available and commercially available alignment algorithms and programs, such as, but not limited to, CLUSTERalΩ, MAFFT, Probcons, T-Coffee, Probalign, and BLAST, which can be reasonably selected and used by those skilled in the art. Those skilled in the art can determine suitable parameters for the alignment sequences, including, for example, any algorithm required to achieve a better or optimal alignment of the entire length of the compared sequences, and any algorithm required to achieve a better or optimal alignment of a local portion of the compared sequences.
[0376] CRISPR-Cas12 system
[0377] As used herein, the terms “regularly clustered interspaced short palindromic repeats (CRISPR)-CRISPR-related (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements relating to the expression of CRISPR-related (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes. Such transcripts or other elements may contain sequences encoding Cas effector proteins and guide polynucleotides.
[0378] In 2015, Zhang Feng's research group discovered Cas12a, classifying it as type V in the Class I ICRISPR-Cas system. Following a detailed study of the VA subtype (Cas12a), Zhang Feng's group also reported Cas12b (C2C1) in 2015. In 2017, Burstein et al. reported the Cas12e (CasX) nuclease. In 2019, Winston X. Yan et al., through bioinformatics analysis, reported in detail the newly discovered type V Cas effector proteins Cas12c, Cas12h, Cas12i, and Cas12g.
[0379] In some embodiments, the Cas12 protein described herein refers to a protein whose amino acid sequence includes or has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of SEQ ID NO: 1-53, 696, 728.
[0380] In this disclosure, the CRISPR-Cas12 system comprises a Cas12 protein having at least 50% sequence identity with any one of SEQ ID NO: 1-53, 696, 728, or a nucleic acid encoding the Cas12 protein; and a guide polynucleotide or a nucleic acid encoding the guide polynucleotide; the guide polynucleotide comprising a homologous repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize with a target nucleic acid, the guide polynucleotide being capable of forming a complex with the Cas12 protein and guiding the complex to bind sequence-specifically to the target nucleic acid.
[0381] Guided polynucleotides
[0382] As used herein, the term "guide polynucleotide" refers to a molecule in the CRISPR-Cas system that forms a complex with the Cas protein and guides the complex to a target sequence. Typically, a guide polynucleotide contains a backbone sequence linked to a guide sequence that can hybridize with the target sequence. The backbone sequence typically contains a direct repeat sequence and may sometimes contain a tracrRNA sequence. In some embodiments, the guide polynucleotide does not contain a tracrRNA sequence. In some embodiments, the guide polynucleotide contains a tracrRNA sequence.
[0383] In some embodiments, the guide polynucleotide of the CRISPR-Cas12 system is guide RNA. In some embodiments, the guide polynucleotide is a chemically modified guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least one chemically modified nucleotide.
[0384] In some embodiments, the chemically modified nucleotides include base-modified nucleotides, phosphate-modified nucleotides, and ribose-modified nucleotides.
[0385] In some preferred embodiments, the base-modified nucleotide is selected from nucleotides with non-natural bases.
[0386] In some preferred embodiments, the phosphate-modified nucleotide is selected from aminophosphate nucleotides, thiophosphate nucleotides, dithiophosphate nucleotides, methylphosphonate nucleotides, 5′-phosphate nucleotides, alkylphosphate nucleotides, and borane phosphate nucleotides.
[0387] In some preferred embodiments, the ribose-modified nucleotide is selected from deoxynucleotides, 3′-terminal deoxythymidine (dT) nucleotides, 2′-O-methyl modified nucleotides, 2′-fluorine modified nucleotides, 2′-deoxy modified nucleotides, 2′-amino modified nucleotides, 2′-O-allyl modified nucleotides, 2′-C-alkyl modified nucleotides, 2′-hydroxy modified nucleotides, 2′-methoxyethyl modified nucleotides, 2′-O-alkyl modified nucleotides, and morpholinonucleotides.
[0388] In some preferred embodiments, the guiding nucleotide comprises a chemically modified nucleotide at the 5' end and / or the 3' end.
[0389] In some specific embodiments, the guide nucleotide includes a deoxynucleotide at its 5' end; the deoxynucleotide is 10-25 nt in length, for example, 14 or 25 nt.
[0390] In some specific embodiments, the guiding nucleotide contains a phosphate thioester group at the 5' or 3' end and is a nucleotide modified with 2'-OMe.
[0391] In some embodiments, the guiding polynucleotide comprises at least one guide sequence (also called a spacer sequence) linked to at least one direct repeat (DR) sequence. In some embodiments, the guide sequence is located at the 3′ end of the direct repeat sequence. In some embodiments, the guide sequence is located at the 5′ end of the direct repeat sequence.
[0392] In some implementations, the tracrRNA sequence is linked to the same-direction repeat sequence.
[0393] In some embodiments, the tracrRNA sequence is located at the 5′ or 3′ end of the same-direction repeat sequence. In some embodiments, the tracrRNA sequence is located at the 5′ end of the same-direction repeat sequence. In some embodiments, the tracrRNA sequence is located at the 3′ end of the same-direction repeat sequence.
[0394] In some embodiments, the nucleotide sequence of the guiding polynucleotide comprises, from 5′ to 3′, the following sequence: tracrRNA, a direct repeat sequence, and a guiding sequence.
[0395] In some embodiments, the nucleotide sequence of the guiding polynucleotide comprises, from 5′ to 3′, the following sequence: tracrRNA, a linker sequence, a direct repeat sequence, and a guiding sequence.
[0396] In some embodiments, the nucleotide sequence of the guiding polynucleotide comprises, from 5′ to 3′, the following sequence: tracrRNA, loop sequence, direct repeat sequence, and guiding sequence.
[0397] In some implementations, the structure of the guiding polynucleotide is 5′-tracrRNA-loop-directed repeat sequence-guiding sequence-3′.
[0398] In some implementations, the guiding polynucleotide tracrRNA and the same-direction repeat sequence are linked by a nucleotide sequence.
[0399] In the specific embodiments disclosed herein, the tracrRNA sequence and the recurrent repeat sequence are linked by a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In the specific embodiments disclosed herein, the tracrRNA sequence and the recurrent repeat sequence are linked by a nucleotide sequence consisting of 4 nucleotides. In the specific embodiments disclosed herein, the tracrRNA sequence and the recurrent repeat sequence are linked by a 5'-GAAA-3' sequence.
[0400] In some embodiments, the guide sequence is sufficiently complementary to the target nucleic acid sequence to hybridize with the target nucleic acid and guide the CRISPR-Cas12 complex to bind sequence-specifically to the target nucleic acid. In some embodiments, the guide sequence is 100% complementary to the target nucleic acid, but the guide sequence may be less than 100% complementary to the target nucleic acid, for example, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% complementary.
[0401] In some embodiments, the guide sequence is engineered to hybridize with the target nucleic acid, with a mismatch of no more than two nucleotides. In some embodiments, the guide sequence is engineered to hybridize with the target nucleic acid, with a mismatch of no more than one nucleotide. In some embodiments, the guide sequence is engineered to hybridize with the target nucleic acid, with or without a mismatch.
[0402] In some specific embodiments disclosed herein, the guiding sequence, such as any one of SEQ ID NO: 783-805, 775, 825-877, has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity. In some specific embodiments disclosed herein, the guiding sequence is as shown in any one of SEQ ID NO: 783-805, 775, 825-877.
[0403] In some embodiments, the CRISPR-Cas12 system comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different guide polynucleotides. In some embodiments, the guide polynucleotides target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different target nucleic acid molecules, or target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different regions of one or more target nucleic acid molecules.
[0404] In some embodiments, the guiding polynucleotide comprises a constant, unidirectional repeat sequence upstream of a variable guiding sequence. In some embodiments, multiple guiding polynucleotides are part of an array (which may be part of a vector, such as a viral vector or plasmid). For example, a guiding array comprising the sequence DR-spacer-DR-spacer-DR-spacer-...-DR-spacer may comprise multiple unique, unprocessed guiding polynucleotides (one for each DR-spacer or spacer-DR sequence). Once introduced into a cell or cell-free system, the array is processed by the Cas12 protein into several individual mature guiding polynucleotides. This allows for multiplexing, such as delivering multiple guiding polynucleotides to a cell or system to target multiple target nucleic acids or multiple regions within a single target nucleic acid.
[0405] The ability of a CRISPR guide polynucleotide complex (CRISPR complex) to bind specifically to a target nucleic acid sequence can be assessed by any suitable assay. For example, components of a CRISPR system sufficient to form the complex (CRISPR complex), including the guide polynucleotide to be tested, can be provided to a host cell containing the corresponding target nucleic acid molecule, for example, through transfection with a vector encoding a component of the CRISPR complex, and then the preferential cleavage within the target sequence can be assessed. Similarly, the cleavage of the target nucleic acid sequence can be assessed in vitro by providing the target nucleic acid, components of the CRISPR complex, including the guide polynucleotide to be tested and a control guide polynucleotide different from the test guide polynucleotide, and comparing the ability of the test and control guide polynucleotides to bind to the target nucleic acid or the rate of cleavage of the target nucleic acid. The ability of the CRISPR complex to cleave the target nucleic acid or the target nucleic acid itself can also be assessed by the assays described above.
[0406] Cas12 mutant
[0407] As described herein, when referring to "the corresponding position of the sequence shown in SEQ ID NO: XX" or using similar wording, the corresponding position is determined by amino acid sequence alignment. Typically, two sequences are compared to produce maximum sequence identity. Such alignments can be performed using publicly available and commercially available alignment algorithms and programs, such as, but not limited to, ClustalΩ, MAFFT, Probcons, T-Coffee, Probalign, and BLAST, which can be reasonably selected by those skilled in the art. Those skilled in the art can determine suitable parameters for aligning sequences, including, for example, any algorithm required to achieve a better or optimal alignment of the full length of the compared sequences, and any algorithm required to achieve a better or optimal alignment of a portion of the compared sequences.
[0408] Optionally, the corresponding position is determined by performing an online sequence alignment of the amino acids of the Cas12 protein with any of the sequences shown in SEQ ID NO: 1-53, 696, 728 using the MAFFT version 7 tool (https: / / mafft.cbrc.jp / alignment / server / index.html), selecting the following parameters: G-INS-i (Very slow; recommended for <200 sequences with global homology; 2iterative cycles only), Try to align gappy regions anyway, Scoring matrix for amino acid sequences-BLOSUM62, Gap opening penalty 1.53, Offset value 0.0, Mafft-homologs-Use UniRef50 (more comprehensive and requires longer search time).
[0409] Optionally, the corresponding position is determined by performing an online sequence alignment of the amino acids of the Cas12 protein with the sequence shown in SEQ ID NO: 696 using the MAFFT version 7 tool (https: / / mafft.cbrc.jp / alignment / server / index.html), selecting the following parameters: G-INS-i (Very Slow; recommended for <200 sequences with global homology; 2iterative cycles only), Try to align gappy regions anyway, Scoring matrix for amino acid sequences -- BLOSUM62, Gap opening penalty 1.53, Offset value 0.0, Mafft-homologs -- Use UniRef50 (more comprehensive and requires longer search time).
[0410] In some embodiments, the Cas12 protein provided herein contains one or more mutations compared to the Cas12 protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728, such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof. In some examples, the Cas12 protein, compared to the Cas12 protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728, contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 3 4, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.In some instances, the Cas12 protein, compared to the Cas12 protein represented by any of the sequences SEQ ID NO: 1-53, 696, 728, comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 7 3, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide.
[0411] One type of modification or mutation involves replacing amino acid residues with similar biochemical properties; this is known as conserved substitution. Typically, conserved substitutions have little or no effect on the activity of the resulting protein or peptide. For example, a conserved substitution is an amino acid substitution in the Cas12 protein that essentially does not affect the binding of the Cas12 protein to target nucleic acid molecules complementary to the guide sequence of the gRNA molecule, and / or the processing of the guide array RNA transcript into gRNA molecules.
[0412] More substantial alterations can be achieved by using less conserved substitutions, for example, by selecting residues that differ more significantly in maintaining the following effects: (a) the structure of the polypeptide backbone in the region where the substitution occurs, for example, as a helical or folded conformation; (b) the charge or hydrophobicity of the region interacting with the target site; or (c) the volume of the side chain. Substitutions that are generally expected to produce the greatest changes in polypeptide function are (a) substitutions between hydrophilic residues (e.g., serine or threonine) and hydrophobic residues (e.g., leucine, isoleucine, phenylalanine, valine, or alanine); (b) substitutions between cysteine or proline and any other residue; (c) substitutions between residues with positively charged side chains (e.g., lysine, arginine, or histidine) and negatively charged residues (e.g., glutamic acid or aspartic acid); or (d) substitutions between residues with large side chains (e.g., phenylalanine) and residues without side chains (e.g., glycine).
[0413] Cas12 active fragment
[0414] In this disclosure, the Cas12 protein may contain only the WED-I domain, Helical-I1 domain, PI domain, Helical-I2 domain, Helical-II domain, WED-II domain, Ruvc-I domain, Helical-III domain, BH domain, R-uvc-II domain, Nuc domain and / or Ruvc-III domain.
[0415] The Cas12 protein disclosed herein may include, in addition to the aforementioned domain, other Cas12 protein domains from the prior art, which together form the complete structure of the Cas12 protein to achieve the functions of the Cas12 protein disclosed herein, including but not limited to retaining the ability of the Cas12 protein to form a complex with gRNA, retaining the ability of the Cas12 protein to form a complex with gRNA and target target nucleic acids, retaining the ability of the complex formed by the Cas12 protein and gRNA to target and regulate the expression of target nucleic acids, retaining the ability of the complex formed by the Cas12 protein and gRNA to target and cleave single-stranded or double-stranded target nucleic acids, retaining the ability of the Cas12 protein to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.
[0416] Cas12 inactivated variant
[0417] By inactivating the RuvC domain of Cas12 through point mutation, the Cas12 protein will lose its endonuclease activity, and the resulting dCas12 (dead Cas12) can only bind to target genes under the mediation of guide polynucleotides, but does not have the function of cutting DNA.
[0418] Alternatively, point mutations can be used to partially deactivate the RuvC domain of Cas12, forming Cas12 nickase (nCas12). This nickase binds to the target gene under the guidance of a guide polynucleotide, cleaving one single strand of the double-stranded nucleic acid without cleaving the other single strand.
[0419] Therefore, dCas12 or nCas12 can be fused or conjugated with other domains (including but not limited to deaminase domains, transcription activation domains, transcription repression domains, methylation domains, demethylation domains, histone acetylation domains, and histone deacetylation domains) to guide polynucleotides to the target sequence of the target nucleic acid, and then perform the corresponding functions with the help of the other domains; for example, C→T conversion of the target nucleic acid bases can be achieved by deamination of cytosine bases, A→G conversion of the target nucleic acid bases can be achieved by deamination of adenine bases, transcriptional repression of the target nucleic acid can be achieved by the transcription repression domain KRAB, and transcriptional activation domain VP64 can promote the transcription of the target nucleic acid.
[0420] Functional structural domain
[0421] In some embodiments, the Cas12 protein or Cas12 inactivated variant is covalently linked or fused with homologous or heterologous functional domains.
[0422] In some embodiments, the functional domain has enzymatic activity that modifies the target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristylation activity, and / or demyristylation activity.
[0423] In some embodiments, the functional domain is optionally selected from one or more of the following: nucleases (e.g., FokI), methyltransferases, demethylases, DNA repair enzymes, DNA damage enzymes, deaminases, superoxide dismutases, alkylating enzymes, depurinases, oxidases, pyrimidine dimer forming enzymes, integrases, transposases, recombinases, polymerases, ligases, helicases, photolyases, glycosylation enzymes, deglycosylation enzymes, acetyltransferases, deacetylases, kinases, phosphatases, ubiquitin ligases, deubiquitinating enzymes, adenylate acylases, deadenylate acylases, SUMOylating enzymes, deSUMOylating enzymes, myristylases, and / or demyristylases.
[0424] In some embodiments, the functional domains are selected from one, two, three, four or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase repression domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, epitope tag and / or reporter domain.
[0425] In some embodiments disclosed herein, the deaminase domain may be selected from: APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, activation-induced cytidine deaminase (AID), CDA from lamprey, or a mutant of adenosine deaminase (TadA) engineered to act on DNA.
[0426] In some embodiments, the transcriptional activation domain is optionally selected from: p65, VPR, VP16, VP64, VTR1, VTR2, VTR3, p65, MyoD1, HSF1, RTA, SET7 / 9, and histone acetyltransferase. In some embodiments, the transcriptional activation domain is optionally selected from: the sequence ETFSDLWKL from p53 TAD1, the sequence DDIEQWFTE from p53 TAD2, the sequence SDIMDFVLK from MLL, the sequence DLLDFSMMF from E2A, the sequence ETLDFSLVT from Rtg3, the sequence RKILNDLSS from CREB, the sequence EAILAELKK from CREBaB6, the sequence DDVVQYLNS from Gli3, the sequence DDVYNYLFD from Gal4, the sequence DLFDYDFLV from Oaf1, the sequence DFFDYDLLF from Pip2, the sequence EDLYSILWS from Pdr1, and the sequence TDLYHTLWN from Pdr3.
[0427] In some embodiments, the transcriptional repressor domain is optionally selected from: KOX1, KAP-1, MAD, FKHR, EGR-1, ERD, SID, a tandem of SID (e.g., SID4X), TIEG, v-ERB-A, MBD2, MBD3, TRa, histone methyltransferase, histone deacetylase (HDAC), nuclear hormone receptor (e.g., estrogen receptor or thyroid hormone receptor), DNMT family members (e.g., DNMT1, DNMT3A, DNMT3B), the KRAB domain of MeCP2, ROM2, and AtHD2A.
[0428] In some embodiments, the transcriptional repressor domain is the KRAB domain from the KOX1 protein.
[0429] In some embodiments, the nuclease domain is optionally selected from FokI, a polypeptide with ssDNA cleavage activity, or a polypeptide with dsDNA cleavage activity.
[0430] In some embodiments, the methyltransferase domain is selected from DNA methyltransferases, including but not limited to DNMT1, DNMT3a, and DNMT3b.
[0431] In some embodiments, the demethylase is selected from TET1CD, TET1, ROS1, DME, DML2, and DML3.
[0432] Methylation and demethylation are widely recognized in the field as important mechanisms of epigenetic gene regulation.
[0433] In some embodiments, the homologous or heterologous functional domains are sequence tags useful for the dissolution, purification, or detection of the fusion protein or conjugate. Suitable protein tag sequences are provided herein, including but not limited to biotinylate carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, multihistidine tags (also known as His tags), maltose-binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S-tags, softtags (e.g., Softag1, Softag3), strep-tags, biotin ligase tags, FLASH tags, V5 tags, and SBP tags. Other suitable sequences will be readily apparent to those skilled in the art.
[0434] In some embodiments disclosed herein, a single-base editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with a deaminase domain and a nuclear localization signal. In other embodiments disclosed herein, a single-base editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with an APOBEC3A domain and an SV40 NLS domain.
[0435] In some embodiments disclosed herein, a transcriptional repression epigenetic editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with KRAB and SV40 NLS. In some embodiments disclosed herein, a transcriptional activation epigenetic editor is constructed by fusing the Mut-02-1-426-846-858-860-D651A-E891A-D1082A mutant with the VP64 domain and SV40 NLS.
[0436] Subcellular localization signals
[0437] In some embodiments, the Cas12 protein is fused with at least one homologous or heterologous subcellular localization signal. Exemplary subcellular localization signals include organelle localization signals, such as nuclear localization signals (NLS), nuclear export signals (NES), or mitochondrial localization signals.
[0438] Non-limiting examples of NLS include NLS sequences derived from: NLS of the SV40 viral large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 738); NLS from nucleoplasmic proteins (e.g., sequence KRPAATKKAGQAKKKK, SEQ ID NO: 739); c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 740) or RQRRNELKRSP (SEQ ID NO: 741); hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 742); sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 743) from the IBB domain; and sequences VSRKRPRP (SEQ ID NO: 744) and PPKKARED (SEQ ID NO: 745) from the fibroid T protein. Sequences of human p53 (SEQ ID NO: 745); human p53 sequence PQPKKKPL (SEQ ID NO: 746); mouse c-ablIV sequence SALIKKKKKMAP (SEQ ID NO: 747); influenza virus NS1 sequences DRLRR (SEQ ID NO: 748) and PKQKKRK (SEQ ID NO: 749); hepatitis virus delta antigen sequence RKLKKKIKKL (SEQ ID NO: 750); mouse Mx1 protein sequence REKKKFLKRR (SEQ ID NO: 751); human PARP enzyme sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 752); and steroid hormone receptor sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 753). In some embodiments, the nuclear localization sequences are strong enough to drive the fusion proteins or conjugates disclosed herein to accumulate in detectable amounts in the nucleus of eukaryotic cells. In general, the intensity of nuclear localization activity can be derived from the number of NLSs, one or more specific NLSs used, or a combination of these factors. Accumulation in the nucleus can be detected using any suitable technique. For example, detectable markers can be fused to Cas proteins to visualize their intracellular location, such as in combination with means for detecting nuclear location (e.g., nucleus-specific dyes, such as DAPI). The nucleus can also be isolated from the cell, and its contents can then be analyzed using any suitable method for protein detection, such as immunohistochemistry, Western blotting, or enzyme activity assays.Accumulation in the nucleus can also be determined indirectly by means such as measuring the role of nucleic acid targeting complex formation (e.g., measuring DNA or RNA cleavage or mutation at the target sequence, or measuring altered gene expression activity due to the effects of DNA or RNA targeting complex formation and / or DNA or RNA targeting Cas protein activity), compared with controls that are not exposed to nucleic acid targeting Cas protein or nucleic acid targeting complex, or exposed to nucleic acid targeting Cas protein lacking one or more NLS.
[0439] carrier system
[0440] Another aspect of this disclosure relates to a vector system comprising the CRISPR-Cas12 system described herein, the vector system comprising one or more recombinant vectors, the recombinant vectors comprising a polynucleotide sequence encoding the Cas12 protein and a polynucleotide sequence encoding the guide polynucleotide.
[0441] In some embodiments, the vector system comprises at least one plasmid or viral recombinant vector (e.g., retrovirus, lentivirus, adenovirus, adeno-associated virus, or herpes simplex virus). In some embodiments, the polynucleotide sequence encoding the Cas12 protein and the polynucleotide sequence encoding the guide polynucleotide are located on the same recombinant vector. In some embodiments, the polynucleotide sequence encoding the Cas12 protein and the polynucleotide sequence encoding the guide polynucleotide are located on multiple recombinant vectors.
[0442] In some embodiments, the polynucleotide sequence encoding the Cas12 protein and / or the polynucleotide sequence encoding the guiding polynucleotide is operatively linked to a regulatory sequence (also called a regulatory element). The regulatory element includes promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Regulatory elements include those that constitutively express the nucleotide sequence in many types of host cells, and those that express the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters may be expressed directly primarily in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). Regulatory elements may also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, and may or may not be tissue- or cell type-specific. In some embodiments, the regulatory element is an enhancer element, such as WPRE, CMV enhancer, R-U5 segment in the LTR of HTLV-1, SV40 enhancer, or intron sequence between exons 2 and 3 of rabbit β-globin.
[0443] In some embodiments, the recombinant vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., a retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), an SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, a glycerol phosphokinase (PGK) promoter, or an EF1 promoter), or a pol III promoter and a pol II promoter.
[0444] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1α, CAG, and β-actin promoters. In some embodiments, the promoter is an inducible promoter regulated by external signals or molecules (e.g., transcription factors).
[0445] In some embodiments, the promoter is a tissue-specific promoter that can be used to drive tissue-specific expression of the Cas12 protein. Suitable muscle-specific promoters include, but are not limited to, CK8, MHCK7, myoglobin promoter (Mb), desmin promoter, muscle creatine kinase promoter (MCK) and its variants, and SPc5-12 synthesis promoter. Suitable immune cell-specific promoters include, but are not limited to, the B29 promoter (B cells), the CD14 promoter (monocytes), the CD43 promoter (leukocytes and platelets), CD68 (macrophages), and the SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include, but are not limited to, the CD43 promoter (leukocytes and platelets), the CD45 promoter (hematopoietic cells), INF-β (hematopoietic cells), the WASP promoter (hematopoietic cells), the SV40 / CD43 promoter (leukocytes and platelets), and the SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include, but are not limited to, the elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, the Fit-1 promoter and the ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, the GFAP promoter (astrocytes), the SYN1 promoter (neurons), and the NSE / RU5′ promoter (mature neurons). Suitable kidney-specific promoters include, but are not limited to, the NphsI promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, the OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, the SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, the SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, α-MHC.
[0446] AAV carrier
[0447] Another aspect of this disclosure relates to an adeno-associated virus (AAV) vector comprising the CRISPR-Cas12 system described herein, wherein the AAV vector comprises DNA encoding the Cas12 protein and the guide polynucleotide described herein.
[0448] Delivery of the CRISPR-Cas system via an AAV vector was described in Maeder et al., Nature Medicine 25:229-233 (2019). Clinically demonstrated safety and efficacy of subretinal AAV delivery were described. Local delivery via subretinal injection, the natural tropism of AAV5 for photoreceptor cells, and the use of the photoreceptor-specific GRK1 promoter all contribute to limiting CRISPR / Cas system expression to therapeutic target tissues and cell types, which is incorporated herein by reference in its entirety. In some embodiments, the AAV vector comprises an ssDNA genome containing the coding sequences for the Cas12 protein and the guide polynucleotides flanking the ITR.
[0449] In some embodiments, the CRISPR-Cas12 system described herein is packaged in an AAV vector, such as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, and AAVrh74. In some embodiments, the CRISPR-Cas12 system described herein is packaged in an AAV vector containing a tissue-tropy engineered capsid, such as an engineered muscle-tropy capsid. Tabebordbar et al., Cell 184:4919-4938 (2021) describes the engineering of tissue-tropy AAV capsids through directed evolution, identifying a class of capsids containing RGD motifs, and demonstrating efficient transduction of primate muscle via systemic injection of MyoAAV. This paper is incorporated herein by reference in its entirety.
[0450] Lipid nanoparticles
[0451] Another aspect of this disclosure relates to lipid nanoparticles (LNPs) comprising the CRISPR-Cas12 system described herein, wherein the LNP comprises the guiding polynucleotide described herein and mRNA encoding the Cas12 protein described herein.
[0452] The delivery of LNPs using the CRISPR-Cas system is described in Gillmore et al., N. Engl. J. Med., 385: 493-502 (2021). The lipid nanoparticles (LNPs) consist of four lipids, including the proprietary ionizable lipid LP000001; DSPC; cholesterol; and DMG-PEG2k. The LNP suspension was prepared in an aqueous buffer of Tris, NaCl, and sucrose at pH 7.4. The full text of this paper is incorporated herein by reference. In some embodiments, in addition to the RNA payload (Cas12 mRNA and guide polynucleotide), the lipid nanoparticles (LNPs) also contain four components: cationic or ionizable lipids, cholesterol, cofactor lipids, and PEG-lipids. In some embodiments, the cationic or ionizable lipids include cKK-E12, C12-200, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA, FTT5, Moderna SM-102, and Intellia LP01. In some embodiments, the PEG-lipids comprise PEG-2000-C-DMG, PEG-2000-DMG, or ALC-0159. In some embodiments, the accessory lipids comprise DSPC. The components of the LNP are described in Paunovska et al., Nature Reviews Genetics 23:265-280 (2022), and the FDA-approved LNP contains variants of four basic components: cationic or ionizable lipids, cholesterol, accessory lipids, and polyethylene glycol (PEG) lipids, which are incorporated herein by reference in their entirety.
[0453] Lentiviral vector
[0454] Another aspect of this disclosure relates to a lentiviral vector comprising the CRISPR-Cas12 system described herein, wherein the lentiviral vector comprises the guide polynucleotide described herein and mRNA encoding the Cas12 protein described herein. In some embodiments, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein such as VSV-G. In some embodiments, the mRNA encoding the Cas12 protein is linked to an aptamer sequence.
[0455] RNP complex
[0456] Another aspect of this disclosure relates to a ribonucleoprotein complex comprising the CRISPR-Cas12 system described herein, wherein the ribonucleoprotein complex is formed from the guide polynucleotide and Cas12 protein described herein. In some embodiments, the ribonucleoprotein complex can be delivered to eukaryotic cells, mammalian cells, or human cells by microinjection or electroporation. In some embodiments, the ribonucleoprotein complex can be packaged in virus-like particles and delivered in vivo to mammalian or human subjects.
[0457] Virus-like particles
[0458] Another aspect of this disclosure relates to virus-like particles (VLPs) comprising the CRISPR-Cas12 system described herein, wherein the virus-like particles comprise the guide polynucleotide and Cas12 protein described herein, or a ribonucleoprotein complex consisting of the guide polynucleotide and Cas12 protein.
[0459] Banskota et al. Cell 185(2):250-265 (2022) reported the development and application of DNA-free virus-like particles (eVLPs) for efficient packaging and delivery of base editors or Cas9 ribonucleoprotein; Mangeot et al., Nature Communications 10(1):1-15 (2019) induced efficient genome editing in cell lines and primary cells (including human induced pluripotent stem cells, human hematopoietic stem cells, and mouse bone marrow cells) using engineered mouse leukemia virus-like particles (Nanoblades) loaded with Cas9-sgRNA ribonucleoprotein; Campbell, et al., Molecular Therapy27:151-163 (2019) utilizes a specialized extracellular vesicle called a “gesicle” to efficiently but transiently deliver Cas9, targeting the HIV long terminal repeat (LTR), in the form of a ribonucleoprotein. Gesicles are produced by expressing vesicular stomatitis virus glycoproteins and packaging proteins (as its cargo), thus eliminating the need for transgenic delivery and allowing for more precise control over Cas9 expression. Mangeot et al. Molecular Therapy, 19(9):1656-1666 (2011) reported that overexpression of the spike glycoprotein of vesicular stomatitis virus (VSV-G) in human cells induced the release of fusion vesicles called gesicles. Biochemical and functional studies showed that glial cells bind proteins from producing cells and can transport them to recipient cells. This protein transduction method allows for the direct transport of cytoplasmic, nuclear, or surface proteins in target cells. These references all describe engineered VLPs, the full text of which is incorporated herein by reference.
[0460] In some embodiments, the engineered virus-like particles (VLPs) are pseudotyped with homologous or heterologous envelope proteins such as VSV-G. In some embodiments, the Cas12 protein is fused to a gag protein (e.g., MLVgag) via a cleavable linker, wherein cleavage of the linker in the target cell exposes the NLS located between the linker and the Cas12 protein. In some embodiments, the fusion protein or conjugate comprises (e.g., from 5′ to 3′) a gag protein (e.g., MLVgag), one or more NES, a cleavable linker, one or more NLS, and Cas12, as described in Banskota et al. Cell 185(2):250-265(2022).
[0461] In some embodiments, the Cas12 protein is fused with a first dimerizing domain that is capable of dimerizing or heterodimerizing with a second dimerizing domain fused to a membrane protein, wherein the presence of a ligand promotes the dimerization and enriches the Cas12 protein or fusion protein or conjugate into the VLP, as described in Campbell, et al., Molecular Therapy 27:151-163 (2019).
[0462] cell
[0463] Another aspect of this disclosure relates to cells comprising the CRISPR-Cas12 system described herein. Cells (e.g., those that can be used to generate cell-free systems) can be eukaryotic or prokaryotic. Examples of such cells include, but are not limited to, bacterial, archaea, plant, fungal, yeast, insect, and mammalian cells, such as Lactobacillus, Lactococcus, Bacillus (e.g., Bacillus subtilis), Escherichia (e.g., Escherichia coli), Clostridium, Yeast, or Pichia pastoris (e.g., Saccharomyces cerevisiae or Pichia pastoris), Kluyveromyces lactis, Salmonella typhimurium, Drosophila cells, Caenorhabditis elegans cells, Xenopus laevis cells, SF9 cells, C129 cells, 293 cells, Neurospora, and immortalized mammalian cell lines (e.g., HeLa cells, bone marrow cell lines, and lymphoid cell lines).
[0464] In some embodiments, the cells are prokaryotic cells, such as bacterial cells, such as *Escherichia coli*. In some embodiments, the cells are eukaryotic cells, such as mammalian cells or human cells. In some embodiments, the cells are primary eukaryotic cells, stem cells, tumor / cancer cells, circulating tumor cells (CTCs), blood cells (e.g., T cells, B cells, NK cells, Tregs, etc.), hematopoietic stem cells, specialized immune cells (e.g., tumor-infiltrating lymphocytes or tumor suppressor lymphocytes), or stromal cells in the tumor microenvironment (e.g., cancer-associated fibroblasts, etc.). In some embodiments, the cells are brain or neuronal cells of the central or peripheral nervous system (e.g., neurons, astrocytes, microglia, retinal ganglion cells, rod / cone cells, etc.).
[0465] Target nucleic acid or target DNA
[0466] In some embodiments disclosed herein, the target nucleic acid is target DNA.
[0467] The CRISPR-Cas12 system described herein can be used to target one or more target nucleic acid molecules, such as those present in biological samples, environmental samples (e.g., soil, air, or water samples).
[0468] In some embodiments disclosed herein, the target nucleic acid is a disease or symptom-related gene. In some embodiments disclosed herein, the target nucleic acid is a disease-related gene. In some embodiments disclosed herein, the disease-related gene is a pathogenic gene that directly causes the disease. In some embodiments disclosed herein, the disease-related gene is an abnormal gene that directly causes the disease or a gene whose expression is abnormal. For example, an unfavorable mutation in the gene leads to the occurrence of the disease. As another example, overexpression or underexpression of the gene leads to the occurrence of the disease. In some embodiments disclosed herein, overexpression of the gene leads to the occurrence of the disease. In some embodiments disclosed herein, underexpression of the gene leads to the occurrence of the disease. In some embodiments disclosed herein, overexpression of the gene is associated with the occurrence of the disease. In some embodiments disclosed herein, underexpression of the gene is associated with the occurrence of the disease.
[0469] In some embodiments disclosed herein, the disease or condition is a hematologic disease or condition, an ophthalmic disease or condition, a neurological disease or condition, a respiratory disease or condition, a liver disease or condition, a metabolic disease or condition, cancer, or an infectious disease.
[0470] In some embodiments disclosed herein, the target nucleic acid is selected from genes listed in Table 27, and the disease or symptom is one listed in Table 27. Table 27 shows the target nucleic acids and the corresponding diseases or symptom for each target nucleic acid.
[0471] In some embodiments disclosed herein, the diseases or conditions mentioned are selected from: hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes mellitus, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bied syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, X-pharyngeal cancer, Bietti lens dystrophy Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle-invasive bladder cancer, non-alcoholic... Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent... Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, canavan disease, cocaine addiction, Clabberg disease.Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Ritter syndrome. Triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignant osteosclerosis. Nutritional bullous epidermolysis, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0472] The genes associated with thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0473] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0474] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0475] The genes related to AATD lung disease include, but are not limited to, AATD;
[0476] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0477] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0478] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0479] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0480] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0481] The genes related to hemophilia B include, but are not limited to, factor IX;
[0482] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0483] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0484] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0485] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0486] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0487] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0488] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0489] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0490] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0491] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0492] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0493] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0494] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0495] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0496] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B;
[0497] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0498] The epilepsy-related genes include, but are not limited to, GAT1.
[0499] In some specific embodiments disclosed herein, the sequence of the target nucleic acid is shown as any one of SEQ ID NO: 761-782.
[0500] Non-limiting examples of such target nucleic acids also include those listed in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed December 12, 2012 and January 2, 2013, respectively, and International Application PCT / US2013 / 074667, filed December 12, 2013, all of which are incorporated herein by reference.
[0501] In some embodiments, the target nucleic acid is a reporter gene. Examples of reporter genes include, but are not limited to, glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP).
[0502] Applications for treating or preventing diseases
[0503] Another aspect of this disclosure relates to a pharmaceutical composition comprising, for example, the Cas12 protein as described herein, the guide polynucleotide as described herein, the Cas12 inactivating variant as described herein, the fusion protein or conjugate as described herein, the nucleic acid as described herein, the CRISPR-Cas12 system as described herein, the vector system as described herein, the delivery system as described herein, or the cell as described herein. The pharmaceutical composition may comprise, for example, an AAV vector encoding, the Cas12 protein or Cas12 inactivating variant described herein, and the guide polynucleotide. The pharmaceutical composition may comprise, for example, lipid nanoparticles containing, the guide polynucleotide described herein, and mRNA encoding the Cas12 protein. The pharmaceutical composition may comprise, for example, a lentiviral vector containing, for example, the guide polynucleotide described herein and mRNA encoding the Cas12 protein. The pharmaceutical composition may comprise, for example, virus-like particles containing, for example, the guide polynucleotide described herein and the Cas12 protein, or a ribonucleoprotein complex formed from said guide polynucleotide and the Cas12 protein.
[0504] Another aspect of this disclosure relates to the use of the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, the cell as described in this disclosure, the pharmaceutical composition as described in this disclosure, or the kit as described in this disclosure in cutting or editing target nucleic acids in mammalian cells.
[0505] Another aspect of this disclosure relates to the use of the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, the cell as described in this disclosure, the pharmaceutical composition as described in this disclosure, or the kit as described in this disclosure for any of the following purposes: cleaving or creating a nick in one or more target nucleic acid molecules, activating or upregulating the expression of one or more target nucleic acid molecules, activating or inhibiting the transcription of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling, or detecting one or more target nucleic acid molecules, binding one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules.
[0506] Another aspect of this disclosure relates to the use of modifying one or more target nucleic acid molecules with the Cas12 protein, the guide polynucleotide, the inactivated variant, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas12 system, the vector system, the delivery system, the cell, the pharmaceutical composition, or the kit, as described in this disclosure, including one or more of the following: nucleic acid base substitution, nucleic acid base deletion, nucleic acid base insertion, target nucleic acid breakage, nucleic acid methylation, and nucleic acid demethylation.
[0507] Another aspect of this disclosure relates to the use of the Cas12 protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the Cas12 inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the CRISPR-Cas12 system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, the cell as described in this disclosure, the pharmaceutical composition as described in this disclosure, or the kit as described in this disclosure in the diagnosis, treatment, or prevention of diseases or conditions related to the target nucleic acid.
[0508] Another aspect of this disclosure relates to the use of the Cas12 protein, the guide polynucleotide, the inactivated variant of Cas12, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas12 system, the vector system, the delivery system, the cell, the pharmaceutical composition, or the kit as described in this disclosure in the preparation of a medicament for the diagnosis, treatment, or prevention of a disease or condition associated with a target nucleic acid.
[0509] In some embodiments disclosed herein, the target nucleic acid is selected from genes listed in Table 27, and the disease or symptom is one listed in Table 27. Table 27 shows the target nucleic acids and the corresponding diseases or symptom for each target nucleic acid.
[0510] In some embodiments disclosed herein, the CRISPR-Cas12 system described herein is used to target and cleave specific genes listed in Table 27, thereby preventing, diagnosing, or treating diseases or conditions corresponding to those genes in Table 27. For example, after targeted cleavage, indels are formed through cell repair, and the target gene is knocked out, thereby inhibiting its function.
[0511] In some embodiments disclosed herein, the CRISPR-Cas12 system described herein is used to target and modify specific genes listed in Table 27, thereby preventing, diagnosing, or treating diseases or conditions corresponding to the genes in Table 27.
[0512] In some embodiments disclosed herein, the expression of specific genes listed in Table 27 is targeted and regulated by the CRISPR-Cas12 system described herein, thereby preventing, diagnosing, or treating the diseases or conditions corresponding to the genes in Table 27.
[0513] In some embodiments, the pharmaceutical composition is delivered in vivo to a human subject. The pharmaceutical composition can be delivered via any effective route. Exemplary routes of administration include, but are not limited to, intravenous infusion, intravenous injection, intraperitoneal injection, intramuscular injection, intratumoral injection, subcutaneous injection, intradermal injection, intraventricular injection, intravascular injection, cerebellar injection, intraocular injection, subretinal injection, intravitreal injection, intra-anterior chamber injection, intratympanic injection, intranasal administration, and inhalation.
[0514] Diagnostic applications
[0515] Another aspect of this disclosure relates to an in vitro composition comprising the CRISPR-Cas12 system described herein and detector DNA that cannot be used with the markers for guiding polynucleotide hybridization described herein.
[0516] Another aspect of this disclosure relates to the use of the CRISPR-Cas12 system described herein in detecting target nucleic acids in nucleic acid samples suspected of containing target nucleic acids.
[0517] Another aspect of this disclosure relates to the use of the CRISPR-Cas12 system described herein in detecting target nucleic acids in nucleic acid samples containing target nucleic acids.
[0518] In some implementations, the target nucleic acid to be detected is the target RNA.
[0519] In some embodiments, the target nucleic acid to be detected is target DNA. In some embodiments, the method for detecting target DNA includes a Cas12 protein fused to a fluorescent protein or other detectable marker, and a guide polynucleotide containing a guide sequence specific to the target DNA. The binding of Cas12 to the target DNA can be visualized by microscopy or other imaging methods.
[0520] In some implementations, methods for detecting target nucleic acids in cell-free systems result in the generation of detectable markers or enzyme activity. For example, by using the Cas12 protein, a guide polynucleotide containing a guide sequence specific to the target nucleic acid, and a detectable marker, the target nucleic acid will be recognized by Cas12. The binding of Cas12 to the target nucleic acid triggers its DNase activity, which leads to the cleavage of both the target nucleic acid and the detectable marker.
[0521] In some implementations, the detectable marker is DNA linked to a fluorescent probe and a quencher. The intact detectable DNA is linked to the fluorescent probe and the quencher, suppressing fluorescence. After the detectable DNA is cleaved by Cas12, the fluorescent probe is released from the quencher and exhibits fluorescent activity. This method can be used to determine the presence of target DNA in lysed cell samples, lysed tissue samples, blood samples, saliva samples, environmental samples (e.g., water, soil, or air samples), or other lysed cell or cell-free samples. This method can also be used to detect pathogens, such as viruses or bacteria, or to diagnose disease states, such as cancer.
[0522] In some implementations, the detection of target nucleic acids helps diagnose diseases and / or pathological conditions, or the presence of viral or bacterial infections.
[0523] Table 27: Target Nucleic Acids / Target Genes and Corresponding Diseases or Symptoms (When there are two or more target points in a specific cell, it indicates that two or more target genes are being targeted simultaneously; without limitation, the targeting can be targeted knockout, single-base editing, homologous recombination, targeted enhancement of transcription, targeted inhibition of transcription, etc.):
[0524]
[0525]
[0526]
[0527]
[0528]
[0529]
[0530]
[0531]
[0532]
[0533]
[0534]
[0535]
[0536]
[0537]
[0538]
[0539]
[0540]
[0541]
[0542]
[0543]
[0544]
[0545]
[0546]
[0547]
[0548]
[0549]
[0550]
[0551]
[0552]
[0553]
[0554]
[0555]
[0556]
[0557]
[0558]
[0559]
[0560]
[0561]
[0562]
[0563]
[0564]
[0565]
[0566]
[0567]
[0568]
[0569]
[0570]
[0571]
[0572]
[0573]
[0574]
[0575]
[0576]
[0577]
[0578]
[0579]
[0580]
[0581]
[0582]
[0583]
[0584]
[0585]
[0586]
[0587]
[0588]
[0589]
[0590]
[0591]
[0592]
[0593]
[0594]
[0595]
[0596]
[0597]
[0598]
[0599]
[0600]
[0601]
[0602]
[0603] Example
[0604] The present disclosure is further illustrated below by way of examples, but these examples are not intended to limit the scope of the disclosure to the examples described. Experimental methods in the following examples that do not specify specific conditions were performed according to conventional methods and conditions, or as selected in accordance with the product instructions.
[0605] Experimental Example 1: Screening of C12-279 Protein
[0606] like Figure 1 A and Figure 1 As shown in B, the Cas12i protein C12-279 was finally obtained through screening, analysis and design using complex bioinformatics methods.
[0607] The amino acid sequence of the C12-279 protein (SEQ ID NO: 696) is as follows:
[0608]
[0609] The proposed structure of the C12-279 protein is shown in Figure 24.
[0610] Experimental Example 2: Preparation and Purification of C12-279 Protein
[0611] 1. Carrier Construction
[0612] The pET28a vector plasmid was double-digested with BamHI and XhoI, and the linearized vector was recovered by agarose gel electrophoresis. Using the prepared pXC12-279-GFPPAM-DR5 plasmid (SEQ ID NO: 697) as a template, a DNA fragment containing the coding sequence of the C12-279 protein was obtained by PCR amplification using primers ChkCas-pET28-PF1 (SEQ ID NO: 699) and ChkCas-pET28-PR1 (SEQ ID NO: 700). Homologous recombination (NEB, Gibson) was then used to amplify the fragment. MasterMix was inserted into the cloning region of vector pET28a to construct the recombinant vector C12-279-pET28a-01 (SEQ ID NO: 698). Stbl3 competent cells were transformed with the reaction solution, plated on LB plates containing kanamycin sulfate, and incubated overnight at 37°C. Clones were then picked and sequenced for identification.
[0613] Positive clones with correct sequences were selected and cultured overnight. After plasmid extraction, the clones were transformed into the expression strain Rosetta(DE3), plated on LB agar plates containing kanamycin sulfate, and cultured overnight at 37°C.
[0614] 2. Protein expression
[0615] Select a single clone and inoculate it into 5 mL of LB medium containing kanamycin sulfate, and incubate overnight at 37°C.
[0616] Transplant the culture medium into 500 mL of LB medium containing kanamycin sulfate at a ratio of 1:100, incubate at 37 °C with a rotation speed of 220 rpm until OD 0.6, add IPTG to a final concentration of 0.2 mM, and induce at 16 °C for 24 h.
[0617] Wash with 15ml PBS, centrifuge to collect bacterial cells, add lysis buffer and sonicate to disrupt, centrifuge at 10000g for 30min to obtain supernatant containing recombinant protein, filter the supernatant through a 0.45μm filter membrane and then purify by column.
[0618] 3. Protein purification
[0619] The expressed recombinant protein has 1240 amino acids and its structure is His tag-NLS-C12-279-SV40 NLS-nucleoplasmin NLS. Purification was performed using the six N-terminal His molecules as purification tags via IMAC (Ni Sepharose 6Fast Flow, Cytiva) followed by heparin affinity chromatography (POROST). M Heparin 50μm column, Thermo Scientific TM The C12-279 recombinant protein was obtained after purification. The purified recombinant protein was analyzed by SDS-PAGE electrophoresis, and the results are as follows: Figure 2 As shown.
[0620] Experimental Example 3: In vitro cleavage of PAM library by C12-279 protein and capture of PAM
[0621] In this experimental example, sgRNA containing a specific guide sequence and C12-279 recombinant protein prepared and purified as described in Example 2 were mixed and used to cleave the in vitro cleavage substrate (containing a spacer sequence and a 7nt random sequence) (e.g. Figure 3 (As shown), after incubation at 37℃, the sample was purified, a library was constructed, and NGS sequencing and analysis were performed to determine the PAM sequence of C12-279. The specific steps are as follows:
[0622] A. External cutting substrate
[0623] The designed in vitro cutting substrate sequence is as follows:
[0624] ggagtteagacgtgtgetcttccgatctcageacaaaaggaaacteaecctaactgtaaagtaattgtgtgttttgagactataaatatgcatgcgagaaaagccttgtttgccaccatGGAACGGCTCGGAGATCAT CATTGCGNNNNNNNgtgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcagatcggaagagcacacgtctgaactcc(SEQ ID NO: 701).
[0625] In the sequence, N represents any one of A, T, C, or G.
[0626] Double-stranded DNA containing the above sequence was prepared using PCR amplification and used as an in vitro cleavage substrate.
[0627] The cleavage substrate was sent to a sequencing company for PCR-Free library construction and NGS sequencing. Complexity and abundance were analyzed for the PAM library composed of 7nt random sequences. The results are as follows:
[0628] The four bases A, T, G, and C have essentially the same composition; meanwhile, the PAM library, composed of 7nt random sequences, contains 4^7 = 16384 different combinations, all of which were detected. The complexity and abundance of the PAM library are satisfactory.
[0629] Preparation of B.sgRNA
[0630] A specific guide sequence-containing sgRNA (C12-279-sgRNA) was synthesized in vitro at 37°C using a DNA template system containing T7 RNA transcriptase, four ribonucleotide triphosphates, and a T7 promoter. The transcription product was precipitated and purified using LiCl. The sgRNA sequence is as follows:
[0631] >C12-279-sgRNA
[0632] 5'-gtaatgcgtctcccattgaegcGTGAGCAAGGGCGAGGAGCTGTTC-3' (SEQ ID NO: 702).
[0633] >C12-279-sgRNA-Rev
[0634] 5'-gtaatgcgtctcccattgacgcCGCAATGATGATCTCCGAGCCGTTCC-3' (SEQ ID NO: 703).
[0635] The sgRNA backbone sequence (direct repeat sequence, DR sequence) is: 5'-gtaatgcgtctcccattgacgc-3' (SEQ ID NO: 704).
[0636] The uppercase bases are the guide sequence for sgRNA.
[0637] C.NGS library construction and PAM analysis
[0638] 1) PAM library cleavage and T4 DNA Polymerase treatment
[0639] Reaction systems containing C12-279 protein, two different sgRNAs, in vitro cleavage substrates, and buffer were prepared (as shown in Table 1). The reaction was carried out at 37°C for 3 h and 75°C for 15 min.
[0640] Table 1. Reaction system for in vitro cleavage reaction
[0641] <![CDATA[10xCut Buffer(500mM Tris-HC1 PH8.0,2M NaCl,100mM MgCl2,10mM DTT)]]> 5 Cutting substrate (59.5 ng / μL) 36.5 C12-279-sgRNA or C12-279-sgRNA-Rev 2.8 μg C12-279 protein (10 mg / mL) 0.5
[0642] 2) T4 DNA Polymerase treatment to fill in the cleavage products
[0643] Add T4 DNA Polymerase (Thermo Scientific) to the digested product. The specific reaction system is shown in Table 2. After addition, react at 37°C for 20 min and then at 85°C for 10 min.
[0644] Table 2. Reaction system for C12-279 cutting product compensation
[0645] C12-279 Cutting Products 50 5xT4 DNA Polymerase Buffer 13 T4 DNA Polymerase 1 dNTPs (10mM) 0.65 <![CDATA[ddH2O]]> 0.35
[0646] 3) Adding an A to the 3' end and adding biotin-labeled adapters
[0647] a. Add 78 μL of SPRISelect Beads (Beckman COULTER) to the T4 DNA Polymerase reaction product, mix well, incubate at room temperature for 5 min, transfer the product to a magnetic rack for 5 min, transfer the supernatant to a new 1.5 mL tube; add 39 μL of SPRISelect Beads (Beckman COULTER), mix well, incubate at room temperature for 5 min, transfer the product to a magnetic rack for 5 min, discard the supernatant, wash twice with 85% ethanol, air dry at room temperature for 10 min, and elute with 50 μL of ddH2O.
[0648] b. Using the SynplSeq DNA Library Prep Kit for Illumina, perform 3' addition of A on the product in a according to the system in Table 3, 37℃ for 10 min, 65℃ for 20 min, and then set aside at 4℃.
[0649] Table 3. C12-279 Cutting Product 3' with A
[0650]
[0651]
[0652] c. Add Adapter 1 (obtained by annealing upstream primer: 5'Biosg / gttgacatgctggattgagacttcctacactctttccctacacgacgctcttccgatc*t (SEQ ID NO: 705) and downstream primer: gatcggaagagcgtcgtgtagggaaaga gtgtaggaagtctcaatccagcatgtcaac (SEQ ID NO: 706)) according to the system in Table 4, and react at 20℃ for 30 min, then overnight at 16℃. The reaction product was purified using SPRISelect Beads.
[0653] Biosg represents biotin modification, and * indicates thiolation.
[0654] Table 4. Reaction system with Adapter 1 added
[0655] The reaction products in b 60 Adapter 1 2 DNA Ligase 1.2 3x Ligation Bufer 1 33 <![CDATA[ddH2O]]> 3.8 Total 100
[0656] d. Magnetic beads labeled with streptavidin The reaction product was purified using M-280 Streptavidin (Invitrogen).
[0657] e.Recover PCR
[0658] Design the primers in Table 5, and utilize the system in Table 6 and the reaction procedure in Table 7. Hot Start High-Fidelity 2x Master Mix (NEB) was used for Recover PCR reaction.
[0659] Table 5. Recover PCR Primers
[0660] Recovery PCR Forward ggagttcagacgtgtgctc (SEQ ID NO: 707) Recovery PCR Reverse gttgacatgctggattgagacttc (SEQ ID NO: 708)
[0661] Table 6. Recover PCR reaction system
[0662] streptavidin-labeled magnetic bead purification products 22.5 Recovery PCR Forward (10uM) 2.5 Recovery PCR Reverse (10uM) 2.5 Q5Hot-Start 2x Master Mix 22.5 <![CDATA[ddH2O]]> Up to 50
[0663] Table 7. Recover PCR reaction procedure
[0664]
[0665] f. Transfer the Recovery PCR product to a magnetic rack and let it adsorb for 5 min. Transfer the supernatant to a new 1.5 ml centrifuge tube, take 3 μL of Recovery PCR product, and dilute with 148.5 μL of ddH2O.
[0666] g.Index PCR
[0667] Using the primers listed in Table 8, perform Index PCR according to the system in Table 9 and the reaction procedure in Table 10.
[0668] Table 8. Index PCR Primers
[0669] IF501 aatgataeggcgaccaccgagatctacactatagcctacactctttccctacacgacg (SEQ ID NO: 709) IR701 caagcagaagacggcatacgagatcgagtaatgtgactggagttcagacgtgtgctc (SEQ ID NO: 710)
[0670] Table 9. Index PCR reaction system
[0671] Recovery PCR dilution products 12 IF501 (10uM) 4 IR701 (10µM) 4 Q5 Hot-Start 2x Master Mix 20 Total 40
[0672] Table 10. Index PCR Reaction Procedure
[0673]
[0674] h.Index PCR products were purified by adding 0.7x SPRISelect Beads, eluted with 38ul ddH2O, and then sent for NGS sequencing after concentration determination using Qubit.
[0675] i. Analysis of NGS results: Following NGS sequencing and the method described in the reference (A compact Cas9 ortholog from Staphylococcus Auricularis (SauriCas9) expands the DNA targeting scope. PLoSbiology, 2020, 18(3), e3000686.), the results were analyzed using WebLogo software, yielding the following results: Figure 4 and Figure 5 The sequence shown is the one that has been captured. Figure 4 and Figure 5 Both studies have demonstrated that C12-279 can recognize PAM sequences of 5'-TTN-3'.
[0676] Experiment Example 4: Verifying PAM based on the bacterial in vivo editing activity of C12-279 protein
[0677] In this experimental example, a plasmid library containing a 7-nt random sequence was first constructed, followed by a bacterial expression plasmid containing the C12-279 protein-coding sequence. After transforming the expression plasmid into competent bacterial cells, the 7-nt random sequence plasmid library was electroporated. If the plasmid in the 7-nt random sequence library could be recognized and targeted by C12-279, it would be removed from the library, and the corresponding bacteria would not grow. Figure 6 As shown.
[0678] The specific steps are as follows:
[0679] 1.7nt random sequence plasmid library construction
[0680] The pLVX-EF1a-BSD vector plasmid (SEQ ID NO: 711) was digested with EcoRV and XhoI, and the linearized vector was recovered by agarose gel electrophoresis. Using the prepared pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid (SEQ ID NO: 712) as a template, the DNA fragment containing the coding sequence of the Puro resistance gene was amplified by PCR using primers Puro-PF1 (SEQ ID NO: 714) and Puro-PR1 (SEQ ID NO: 715). Homologous recombination (NEB, Gibson) was then employed to amplify the DNA. The Master Mix was inserted into the enzyme-digested pLVX-EF1a-BSD vector to construct the recombinant vector pLVX-7NN-Puro library plasmid (SEQ ID NO: 713) containing the 7NN random sequence. The reaction solution was transformed into Stbl3 competent cells, plated on LB agar plates containing ampicillin, and incubated overnight at 37°C. All colonies were scraped off for plasmid extraction.
[0681] 2. Construction of the bacterial expression plasmid P15A-C12-279 for the C12-279 protein, and preparation of competent cells containing this plasmid.
[0682] a. Construction of bacterial expression plasmid for C12-279 protein
[0683] The p15A-Cas-NC03 vector plasmid (SEQ ID NO: 716) was digested with SalI, and the linearized vector was recovered by agarose gel electrophoresis. Using the prepared pXC12-279-GFPPAM-DR5 plasmid (SEQ ID NO: 697) as a template, primers ChkNLS-PF1 (SEQ ID NO: 718) and the synthesized C12-279-sgRNA fragment (SEQ ID NO: 717) were used to amplify the DNA fragment encoding the C12-279 protein and the fusion fragment expressing the sgRNA via PCR. Homologous recombination (NEB, Gibson) was then employed. The Master Mix was inserted into the enzyme-digested P15A-Cas-NC03 vector to construct the recombinant vector P15A-C12-279 (SEQ ID NO: 719). The reaction solution was used to transform Stbl3 competent cells, plated on LB plates containing chloramphenicol, and incubated overnight at 37°C. Single clones were then picked and sequenced for verification.
[0684] b. Preparation of competent cells containing bacterial expression plasmid P15A-C12-279
[0685] The P15A-C12-279 plasmid, which was verified by sequencing, was transformed into DH5a competent cells (Weidi Bio, CAT#: DL1001). Single clones were picked and inoculated into LB medium containing chloramphenicol and cultured overnight at 37°C.
[0686] Electrocompetent states were prepared according to the following steps:
[0687] The bacterial culture was inoculated into 100 mL of fresh LB medium containing chloramphenicol at a ratio of 1:100 and cultured at 37°C and 220 rpm for scale-up.
[0688] Incubate until OD600 = 0.5, then transfer the bacterial culture to a 50 mL centrifuge tube and pre-cool on ice for 30 min.
[0689] Centrifuge at 4000 rpm and 4℃ for 10 min, collect the cells, and resuspend them in an equal volume of pre-cooled sterile water;
[0690] Repeat the above steps;
[0691] Resuspend cells in 1 / 10 volume of pre-cooled sterile water containing 10% glycerol, aliquot into 50 μL tubes, and store at -80°C to obtain P15A-C12-279 competent cells.
[0692] c. Plasmid elimination to identify the PAM sequence of the C12-279 protein.
[0693] 100 ng of pLVX-7NN-Puro library plasmid was used to electroporate P15A-C12-279 competent cells and DH5α competent cells, which were labeled as Lib1 (electroplated P15A-C12-279 competent cells) and Lib2 (electroplated DH5α competent cells), respectively.
[0694] After electroporation, add 10 mL of LB medium and incubate at 37°C and 220 rpm for 2 h to recover and culture.
[0695] After resuscitation, the bacterial culture was centrifuged at 4000 rpm for 2 min to collect the bacterial cells. After resuspending in 400 μL LB, the cells were spread on LB plates. The DH5α bacterial culture was spread on LB plates containing ampicillin, and the P15A-C12-279 competent cells were spread on LB plates containing chloramphenicol and ampicillin. The cells were incubated overnight at 37°C.
[0696] Bacterial cells were scraped from the culture plate and plasmid DNA was extracted using the alkaline lysis method.
[0697] 100 ng of each of the two extracted plasmid DNA samples was used as PCR templates. PCR amplification was performed using primers SiteSeq-PF1 (SEQ ID NO: 720) and SiteSeqPuro-PR (SEQ ID NO: 721). The obtained fragments were used to construct an amplicon library using an NGS library construction kit (SynplSeq DNA Library Prep Kit for Illumina). The constructed library was then subjected to NGS sequencing.
[0698] The differences in NGS sequencing results after transforming Lib1 and Lib2 cells were compared and analyzed, such as... Figure 7 As shown. This demonstrates that C12-279 can recognize PAM sequences of 5'-TTN-3'.
[0699] Experimental Example 5: Cleavage activity of C12-279 protein on target nucleic acids in 293T cells
[0700] In this experiment, an sgRNA targeting the TTK gene in HEK293T cells was first designed and constructed into the plasmid pXC12-279-GFPPAM-DR5 (SEQ ID NO: 697), resulting in the plasmid pXC12-279-TTR01 targeting the TTR gene. After transfection into HEK293T cells, the indel ratio was verified by NGS, and the cleavage activity of C12-279 protein in 293T cells was verified. The specific steps are as follows:
[0701] 1. Construction of sgRNA plasmid
[0702] Based on the TTR gene sequence information in the HEK293T cell line, a sgRNA target catgagcatgcagaggtgagtat (SEQ ID NO: 722) was designed.
[0703] Table 11. sgRNA annealing primers
[0704]
[0705] Annealing primers were designed based on the sgRNA target information. Plasmid pXC12-279-GFPPAM-DR5 (SEQ ID NO: 697) was annealed using BsmBI (Thermo Scientific). TM , ER0451) and Acc65I (Thermo Scientific TM After digestion with ER0901 enzyme, the primers (Table 11) were annealed and ligated with the vector to obtain the pXC12-279-TTR01 expression clone.
[0706] 2. TTR gene editing efficiency detection
[0707] Plating: The 293T cell line was plated when the confluence reached 70-80%, and the number of cells seeded in a 24-well plate was 5*10^5 cells / well.
[0708] Transfection: After 12-14 hours of plating, transfect the cells by adding 100 μL of Opti-MEM, 1.5 μL of PEI (Yisheng Bio, Polyethylenimine Linear (PEI) MW25000), and 500 ng of pXC12-279-TTR01 plasmid to each well of a 24-well plate. Mix well, incubate at room temperature for 20 minutes, and then add the mixture to 293T cells for transfection. After transfection overnight, replace the medium with fresh medium and continue culturing.
[0709] DNA extraction, PCR amplification, and NGS library construction: After culturing for 72 hours, cells were washed with PBS, and then 100 μL of cell lysis buffer (Viagen) was added. Lysis was performed using Lysis Reagent (Cell) to obtain a lysate containing genomic DNA. The genomic DNA was then amplified near the target sequence. The PCR products were subjected to NGS library construction, sequencing, and analysis of the sequencing results. Indels were higher than 14%. For example... Figure 8 As shown.
[0710] Experiment Example 6: PAM Recognition of CI1062732
[0711] Protein CI1062732 (SEQ ID NO: 46) has dozens of amino acid residues missing from its N-terminus compared to protein C12-279 (SEQ ID NO: 696).
[0712] The inventors previously used essentially the same method as in Experimental Example 4 to identify the PAM sequence identified by CI1062732. This was done using a DR sequence containing "false" DR sequences. GTAATGCGTCTCCCATTGACGC The gRNA of C (SEQ ID NO: 529) combined with CI1062732 targets a 7nt random sequence plasmid library in bacteria. Sequencing analysis (e.g.) Figure 9 As shown in the figure, the “fake” PAM motif captured by CI1062732 is 5'-TTNC-3'.
[0713] Subsequently, the inventors analyzed the secondary structure of the "pseudo" DR sequence and speculated that the 3' C base of the captured "pseudo" PAM motif TTNC might be due to an extra C at the 3' end of the DR sequence (Figure 10). Therefore, in the subsequent CI1062732 and C12-279 tests, the DR sequence was used... GTAATGCGTCTCCCATTGACGC (SEQ ID NO: 704).
[0714] Experiment Example 7: Screening, design, preparation, and testing of C12-101-07 protein
[0715] (1) The C12-101-07 protein was finally obtained by screening and designing using complex bioinformatics methods.
[0716] The amino acid sequence of the C12-101-07 protein is as follows:
[0717]
[0718] The recombinant vector C12-101-07-pET28a-01 (SEQ ID NO: 725) was constructed using essentially the same method as in Experimental Example 2. The recombinant protein C12-101-07, with 1122 amino acids and the structure His tag-NLS-C12-101-07-SV40 NLS-nucleoplasmin NLS, was expressed and purified. The purified recombinant protein was analyzed by SDS-PAGE electrophoresis, and the results are shown in Figure 11.
[0719] (2) The C12-101-07 protein in vitro cleavage PAM library was prepared using the same method as in Experiment 3, and PAM was captured.
[0720] In vitro transcription synthesizes sgRNA containing a specific guide sequence, the sequence of which is as follows:
[0721] >C12-101-07-sgRNA
[0722] 5'-atcgcaacatctcagaaacccgtcctaagttgacggGTGAGCAAGGGCGAGGAGCTGTTC-3' (SEQ ID NO: 726).
[0723] >C12-101-07-sgRNA-Rev
[0724] 5'-atcgcaacatctcagaaacccgtcctaagttgacggCGCAATGATGATCTCCGAGCCGTTCC-3' (SEQ ID NO: 727).
[0725] The sgRNA backbone sequence is: 5'-atcgcaacatctcagaaacccgtcctaagttgacgg-3' (SEQ ID NO: 534).
[0726] The C12-101-07 protein was edited by combining it with forward and reverse gRNAs, respectively.
[0727] The PAM motifs shown in Figures 12 and 13 were captured in the experiment.
[0728] Prove that C12-101-07 can recognize PAM sequences of 5'-TTN-3'.
[0729] Experiment Example 8: Bacterial in vivo editing activity of C12-101-09 protein, verifying PAM
[0730] A similar protein, C12-101-09, was designed based on C12-101-07. Its sequence is as follows:
[0731] >C12-101-09
[0732]
[0733] The recombinant vector plasmid P15A-C12-101-09 (SEQ ID NO: 729) was constructed using conventional methods. It was then tested using essentially the same method as in Experiment 4.
[0734] The motif identified by C12-101-09 was captured, as shown in Figure 14.
[0735] Prove that C12-101-09 can recognize PAM sequences of 5'-TTN-3'.
[0736] Experiment Example 9: Construction of cell lines containing different PAM reporter systems
[0737] Stable cell lines were prepared using lentiviral infection. We constructed lentiviral expression plasmids containing different PAM sequences, packaged the virus, and then infected 293T cells to build cell lines containing different PAM sequences for subsequent mutant screening.
[0738] (1) Construction of lentiviral expression plasmids for GFP reporter system composed of different PAM sequences
[0739] In addition to the prepared plasmid pCDH-CMV-EGFP-reporter3-EF1-Puro (SEQ ID NO: 712), a plasmid library with PAM composition of NAAN was also constructed. The specific construction scheme is as follows:
[0740] Gene synthesis was performed using a mixed library of EGFP fragments containing the detection system. The fragments were synthesized by digestion with XbaI+NotI enzymes, and then ligated with the XbaI+NotI digestion vector of plasmid pCDH-CMV-EGFP-reporter3-EF1-Puro using T4 DNA ligase. The resulting solution was transformed into Stb13 cells and cultured overnight at 37°C on ampicillin-containing plates.
[0741] For fragment mixing libraries, multiple clones need to be selected for sequencing and identification to obtain plasmids with 16 different sequence compositions (i.e., PAM sequences of AAAA, AAAT, AAAG, AAAC, TAAA, TAAT, TAAG, TAAC, GAAA, GAAT, GAAG, GAAC, CAAA, CAAT, CAAG, CAAC). Then, they are mixed in equal quantities to obtain a plasmid library Plasmid Lib (Puro-NAAN-eGFP-Lib, SEQ ID NO: 730) with PAM sequence composition of NAAN, which is used for subsequent lentivirus packaging and stable cell line construction.
[0742] The unedited reporter system has a non-multiple-of-3 base insertion between the start codon and the normal GFP reading frame (pCDH-CMV-EGFP-reporter3-EF1-Puro has a 32bp base insertion, and the plasmid library Plasmid Lib has a 29bp base insertion, as shown in Figure 15), which interrupts the normal GFP reading frame and prevents GFP expression. By setting the sgRNA target inside the GFP expression frame, indels are generated through Cas12 editing, which have a chance to restore the normal GFP reading frame and enable GFP to be expressed normally. The higher the editing efficiency, the higher the probability that the generated indels will restore the correct GFP reading frame. The editing efficiency of the Cas12 protein can be characterized by detecting the number of cells that can express GFP normally by flow cytometry.
[0743] (2) Packaging of GFP reporter system plasmids with different PAM sequences for lentivirus
[0744] The constructed pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid and library plasmid PlasmidLib were mixed with viral packaging helper plasmids pMD2.G (Miaoling Biotechnology) and psPAX2 (Miaoling Biotechnology) at a molar ratio of 1:1:1, and then transfected into 293T cells using PEI. After 48 h of transfection, the culture supernatant was collected and filtered through a 0.45 μm filter to obtain two crude viruses: pCDH-CMV-EGFP-reporter3-EF1-Puro and PlasmidLib.
[0745] (3) 293T cells were infected with crude pCDH-CMV-EGFP-reporter3-EF1-Puro and Plasmid Lib viruses to construct and detect cell lines.
[0746] 293T cells were infected by adding 1 / 4 volume of pCDH-CMV-EGFP-reporter3-EF1-Puro and PlasmidLib crude viruses to the culture medium. After 48 hours of infection, the medium was changed and 2 μg / ml of Puromycin was added for selection.
[0747] For 293T cells infected with pCDH-CMV-EGFP-reporter3-EF1-Puro, the selected cells were subjected to monoclonal screening using limiting dilution, and the selected monoclonal cells were used as the detection cell line (called the Reporter3 cell line).
[0748] For 293T cells infected with Plasmid Lib, the cell pool obtained after drug screening is the cell line used for detection (called NAAN cell line).
[0749] Experiment Example 10: Design and Editing Efficiency Detection of C12-279 Protein Mutants
[0750] a. Determination of mutation sites
[0751] The 3D structure of the C12-279 protein was predicted and simulated using bioinformatics analysis and AI techniques. The possible DNA binding, recognition, and cleavage sites of C12-279 were analyzed based on the 3D structure. Mutant clones were constructed using molecular cloning point mutagenesis methods targeting these sites.
[0752] First, the mutants shown in Table 12 were designed.
[0753] Table 12. Editing efficiency of different C12-279 mutation sites in NAAN cell lines
[0754]
[0755]
[0756] b. Construction of mutant clones
[0757] After identifying the specific mutation site, the mutated base is introduced using primers to construct expression clones containing different mutation sites. The following explanation uses the construction of the D352R mutant clone as an example.
[0758] Table 13. Primer sequences for constructing the C12-279 mutant clone C12-279-GFPPAM-01
[0759]
[0760] Primers were designed targeting the D352R mutation site (as shown in Table 13), and the mutation site was introduced using primers C279-D352R-PF1 and C279-D352R-PR1. PCR amplification was performed using pXC12-279-GFPPAM-DR5 (SEQ NO: 1) as a template and ChkCas12-PF1+C279-D352R-PR1 (Yijin Biotechnology, PC019, UltraHiPF). TM The DNA Polymerase Kit was used to obtain the fragment C12-279-D352R-F1. ChkCas12-PR1+C279-D352R-PF1 was used for PCR amplification to obtain the fragment C12-279-D352R-F2. The plasmid pXC12-279-GFPPAM-DR5 (SEQ ID NO: 697) was digested with HindIII+KpnI and the 5646bp vector fragment was recovered via gel extraction (Guangzhou Meiji Biotechnology Co., Ltd., D2110, HiPure Gel Pure Micro Kit). In vitro recombination was then performed with fragments C12-279-D352R-F1 and C12-279-D352R-F2 (NEB, E2611L, Gibson). The mutant cloning plasmid C12-279-GFPPAM-01 was obtained by heat shock transformation of E. coli (Master Mix).
[0761] c. Editing efficiency test of mutants in the NAAN reporting system
[0762] Plating: When the NAAN cell line reaches 70-80% confluence, it is plated at a rate of 5*10^5 cells / well in a 24-well plate.
[0763] Transfection: After 12-14 hours of plating, transfection was performed by adding 1.5 μL of PEI (Yisheng Bio, 40815ES03, Polyethylenimine Linear (PEI) MW25000) + 500 ng of mutant cloning plasmid to 100 μL of Opti-MEM per well in a 24-well plate. The mixture was incubated at room temperature for 20 minutes and then added to the NAAN cell line for transfection. After transfection overnight, the medium was replaced with fresh medium and cultured for 72 hours. Flow cytometry was used to detect the cell editing efficiency of different mutant clones based on the proportion of GFP-positive cells.
[0764] The results are shown in Table 16 and Figure 16. The editing efficiency of different C12-279 mutants in NAAN cells is expressed as a multiple of the editing efficiency compared to C12-279. Several mutants improved the editing efficiency.
[0765] d. Based on the results of the first round of mutations, some mutation sites that significantly improve editing efficiency were combined to try multiple mutation site combinations in order to further improve editing efficiency. The designed mutants are shown in Table 14.
[0766] Table 14. Editing efficiency of different combinations of C12-279 mutation sites in NAAN cell lines
[0767] D352R Q186R C12-279-GFPPAM-28 D352R L260R C12-279-GFPPAM-29 D352R A355R C12-279-GFPPAM-30 A355R L260R C12-279-GFPPAM-31 P386R C385R C12-279-GFPPAM-32 E485R Q462R C12-279-GFPPAM-33
[0768] The vector was constructed using the same method as described above, and the editing efficiency was tested in the NAAN cell line. The results are shown in Figure 17.
[0769] e. Construction of PAM combinatorial mutant plasmids for targeted fixation
[0770] Since the PAM in the NAAN cell line is a mixed library, it may theoretically have some impact on the actual editing efficiency. In order to more intuitively and efficiently demonstrate the impact of mutations on editing efficiency, based on the sequence between the start codon and the GFP reading frame in the Reporter3 cell line, the PAM sequence was selected to form a TTG, and the target sequence was CTCACCTCGCGACGCAATGATG (SEQ ID NO: 735) for subsequent editing efficiency testing.
[0771] Construction of mutant clones targeting Reporter3 cell lines
[0772] Based on C12-279-GFPPAM-DR5, the sgRNA was modified to target CTCACCTCGCGACGCAATGATG (SEQ ID NO: 735). Primers pCDH-PF1: GTACCGAAAAACATCATTGCGTCGCGAGGTGAGGCGTC (SEQ ID NO: 754) and pCDH-PR1: CATTGACGCCTCACCTCGCGACGCAATGATGTTTTTCG (SEQ ID NO: 755) were synthesized and annealed. The vector was recovered by digestion with Acc65I (Thermo Scientifc) and BsmBI (Thermo Scientifc) and then ligated with the annealing product to obtain the mutant clone C12-279-pCDH targeting the Reporter3 cell line. The construction method for mutant clones targeting the Reporter3 cell line is the same as that described in this experimental example, but the vector plasmid needs to be changed from C12-279-GFPPAM-DR5 to C12-279-pCDH, and the primer ChkCas12-PR1 needs to be replaced with ChkCas12-PR2: GCGACCAATGATGTTTTTCGGTACC (SEQ ID NO: 736). The rest of the methods and steps are completely the same.
[0773] Editing efficiency was tested on a subset of mutants in the Reporter3 cell line using the same method as the editing efficiency test in the NAAN reporter system, except that the cell line was changed from the NAAN cell line to the Reporter3 cell line.
[0774] Based on the results of the first round of mutations, the entire 3D structure was corrected and labeled. The data from the first round was then placed in a new model for predictive analysis. Finally, potential mutation sites for the second round were analyzed and predicted, and mutations and detections were performed. This process was repeated, involving multiple rounds of mutation, selection, and accumulation, to determine the optimal mutation combination for C12-279. Table 15 shows the C12-279 mutants and their editing efficiency in the Reporter3 cell line. The absolute value (average) of the editing efficiency of the C12-279-pCDH group, i.e., the C12-279 protein, was 8.55%.
[0775] Table 15. Editing efficiency of the C12-279 mutant in the Reporter3 cell line
[0776] (Expressed as a multiple of editing efficiency compared to C12-279)
[0777]
[0778]
[0779] Experimental Example 11: Design and Editing Efficiency Detection of C12-101-07 Protein Mutants
[0780] Mutants were designed for the C12-101-07 protein, as shown in Table 16.
[0781] The vector plasmid of the mutant was constructed using a method essentially the same as that used in Experiment 10, and the editing efficiency was tested.
[0782] a. Editing efficiency for detecting single-point mutations based on the NAAN reporting system
[0783] For the selected mutation sites, C12-101-07-GFPPAM (SEQ ID NO: 737) was used as the cloning template plasmid and the control plasmid for transfection testing. Mutant clones were constructed using molecular cloning point mutagenesis, and editing efficiency was assessed. The results are shown in Table 16. The editing efficiency of different C12-101-07 mutants in the NAAN cell line is expressed as a multiple of the editing efficiency compared to C12-101-07. The absolute value (average) of the editing efficiency of the C12-101-07-GFPPAM group (i.e., the C12-101-07 protein) was 0.23%.
[0784] Table 16. Different mutants of C12-101-07
[0785]
[0786] b. Combining the mutation sites that significantly improved editing efficiency from the first round of mutations, we attempted two point mutations. The mutants shown in Table 17 were designed.
[0787] Mutant clones were constructed using molecular cloning point mutagenesis and their editing efficiency was determined based on the NAAN cell line.
[0788] The results are shown in Table 17. The editing efficiency of different C12-101-07 mutants in the NAAN cell line is expressed as a multiple of the editing efficiency of C12-101-07.
[0789] Table 17. C12-101-07 mutants
[0790]
[0791] c. Construction of mutant clones targeting Reporter3 cell lines and detection of editing efficiency
[0792] Using a method essentially the same as in Experiment 10, the target sequence on the original C12-101-07-GFPPAM vector was replaced, and the target targeting the NAAN cell line was replaced with the target targeting the Reporter3 cell line, resulting in the control plasmid 101-07-sgRNA02. Then, the mutant vector was constructed, and the editing efficiency was tested in the Reporter3 cell line.
[0793] The results are shown in Table 18. The editing efficiency of different C12-101-07 mutants in the Reporter3 cell line is expressed as a multiple of the editing efficiency compared to C12-101-07. The absolute value (mean) of the editing efficiency of the 101-07-sgRNA02 control group, i.e., the C12-101-07 protein, was 10.32%.
[0794] Table 18. Editing efficiency of different mutants of C12-101-07 in the Reporter3 cell line
[0795] (Expressed as a multiple of the editing efficiency compared to C12-101-07)
[0796]
[0797] Experiment Example 12: Constructing Different Mutants from C12-279 Protein by Site-Directed Mutation
[0798] The amino acid sequence of the C12-279 protein is SEQ ID NO: 696. Two mutation primers, F / R, were designed at the mutation site to introduce the desired mutation sequence. PCR amplification was performed using universal primers at both ends of the vector to obtain two mutant fragments, F1 and F2. Mutant fragments F1 and F2 were then homologously recombinated with the linearized vector to obtain a mutant plasmid. In this experimental example, single-point mutations (D426R, L860R) and multi-point mutations (D426R & L860R) were constructed using amino acid sites 426 and 860. The specific steps are as follows:
[0799] The primers for site-directed mutations at sites 426 and 860 are shown in Table 19.
[0800] Table 19. Primers
[0801]
[0802] Cloning of single-point mutations at sites 426 and 860
[0803] PCR amplification was performed using the C12-279-pCDH plasmid as a template and the ChkCas12-PF1+279_426_R plasmid (Yijin Biotechnology, PC019, UltraHiPF). TMThe mutant fragment D426R-F1 was obtained using a DNA Polymerase Kit. ChkCas12-PR2+279_426_F was amplified by PCR to obtain the mutant fragment D426R-F2. Plasmid C12-279-pCDH was digested with HindIII+KpnI, and the 5647bp vector fragment was recovered using a gel (Guangzhou Meiji Biotechnology Co., Ltd., D2110, HiPure Gel Pure Micro Kit). This fragment was then used for in vitro recombination with fragments D426R-F1 and D426R-F2 (NEB, E2611L). Master Mix), heat shock transformation of E. coli yielded plasmid C12-279-pCDH-426, a mutant at the 426 site;
[0804] Using the same steps and methods, replace primers 279_426_F with 279_860_F and 279_426_R with 279_860_R, perform the same amplification and recombination, and heat-transform E. coli to obtain the plasmid C12-279-pCDH-860 for the 860 site mutant. The construction methods for plasmids of Cas protein mutants at other different sites are the same as those for the 426 and 860 sites, except that mutant primers for each different site need to be designed.
[0805] Validation of Cas protein mutant editing activity
[0806] The Reporter3 cell line constructed in the aforementioned experimental example was used for testing. The sequence between the start codon and the GFP reading frame was 32 bases, not a multiple of 3, leading to abnormal GFP reading and no GFP expression. A Cas protein mutant was used to edit the Reporter3 cell line, targeting the PAM sequence TTG and the target sequence CTCACCTCGCGACGCAATGATG (SEQ ID NO: 735), generating indels and restoring the normal GFP reading frame. Flow cytometry was used to detect the proportion of cells with restored GFP expression, thus characterizing the editing efficiency of different mutants. The specific steps are as follows:
[0807] Cell culture and plating: When the cell line reaches 70-80% confluence, it is plated, and the number of cells seeded in a 24-well plate is 5*10^5 cells / well.
[0808] Transfection: After 12-14 hours of plating, transfection was performed by adding 1.5 μL of PEI (Yisheng Biotechnology) + 500 ng of mutant plasmid to 100 μL of Opti-MEM per well in a 24-well plate. After mixing and incubation at room temperature for 20 minutes, the mixture was added to the Reporter3 cell line for transfection. After overnight transfection, the medium was replaced with fresh medium and cultured for 72 hours. Flow cytometry analysis was performed to detect the editing efficiency of different mutant clones based on the proportion of GFP-positive cells, and the average value of multiple batches of data was taken. Specific results are shown in Table 20 and Figure 18. The mean absolute value of the editing efficiency of the wild-type C12-279 group was 9.0%.
[0809] Table 20. Editing efficiency of each C12-279 mutant in the Reporter3 cell line
[0810] (Expressed as a multiple of editing efficiency compared to C12-279)
[0811]
[0812]
[0813]
[0814]
[0815]
[0816]
[0817]
[0818]
[0819]
[0820]
[0821] In the table, nR notation (n is an integer) indicates that the nth amino acid defect mutation is R.
[0822] Construction of clones of Cas variant proteins with different combinations of mutation sites, and verification of editing activity:
[0823] Based on the editing efficiency results of different mutation sites mentioned above, multiple mutation combinations were selected for different mutation sites. The plasmid construction of the Cas mutant with two sites of 426 and 860 was introduced as an example.
[0824] Using C12-279-pCDH plasmid as a template, PCR amplification of ChkCas12-PF1+279_426_R yielded the mutant fragment D426R-F1; PCR amplification of 279_426_F+279_860_R yielded the mutant fragment D426R-L860R-F1; and PCR amplification of ChkCas12-PR2+279_860_F yielded the mutant fragment L860R-F2. Plasmid C12-279-pCDH was digested with HindIII+KpnI, and the 5647bp vector fragment was recovered by gel extraction. It was then recombined in vitro with fragments D426R-F1, D426R-L860R-F1, and L860R-F2. The resulting variant plasmid C12-279-pCDH-426-860 was obtained by heat shock transformation of E. coli. The construction methods of other multi-point mutant plasmids were the same as those for C12-279-pCDH-426-860.
[0825] The amino acid sequence of the Mut-02-1-426-846-858-860 mutant is as follows:
[0826]
[0827] The editing activity of the multipoint mutant was verified using the same method described in the previous experimental example.
[0828] The results showed that the editing efficiency of most multi-point mutants was improved compared with that of the wild type, as shown in Table 21 and Figure 19. Among them, the absolute value of the editing efficiency of the wild type C12-279 group was 8.7%.
[0829] Table 21. Editing efficiency of each C12-279 mutant in the Reporter3 cell line
[0830] (Expressed as a multiple of editing efficiency compared to C12-279)
[0831]
[0832]
[0833]
[0834] In the table, the integer n indicates that the disabled mutation at position n is R. Mut-01 is the Q186R variant, Mut-02 is a Q186R&D352R double-point mutant, and Mut-03 is a G184R+Q186R+D352R triple-point mutant; 5+426+860 indicates that positions 5, 426, and 860 are simultaneously mutated to R; Mut-02+860 indicates a multi-point mutant obtained by introducing an additional mutation at position 860 (mutating to arginine R) on the basis of the Mut-02 mutant; other multi-point mutants follow the same pattern (mutating to arginine R).
[0835] Experimental Example 13: Detection of Editing Efficiency of Different Mutants Targeting Different Sites of TTR and HBG
[0836] In this experiment, we first selected the PAM sequence as the target site for TTN in the TTR gene (GeneBank: NG_009490.1) and HBG gene (GeneBank: NC_000011.10), and designed and constructed sgRNAs targeting different locations (as shown in Table 22) to be combined with different mutants to test the editing efficiency.
[0837] Table 22. sgRNAs targeting TTR and HBG
[0838]
[0839]
[0840]
[0841] Based on the sgRNA sequences in the table above, sgRNA plasmids were constructed respectively: the vector plasmid SpCas9-gRNA-pUC57Kan was linearized using BbsI (Thermofisher) and XhoI (Thermofisher), primers were synthesized for different sgRNA sequences, and after annealing, it was ligated into the linearized vector and transformed into E. coli to obtain the final sgRNA expression vector plasmid.
[0842] Different mutant plasmids were combined with different sgRNA plasmids and HEK293T cells were transfected with PEI. After 48 hours, the cells were collected and lysed using DirectPCR Lysis Reagent (Cell) (VIAGEN: 302-C). Different primers were selected according to different targets for PCR amplification. Sanger sequencing was performed and the editing efficiency was analyzed using TIDE.
[0843] The specific amplification and sequencing primers are shown in Table 23. The experiments were repeated three times, and the editing efficiency results are shown in Figures 20A and 20B. Each sgRNA group produced effective editing in each batch of experiments, with editing efficiency superior to wild-type C12-297 at many target sites.
[0844] Table 23. Primers for amplification and sequencing of different targets
[0845]
[0846] Experimental Example 14: Preparation and Activity Detection of Inactivated Version of C12-279 Mutant
[0847] Based on the multi-point mutants designed in the aforementioned experimental examples, one or more of the D651A, E891A, and D1082A mutations were further introduced. Using the same method as in Experiment 10, mutant cloning plasmids were constructed and edited in the Reporter3 cell line. The proportion of cells that restored GFP expression after editing was analyzed by flow cytometry to determine the mutant editing efficiency. Specific results are shown in Figure 21. In the figure, if any mutant contains at least one of the D651A, E891A, or D1082A mutations, the editing efficiency drops below 1.3%. Introducing any one of D651A, E891A, or D1082A yields deadCas12 (dCas12).
[0848] The amino acid sequence of the dCas mutant Mut-02-1-426-846-858-860-D651A-E891A-D1082A is as follows:
[0849]
[0850] Experimental Example 15: mRNA+gRNA Delivery
[0851] The TTR gene was edited using the encoding mRNA of the Mut-02-3-426-860 mutant (obtained through in vitro transcription) along with modified gRNA (as shown in Table 24).
[0852] Table 24. Modified gRNAs
[0853]
[0854] In the table, "r" represents a natural base, "d" represents deoxygenation modification, "m" represents methylation modification, and "*" represents thiophosphate modification.
[0855] HEK293 cells were transformed with mRNA and gRNA by electroporation. After 48 hours, cells were collected, lysed using DirectPCR LysisReagent (Cell) (VIAGEN: 302-C), and sequenced by Sanger sequencing. Editing efficiency was analyzed using TIDE. The results are shown in Table 24.
[0856] Adding a DNA nucleotide sequence to the 3' end significantly enhances editing activity. For example, adding a sequence...
[0857] dAdTdGdTdGdTdTdTdTdTdGdTdCdAdAdAdAdGdAdCdCdTdTdTdT (SEQ ID NO: 821) or dGdTdTdGdCdAdAdTdCdCdCdAdAdG (SEQ ID NO: 822).
[0858] In addition, PCR amplification was performed using primers TTR-NGS-PF2 (GCGTAACTTAATCCAGACTTTCACACCTT, SEQ ID NO: 823) and TTR-NGS-PR2 (GGTCATTCATCACCTTCCTTAGGACA, SEQ ID NO: 824), followed by library construction and NGS sequencing to detect editing efficiency. The combined electroporation editing efficiency of the mutant Mut-02-3-426-860 and C279-dmTTR01-01 reached 79.64%.
[0859] Subsequently, C279-dmTTR01-02 was tested in combination with different mutants. The specific NGS sequencing results are shown in Table 25 and Figure 22.
[0860] Table 25. Editing efficiency of modified gRNA (C279-dmTTR01-02) with different mutant combinations (NGS assay)
[0861] Mut-02-426-860 C279-dmTTR01-02 24.63% Mut-02-5-426-860 C279-dmTTR01-02 63.48% Mut-02-1-5-426-858-860 C279-dmTTR0 1-02 92.18% Mut-02-1-426-846-858-860 C279-dmTTR0 1-02 88.05% Mut-02-1-3-426-858-860 C279-dmTTR01-02 89.67% Mut-02-3-426-860 C279-dmTTR0 1-02 79.99%
[0862] Experiment Example 16: Validation of Editing Efficiency of Different Endogenous Target Genes
[0863] Different gRNA targets were designed based on different disease-related targets (Table 26). The mutant Mut-02-1-426-846-858-860 protein was edited by combining it with the gRNAs in Table 26. The same methods as in Example 13 were used, including plasmid construction, transfection, and editing efficiency detection. The results are shown in Table 26.
[0864] Table 26. Editing efficiency targeting different endogenous genes
[0865]
[0866]
[0867]
[0868]
[0869] Experimental Example 17: PAM recognition of the C12-279 mutant Mut-02-1-426-846-858-860
[0870] Using the same method as in the previous embodiments, in vivo bacterial editing experiments were conducted, and finally, NGS sequencing confirmed that the PAM recognized by the mutant Mut-02-1-426-846-858-860 was 5'-WTN-3' (W is A or T). As shown in Figure 23.
[0871] While specific embodiments of this disclosure have been described above, those skilled in the art should understand that these are merely illustrative examples, and various changes or modifications can be made to these embodiments without departing from the principles and essence of this disclosure. Therefore, the scope of protection of this disclosure is defined by the appended claims.
Claims
1. A Cas12 protein, characterized in that, The Cas12 protein is any one of the following (a)-(b): (a) The amino acid sequence of the Cas12 protein is SEQ ID NO: 696; or (b) The Cas12 protein in SEQ ID NO: Based on the sequence shown in 696, perform any of the following mutation combinations: 186+352+1+426+846+858+860, 186+352+1+426+846+860, 186+352+1+5+426+858+860, 186+352+1+3+426+858+860, 186+352+3+426+860, 186+352+1+333+426+858+860, 186+352+1+426+485+858+860, 186+352+1+5+426+860, 186+352+3+426+858+860, 186+352+5+42 6+858+860, 186+352+333+426+858+860, 186+352+333+426+860, 186+352+426+846+860, 186+352+5+426+860, 186+352+1+3+426+860, 186+352+426+846+858+860, 186+352+1+426+485+860, 186+352+426+485+858+860, 186+352+426+485+860, 186+352+860, 184+186+352+37 6+1132, 186+352+1+333+426+860, 184+186+352+3+107+426, 186+352+5+426, 186+352+5, 186+352+3+426, 186+352+426+846, 186+352+7, 186+352+3+376+426, 186+352+376+426+860+865, 186+352+376+426, 186+352+858, 184+186+352+860, 184+186+352+3+376, 184+186+352+3+639, 186+352 +333+426, 186+352+426, 184+186+352+376, 186+352+426+485, 186+352+333+352+376+426, 186+352+376+426+485+860, 184+186+352+426, 186+352+639, 184+186+352+846, 184+186+352+639, 184+186+352+426+1132, 186+352+1132, 184+186+352+5, 184+186+352+858, 186+352+426+485+860,186+352+333, 184+186+352, 184+186+352+3, 186+352+426+1132, 184+186+352+639+1132, 184+186+352+333, 186+352+988, 184+186+352+1132, 184+186+352+485, 186+352+485, 184+186+352+7, 184+186+352+333+336, 186+352+376, 186+352+3 and 186+352+333+336+352+376+426; the mutation is a mutation to residue R; The Cas12 protein can form a complex with a guide polynucleotide that can specifically bind to a target nucleic acid; the guide polynucleotide contains a guide sequence that is inversely complementary to the target nucleic acid.
2. A fusion protein, characterized in that, The fusion protein comprises: (1) The Cas12 protein as described in claim 1; and (2) Homologous or heterologous functional structural domains; the functional structural domains are nuclear positioning signals.
3. A guiding polynucleotide, characterized in that, It contains (i) sequences that repeat in the same direction. The nucleotide sequence of the same-direction repeat sequence is shown in SEQ ID NO:
704.
4. An isolated nucleic acid, characterized in that, The nucleic acid encodes the Cas12 protein as described in claim 1, or the fusion protein as described in claim 2.
5. The nucleic acid as described in claim 4, characterized in that, The nucleic acid is codon-optimized for expression in eukaryotes.
6. A CRISPR-Cas12 system, characterized in that, The CRISPR-Cas12 system includes: a. The Cas12 protein as claimed in claim 1, the fusion protein as claimed in claim 2, or the nucleic acid encoding the Cas12 protein or the fusion protein; as well as b. A guiding polynucleotide, or a polynucleotide sequence encoding the guiding polynucleotide; The Cas12 protein, the fusion protein, and the guiding polynucleotide form a complex; the guiding polynucleotide contains a guiding sequence that is engineered to guide the complex to bind to the target nucleic acid in a sequence-specific manner.
7. The CRISPR-Cas12 system as described in claim 6, characterized in that, The guiding polynucleotide comprises a unidirectional repeat sequence linked to the guiding sequence, as shown in SEQ ID NO:
704.
8. A carrier system, characterized in that, The vector system includes one or more recombinant vectors, the recombinant vectors including the CRISPR-Cas12 system as described in claim 6.
9. A delivery system, characterized in that, The delivery system includes: (1) Delivery tools, and (2) The Cas12 protein as described in claim 1, the fusion protein as described in claim 2, or the nucleic acid encoding the Cas12 protein or the fusion protein.
10. The delivery system as claimed in claim 9, characterized in that, The delivery tool is a virus, lipid nanoparticle, exosome, microbubble, or gene gun.
11. The delivery system as claimed in claim 9, characterized in that, The delivery tool is a lipid nanoparticle, which comprises: (a) Guiding polynucleotides; and (b) mRNA encoding the Cas12 protein as described in claim 1 or the fusion protein as described in claim 2.
12. A cell, characterized in that, The cell contains the Cas12 protein as described in claim 1, the fusion protein as described in claim 2, or nucleic acid encoding the Cas12 protein or the fusion protein.
13. A reagent kit, characterized in that, The kit contains the Cas12 protein as described in claim 1, the fusion protein as described in claim 2, or a nucleic acid encoding the Cas12 protein or the fusion protein.
Citation Information
Patent Citations
Cas protein and application thereof
CN117683749A
Targeting system and use thereof
WO2026061453A1