Cas12 protein and use thereof
By providing the Cas12 protein, especially the Cas12h isoform, the shortcomings of targeted binding and cleavage of target nucleic acids in eukaryotic cells have been addressed, enabling highly efficient nucleic acid editing capabilities in eukaryotic cells and expanding the application scope of the CRISPR-Cas system.
Patent Information
- Application Number
- PCT/CN2025/096995
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-05-13
- Filing Date
- 2025-05-23
- Publication Date
- 2025-11-27
AI Technical Summary
Existing technologies lack effective Cas12 proteins for targeted nucleic acid editing in eukaryotic cells, particularly in terms of insufficient targeted binding and cleavage capabilities in the nucleus and mitochondria.
A Cas12 protein, including CLUSTER1 to CLUSTER12 proteins, particularly the Cas12h subtype (VH), is provided that can form a complex with a directing polynucleotide to target and cleave target nucleic acids in eukaryotic cells, particularly those located in the nucleus and mitochondria.
This achievement enables efficient targeted binding and cleavage of the Cas12 protein in eukaryotic cells, expanding the application potential of the CRISPR-Cas system in eukaryotic cell editing.
Smart Images

Figure PCTCN2025096995-FTAPPB-I100001 
Figure PCTCN2025096995-FTAPPB-I100002 
Figure PCTCN2025096995-FTAPPB-I100003
Abstract
Description
Cas12 proteins and applications thereof
[0001] This application claims priority to Chinese Patent Application No. 2024106618374, filed on May 24, 2024, and Chinese Patent Application No. 2025106110569, filed on May 13, 2025. This application incorporates the entire contents of the above-mentioned Chinese Patent Applications by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the field of CRISPR gene editing, in particular to a Cas12 protein and applications thereof. BACKGROUND
[0003] CRISPR-Cas system is an adaptive immune defense formed by bacteria and archaea in the long-term evolution process, which can be used to resist invading viruses and exogenous DNA. Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated protein system (CRISPR-Cas system) can directly change the gene sequence in cells, which is a fast and effective method.
[0004] Many researchers in the field are working to find new Cas12 proteins and CRISPR-Cas12 gene editing systems. SUMMARY
[0005] The present disclosure provides Cas12 proteins and applications thereof.
[0006] In one aspect, the present disclosure provides a technical solution: a Cas12 protein, wherein the Cas12 protein is a CLUSTER1 protein, a CLUSTER2 protein, a CLUSTER3 protein, a CLUSTER4 protein, a CLUSTER5 protein, a CLUSTER6 protein, a CLUSTER7 protein, a CLUSTER8 protein, a CLUSTER9 protein, a CLUSTER10 protein, a CLUSTER11 protein, or a CLUSTER12 protein.
[0007] In some embodiments of the present disclosure, the Cas12 protein is a CLUSTER1 protein.
[0008] In another aspect, the present disclosure provides a Cas12 protein, wherein the Cas12 protein belongs to Cas12h subtype (subtype V-H), and the Cas12 protein can be targeted and combined to a target nucleic acid in a eukaryotic cell.
[0009] In some embodiments of the disclosure, a Cas12 protein is provided, which belongs to Cas12h subtype (subtype V-H), which is capable of forming a complex with a guide polynucleotide, which targets binding to a target nucleic acid within a eukaryotic cell. In some embodiments of the disclosure, the Cas12 protein is capable of forming a complex with a guide polynucleotide, which targets binding and cleaving a target nucleic acid within a eukaryotic cell.
[0010] In some embodiments of the disclosure, the target nucleic acid is located within a nucleus of a eukaryotic cell. In some embodiments of the disclosure, the target nucleic acid is located within a mitochondrion of a eukaryotic cell. In some embodiments of the disclosure, the target nucleic acid is located within a chloroplast of a eukaryotic cell.
[0011] Optionally, the eukaryotic cell is a mammalian cell. Further optionally, the eukaryotic cell is a human cell.
[0012] Optionally, the Cas12 protein is a CLUSTER1 protein.
[0013] Optionally, the amino acid sequence of the Cas12 protein comprises an amino acid sequence that is at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to SEQ ID NO: 3 or 18.
[0014] In another aspect, the disclosure provides a Cas12 protein, the amino acid sequence of which comprises or is an amino acid sequence that is at least 50% identical to any one of SEQ ID NOs: 1-35.
[0015] The Cas proteins with sequences of SEQ ID NOs: 1-35 are listed in Table 1, and the respective direct repeat sequences (DRs) of each Cas protein are listed in Table 2.
[0016] In some embodiments of the disclosure, the at least 50% identity is at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity.
[0017] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to any one of SEQ ID NOs: 1-35.
[0018] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 1-35.
[0019] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1-35.
[0020] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 90% identical to any one of SEQ ID NOs: 1-35.
[0021] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 95% identical to any one of SEQ ID NOs: 1-35.
[0022] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 97% identical to any one of SEQ ID NOs: 1-35.
[0023] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 98% identical compared to any one of SEQ ID NOs: 1-35.
[0024] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 99% identical compared to any one of SEQ ID NOs: 1-35.
[0025] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 99.5% identical compared to any one of SEQ ID NOs: 1-35.
[0026] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 99.7% identical compared to any one of SEQ ID NOs: 1-35.
[0027] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 99.8% identical compared to any one of SEQ ID NOs: 1-35.
[0028] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is 100% identical compared to any one of SEQ ID NOs: 1-35.
[0029] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 1.
[0030] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 2.
[0031] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 3.
[0032] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 4.
[0033] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 5.
[0034] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 6.
[0035] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 7.
[0036] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 8.
[0037] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 9.
[0038] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 10.
[0039] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 11.
[0040] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 12.
[0041] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 13.
[0042] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 14.
[0043] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 15.
[0044] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 16.
[0045] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 17.
[0046] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 18.
[0047] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 19.
[0048] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 20.
[0049] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 21.
[0050] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 22.
[0051] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 23.
[0052] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 24.
[0053] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 25.
[0054] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 26.
[0055] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 27.
[0056] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 28.
[0057] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 29.
[0058] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 30.
[0059] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 31.
[0060] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 32.
[0061] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 33.
[0062] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 34.
[0063] In some embodiments of the disclosure, the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 35.
[0064] In some embodiments of the disclosure, the Cas12 protein retains a function of a protein as set forth in any one of SEQ ID NOs: 1-35.
[0065] In some embodiments of the disclosure, the Cas12 protein forms a complex with a guide polynucleotide. In some embodiments of the disclosure, the Cas12 protein specifically binds to a target nucleic acid with a guide polynucleotide.
[0066] In some embodiments of the disclosure, the Cas12 protein forms a complex with a guide polynucleotide, which specifically binds to a target nucleic acid. In some embodiments of the disclosure, the Cas12 protein forms a complex with a guide polynucleotide, which specifically binds to a target DNA.
[0067] In some embodiments of the disclosure, the Cas12 protein specifically binds to a guide polynucleotide and cleaves a target nucleic acid. In some embodiments of the disclosure, the Cas12 protein specifically binds to a guide polynucleotide and cleaves a target DNA. In some embodiments of the disclosure, the Cas12 protein forms a complex with a guide polynucleotide, which specifically binds to and cleaves a target nucleic acid. In some embodiments of the disclosure, the Cas12 protein forms a complex with a guide polynucleotide, which specifically binds to and cleaves a target DNA.
[0068] In the present disclosure, the retention of a function of a protein as set forth in any one of SEQ ID NOs: 1-35 refers to the retention of the ability to form a complex with a guide polynucleotide, the retention of the ability to bind to a target nucleic acid that is complementary to a guide sequence of a guide polynucleotide, the retention of the ability to target cleavage of a target nucleic acid with a guide polynucleotide, and / or the retention of the ability to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.
[0069] In some embodiments of the disclosure, the retaining the function of the protein as set forth in any one of SEQ ID NOs: 1-35 is retaining the ability to form a complex with a guide polynucleotide.
[0070] In some embodiments of the disclosure, the retaining the function of the protein as set forth in any one of SEQ ID NOs: 1-35 is retaining the ability to bind a target nucleic acid complementary to a guide sequence of a guide polynucleotide.
[0071] In some embodiments of the disclosure, the retaining the function of the protein as set forth in any one of SEQ ID NOs: 1-35 is retaining the ability to target cleavage of a target nucleic acid with a guide polynucleotide.
[0072] In some embodiments of the disclosure, the retaining the function of the protein as set forth in any one of SEQ ID NOs: 1-35 is retaining the ability to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.
[0073] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is the amino acid sequence as set forth in any one of SEQ ID NOs: 1-35.
[0074] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally any one or several of the following:
[0075] A, C, T, G,
[0076] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0077] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,
[0078] AAGG, AAGN, NTNN, TCGT, CNTG, NTGG, CCGN, ATAT, TGCA, NGGT, TGNT, NNTG, NCCG, ACAT, GNTG, CGCG, GACN, NTCG, TCNG, CTGC, TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG, CNTC, TGGN, NATC, GTCG, ACNC, GCNN, GACT, CTNT, NCTT, NAGG, NANC, CTTA, GTCT, ANAG, NGCN, CNNA, TCAG, ACAC, NCGG, TNNT, CAAG, ACCT, CCCA, GTNC, ANTC, GACC, AACG, TTAA, TCCG, CGCC, NCCN, TTNA, NCNT, NGCA, AGNN, AATC, GGGA, GNAN, NAGA, CGNA, GTAT, GTNA, ATNC, ACNA, GGAA, NTCC, GGCG, AATN, CNNT, AGGC, GCGN, GTGC, TTGA, AAGC, GAAG, ATNG, TGCT, TACT, CTAN, GGCT, GNGC, GTCN, CGAA, CNAC, GCCT, TAGG, ANGC, TNAA, GANT, NCNA, NCCT, AGAN, GTAA, TTTN, ATGA, TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG, NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNT,CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG, NCNG, AANG, GGTC, ACTC, TGAA, NAGN, NNCA, ACGG, TGAC, TCCN, ANNN, TCGN, TAAN, CAGG, TTAN, NGAN, NTGC, CCNC, TNTN, ATGN, GTGN, GCAT, NNGN, NNCC, CCNA, CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN, GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN, GAAC, NCCA, TTCC, NCNN, GNNN, ANGT, NTNA, CCCT, GNAA, TTNG, GTNN, GGNG, TCTA, NCAC, GANG, TTCG, CCTC, CNGG, ANNA, TCAN, ATCG, NTGA, CGTA, TTAC, GCTN, GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN,GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NNTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG, ACGN, GGCA, AAAG, TTGT, NGNA, NAAN, TATN, CGGG, CATA, ATGC, ACGC, ACCN, ATTT, TCNA, TNGC, NACA, NACC, CTCN, GGCC, TANG, AGAA, TNGA, TAGC, CAGN, GGCN, ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN, GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA, NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC, ATNN, NNNT, AGAC, AACC, ANNC, ANNG, ACAN, GTTC, TATA, GNTC, NCGT, NGNT, CGTC, CCGC, CGAC, GACG, ATTG, GNNC, CNAA, TATC, AGNA, CTNC, TTCA, ANCA, ACCC, AGTG, CCGT, ANAT, CTGT, GGGC, NTTA, NAAG, AANN, CNAN, NNCN, ANAA, ANAN, CTTG, NGNN, AGAG, TANA, TCNC, GCAA, NGNG, NAGC, NATA, ATCN, CGTT, CNGC, GATN, NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TT TG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;
[0079] The N is A, T, C or G.
[0080] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or more of the following: WYR, BMCTTH, TTN, VNWTV, VNWTC, VNTTC;
[0081] wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, and V is A, C or G.
[0082] In some embodiments of the disclosure, the PAM sequence recognized by the Cas12 protein is WYR; wherein W is A or T, Y is C or T, and R is A or G.
[0083] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the conjugate is optionally selected from any one, two or more degenerate sequences or non-degenerate sequences (non-degenerate sequences refer to any specific sequence encompassed by the degenerate sequences) of the following: WYR, BMCTTH, TTN, VNWTV, VNWTC, VNTTC;
[0084] wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, and V is A, C or G.
[0085] In some embodiments of the disclosure, the PAM sequence recognized by the conjugate is WYR; wherein W is A or T, Y is C or T, and R is A or G.
[0086] In some embodiments of the disclosure, the PAM sequence is immediately adjacent to the target sequence on the target nucleic acid, i.e., the PAM sequence is directly connected to the target sequence on the target nucleic acid by a covalent bond, without nucleotides.
[0087] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the sequences shown in Figure 3. In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the sequences shown in Figure 4. In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the sequences shown in Figure 5. In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the sequences shown in Figure 6. In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the sequences shown in Figure 7.
[0088] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence shown in SEQ ID NO: 18;
[0089] The Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves a target nucleic acid, or, the complex sequence-specifically binds a target nucleic acid, but does not cleave the target nucleic acid;
[0090] Optionally, the guide polynucleotide comprises a guide sequence and a direct repeat sequence; further optionally, the guide sequence is reverse-complementary to the target nucleic acid, the backbone sequence interacts with the Cas12 protein;
[0091] Optionally, the Cas12 protein recognizes a PAM of 5'-WYR-3', wherein W is A or T, Y is C or T, and R is A or G; further optionally, the Cas12 protein recognizes a PAM of 5'-ACA-3', 5'-TCA-3', 5'-ATA-3', 5'-TTA-3', 5'-ACG-3', 5'-TCG-3', 5'-ATG-3', and / or 5'-TTG-3';
[0092] Optionally, the direct repeat sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in any one of SEQ ID NOs: 84-86, 187-195; further optionally, the direct repeat sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 84.
[0093] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the sequence set forth in SEQ ID NO: 18;
[0094] Optionally, the Cas12 protein is non-natural, or, engineered;
[0095] Optionally, the Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves a target nucleic acid, or, the complex sequence-specifically binds a target nucleic acid, but does not cleave the target nucleic acid; optionally, the complex is non-natural, or, engineered;
[0096] Optionally, the guide polynucleotide comprises a guide sequence and a scaffold sequence; optionally, the guide sequence is reverse-complementary to a target nucleic acid, the scaffold sequence interacts with the Cas12 protein; optionally, the scaffold sequence is a direct repeat sequence; optionally, the guide sequence is located at the 5’ end or the 3’ end of the scaffold sequence; optionally, the guide polynucleotide is non-natural, or, engineered;
[0097] Optionally, the Cas12 protein recognition sequence is a PAM of 5'-WYR-3', wherein W = A or T, Y = C or T, and R = A or G; further optionally, the Cas12 protein recognition sequence is a PAM of 5'-ACA-3', 5'-TCA-3', 5'-ATA-3', 5'-TTA-3', 5'-ACG-3', 5'-TCG-3', 5'-ATG-3', 5'-TTG-3', and / or 5'-TTN-3';
[0098] Optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in any one of SEQ ID NOs: 84-86, 187-195; further optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 84;
[0099] Optionally, the Cas12 protein corresponds to, as in SEQ ID, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 19, 21, 22, 23, 24 of the sequence shown in NO:18 ,25,26,27,28,29,30,32,33,34,35,36,37,38,39,40,41,42,43,44,46, 47, 48, 49, 50, 51, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 64, 65, 66, 67, 68, 6 9, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 88, 89, 90 91, 92, 93, 94, 95, 96, 97, 98, 99, 101, 102, 103, 104, 105, 106, 108, 109, 110, 111, 112, 114, 115, 116, 117, 118, 119, 120, 121, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 151, 152, 153, 154, 155, 156, 157, 158 159, 160, 161, 162, 163, 164, 165, 166, 167, 169, 170, 171, 172, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 194, 195, 196, 197, 198, 199, 200, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224 225, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 242, 243, 244, 245, 247, 248, 249, 250, 251, 252, 253, 255, 256, 257, 258, 259, 260, 261, 262, 263, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 278, 279, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292,293、294、295、296、297、298、299、300、301、302、303、305、306、308、309、310、313、315、316、317、318、319、320、321、322、323、324、325、327、328、329、330、331、332、333、334、335、336、337、339、340、341、342、343、344、345、346、347、348、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、431、432、433、435、436、437、439、440、441、442、443、444、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、467、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、496、497、499、500、501、502、503、504、506、507、508、509、510、511、512、513、514、515、516、517、518、519、520、521、522、523、524、525、526、527、528、529、531、532、533、534、535、536、537、538、539、540、541、542、543、544、545、546、547、548、549、550、552、553、555、556、557、558、559、560、561、562、563、564、565、566、567、568、569、570、571、572、573、574、575、576、577、578、579、580、581、582、583、584、585、586、587、589、590、592、593、594、595、596、597、598、599、601、602、603、604、605、606、607、608、609、610、611、612、613、614、615、616、618、619、620、621、622、623、624、625、626、627、628、630、631、632、633、634、635、636、637、638、639、640、641、642、643、644、645、646、647、648、649、650、651、652、653、654、655、656、657、658、659、660、661、662、663、664、665、666、667、668、669、670、671、672、673、674、675、676、678、679、680、681、683、684、685、686、688、689、691、692、693、694、695、696、697、698、699、700、701、702、703、704、705、706、707、708、709、710、711、712、713、715、716、717、719、720、721、722、723、724、725、727、728、729、730、731、732、733、734、736、737、738、739、740、741、742、743、744、745、746、747、748、749、751、752、753、754、755、756、758、759、760、761、762、764、765、766、767、768、769、771、772、773、774、775、776、779、780、781、782、783、784、785、786、787、789、790、791、792、794、795、797、798、800、801、802、804、805、806、807、808、809、810、811、812、813、814、815、817、818、819、821、822、823、824、825、826、827、828、829、830、831、832、833、834、835、836、837、838、839、840, 841, 842, 844, 845, 846, 847, 848, 849, 850, 851, 852, 853, 854, 855, 856, 857, 858, 859, 860, 862, 863, 864, 865, 866, 867, 868, 870, 872, 873, 874, 875, 876, 877, 879, 880, 881, 882, 883, 884, 885, 886, 887, 888, 890, and 891 at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, or at least 12 of the amino acid residues corresponding to any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, or more of the following positions of SEQ ID NO: 18: W8, D9, I10, Q11, R12, C13, Q14, K15, L16, K17, L18, G19, K20, K21, Y38, F42, T54, E62, V93, A94, E95, M96, P97, Q98, A99, S100, A101, S102, S103, F104, Y105, G106, Y109, N111, Y112, S113, C114, N115, D116, K117, A118, K119, W120, T121, Q122, A123, K124, S125, F127, K142, G145, D146, S147, C148, L149, Q151, K171, W174, E175, S178, L181, A182, N183, K184, V185, N186, S187, Y189, R206, E207, S210, E214, R217, L218, Q219, V220, K221, S222, C223, Y224, Q225, K226, N227, L228, D229, H230, V233, T234, L237, S259, L262, Y263, I265, G266, T267, G268, L269, S270, K271, N272, V273, L274, R276, C280, T285, L286, A287, S288, N289, P290, T291, Y292, K293, I294, I296, Y319, K323, D326, Q327, L328, K329, R330, R331, K332, V333, Y334,P335, R336, L337, P338, S339, F340, K341, N342, D343, Y344, K345, M347, F348, L350, S351, S352, L353, K355, L376, F377, M378, N379, S380, H381, Y382, F383, N394, K395, T396, A397, K398, Q399, F404, R405, H406, K407, L408, K409, S410, A415, V416, S417, D418, I420, Y423, V424, K425, Q426, I427, G428, Q430, K431, K432, N433, G434, S435, F436, Y437, V438, T439, L440, M441, F442, T443, M444, E448, E454, R455, F456, F457, K458, T459, A460, S461, P462, D463, K466, Y467, D480, L481, N482, I483, S484, N485, P486, D522, N527, K530, R531, K533, Q534, L535, F537, K538, K540, D541, I543, K544, D545, C546, K547, F548, S549, N550, S551, N552, M556, N557, D558, A559, T560, I561, S562, F563, L564, R566, S569, P570, S571, Q572, S573, P574, R575, C576, M577, I578, Q579, T580, W581, I582, K583, N584, L585, K586, K587, L589, K590, K591, L592, H593, S594, I595, I596, R597, A598, S599, G600, Y601, V602, L608, R609, M610, L611, E612, Q614, D615, A616, M617, K618, S619, L620, I621, S622, S623, Y624, E625, R626, F627, H628, L629, K630, S631, G632, E633, M634, L635, A636, A637, K638, K639, N640, I641, T642, A643, N644, N645, R646, R647, Q648, N649, F650, R651, Q652, F653, I654, S655, R656, K657, I658, A659,S660, K661, I662, V663, Q664, Y665, S666, K667, G668, E675, D676, L677, S678, L679, D680, F681, D682, S683, D684, N685, K686, N687, N688, S689, L690, I691, R692, L693, F694, S695, A696, D697, G698, L699, K701, C702, I703, T704, D705, A706, A707, Y708, K709, A710, G711, I712, L716, P719, M720, G721, T722, S723, K724, R735, N736, L737, K738, N739, K740, N741, A756, D757, A760, H771, S772, I773, Y776, K777, F778, Y779, V780, K781, G782, K784, E794, K795, E796, V797, G798, K799, R800, L801, Q802, R803, F805, E838, N839, A840, F841, Y843, T851, A852, D853, N854, H855, R856; optionally, the mutation is to any other natural amino acid residue; further optionally, the mutation is to residue R, H, K, or A; further optionally, the mutation is to residue R; in some embodiments of the disclosure, the mutation is to residue A;
[0100] optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, or more amino acid residues corresponding to positions N5, D9, E58, S100, N115, K142, C148, S147, K232, S245, I251, Y263, D279, A297, L300, E303, L337, M378, N394, T396, T443, K458, T468, K533, F537, F548, N550, D697, A706, I788 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is to residue R;
[0101] Optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to positions D480, E675, D757 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0102] Some positions that maintain or increase editing activity after mutation (e.g., editing efficiency after mutation is at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 110%, at least 120%, at least 130%, or at least 140% of the wild-type protein of SEQ ID NO: 18) and some positions that can significantly reduce editing activity after mutation (e.g., editing efficiency after mutation is reduced by at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% compared to the wild-type protein) have been listed in Tables 5 and 6 of the disclosure. These positions can be the mutation positions of the Cas12 protein. In some embodiments, the Cas12 protein of the disclosure is obtained by mutating to other amino acid residues at these positions; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue R, H, K, or A; optionally, the mutation is a mutation to residue R; optionally, the mutation is a mutation to residue A.
[0103] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 19;
[0104] The Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves the target nucleic acid, or, the complex sequence-specifically binds the target nucleic acid but does not cleave the target nucleic acid;
[0105] Optionally, the guide polynucleotide comprises a guide sequence and a direct repeat sequence; further optionally, the guide sequence is reverse-complementary to the target nucleic acid, the backbone sequence interacts with the Cas12 protein;
[0106] Optionally, the Cas12 protein recognition sequence is a PAM of 5'-BMCTTH-3';
[0107] Optionally, the direct repeat sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence set forth in any one of SEQ ID NOs: 87-89;
[0108] Optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16 or more amino acid residues corresponding to a position of SEQ ID NO: 19: I1, V2, K3, P4, K5, S6, I7, K8, S9, Y10, S11, S12, M13, L14, D15, V16, D17, H18, R19, K20, N21, T62, D64, L82, P84, N120, F121, D122, K125, Y126, E159, G160, Y161, G163, L164, K165, C166, G167, K168, T169, W170, G171, T172, I173, S174, G175, L176, F177, G178, T179, G180, E181, K182, A183, D184, R185, K188, L192, R207, E224, K227, L228, Y230, G231, N232, I233, G234, R235, A236, S237, F238, V239, I240, V241, R242, E244, D253, K255, Y256, Q259, I260, K263, A266, D267, K270, Q271, D274, L275, Y294, Y295, Q296, P297, S300, E301, S303, N304, N305, L307, P308, I309, I310, Q311, G312, K313, T314, T315, K316, N317, Y318, N319, F320, Q324, Y345, F346, K349, F350, F351, T352, A353, D354, N355, V356, F357, S358, I359, C360, F362, H363, D383, E386, E387, T388, V389, S390, A391, C393, H394, I396, N397, E398, N399, G400, R401, M402, P403, I404, Y405, S406, L407, M409, E431, S434, K435, I436, E437, R438, Q439, K440, L441, N442, P443, I444, V445, E446, G447, K448, A449, S450, F451, N452, W453, G454, N455, V456, S457, K458, I459, S460, G461,C462, I463, I464, S465, K467, E468, K469, E470, K471, H472, I473, V474, S476, K477, H478, N479, H480, D481, S482, S483, I484, W485, I486, E487, T489, W497, K499, H500, H501, F502, R503, M504, F505, N506, T507, R508, F509, Y510, E511, E512, Y514, I529, S531, R532, R533, F534, F536, N537, N538, Q539, V540, V541, L542, S543, E544, D545, Q546, I547, N548, T549, I550, R551, N552, A553, S554, K555, S556, M557, R558, K559, A560, M561, K562, R563, Q564, V565, R566, D583, D584, F585, N586, I587, N588, I589, S590, N591, D592, R594, R597, T598, T599, L600, S601, Y602, K603, I604, E605, R608, V609, E610, T611, F615, D619, Q620, N621, Q622, T623, A624, R625, S659, S660, Q661, L662, V663, N664, D665, K666, S667, F668, D669, Q670, L671, Y673, D674, G675, I676, S677, W678, D679, R680, F681, Q682, S683, W684, C698, V699, S700, K701, N702, R703, K704, A705, Q706, D707, V708, P709, I710, D711, E714, I715, R718, S719, S720, K721, Y722, P724, L726, Y727, D728, R732, C734, G735, I737, K738, K739, I740, M741, K742, G743, K744, Q764, F765, S766, V767, L768, R769, L770, S771, S772, L773, N774, H775, N776, S777, F778, M780, L781, R782, N783, K785, G786, I787, I788, S789, A790, Y791, F792, N793, N794, L795,I796, G797, K798, H799, C800, T801, D802, E803, Q804, K805, F813, R816, I817, E818, L819, E820, E821, K822, R823, Q824, N825, K826, A827, I828, S829, K830, K831, N832, L833, I834, S835, N836, R837, V839, T840, V854, V856, G857, E858, N859, I860, S861, N862, T863, T864, S865, K866, S867, N868, K869, S870, K871, Q872, N873, A874, R875, A876, M877, D878, W879, L880, S881, R882, G883, V884, A885, D886, K887, Q890, M891, T892, E893, M894, H895, R899, F900, R901, D902, I903, N904, P905, A906, Y907, T908, S909, H910, Q911, H916, R917, K921, V924, M926, A928, R929, K931, E937, T939, E940, V941, D942, Y945, Y957, Y958, R1004, S1005, G1006, G1007, R1008, S1034, D1035, N1041, I1042, A1043, L1044, V1045, G1046, I1047, E1048, F1049, E1050;,
[0109] optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D619, E858, D1035 of the sequence set forth in SEQ ID NO: 19; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0110] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 20;
[0111] the Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; alternatively, the complex sequence-specifically binds and cleaves a target nucleic acid, or the complex sequence-specifically binds a target nucleic acid but does not cleave the target nucleic acid;
[0112] alternatively, the guide polynucleotide comprises a guide sequence and a direct repeat sequence; further alternatively, the guide sequence is reverse complementary to the target nucleic acid, and the backbone sequence interacts with the Cas12 protein;
[0113] alternatively, the Cas12 protein recognizes a PAM sequence of 5'-TTN-3';
[0114] alternatively, the direct repeat sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 90-91;
[0115] Optionally, the Cas12 protein has a mutation at an amino acid residue corresponding to any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, or more of the following positions of SEQ ID NO: 20: M1, K2, T3, L4, I5, R6, K7, T8, Y9, V10, M11, L12, V13, K19, Y30, L93, C98, K99, T100, G101, M102, K103, S104, E105, K106, D107, L108, E109, Q110, K111, L112, R113, K114, L115, D117, E118, F136, K140, I141, V143, S144, S145, L146, K147, S148, W149, D150, D151, R152, N153, V155, T156, E166, N170, A179, L180, W183, S186, N187, K188, L189, F190, L191, T192, K193, K194, V195, A196, S197, K198, F199, K200, K201, F202, G203, W204, D205, T206, Y211, V219, N220, S221, D222, A223, S224, Y225, W226, K228, M229, F230, W232, Q233, K236, R243, P244, T245, S246, L247, C248, T249, L250, P251, E252, L253, A254, V255, S256, E257, R258, E259, I260, P261, Y262, G263, V264, R277, A289, L290, R291, T292, L293, Y294, F295, K296, K304, N305, S306, Y307, Y311, R313, G314, N315, N316, A325, V326, L327, K328, E329, I330, T331, Y333, K335, N336, G337, N338, Y339, Y340, V341, G342, L343, S344, L345, N346, L347, Q348, K354, R357, T358, V359, K360, D361, Y362, Y363, F364, F365, K366, D380, L381, G382, I383, T384, N385, P386, V421, K425, K428, A429,S432, F435, A436, I438, T439, E440, L441, H442, P443, I444, K445, K446, Q451, E452, E453, W454, S455, K456, L457, R458, Y459, P460, I461, S462, Q463, M464, I465, E466, K467, L468, S469, K470, E471, M472, R473, Q474, L475, R476, R477, G478, D479, L480, N481, R483, N484, H485, G486, T487, H489, Q491, M492, Q493, F494, L496, Y498, K499, F501, V502, D503, L504, L505, K506, K507, W508, T509, Y510, F511, G512, S513, K514, P515, K518, K519, R521, R522, K523, G524, F525, E526, K527, H528, I529, R530, R531, L532, E533, N534, L535, K536, K537, D538, F539, R540, K541, K542, L543, A544, C545, E546, V548, R549, E562, D563, L564, E565, H566, F567, T568, P569, D570, S571, T572, K573, D574, S575, N576, L577, N578, E579, L580, L581, M582, L583, W584, G585, S586, G587, Q588, I589, G590, K591, W592, E594, H595, F596, Q599, Y600, K606, V607, D608, P609, R610, M611, T612, S613, Q614, I615, R625, S626, K627, Y628, D629, K630, F633, A646, D647, A650, N653, I654, R657, R661, P664, F665, K667, D682, D683, N684, S685, R686, R687, R688, H689, E719, V720, Y721, Y723, G732, K735, Y736, R751;
[0116] Optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D380, E562, D647 of the sequence set forth as SEQ ID NO: 20; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0117] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth as SEQ ID NO: 22;
[0118] The Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves a target nucleic acid, or, the complex sequence-specifically binds a target nucleic acid, but does not cleave the target nucleic acid;
[0119] Optionally, the guide polynucleotide comprises a guide sequence and a direct repeat sequence; further optionally, the guide sequence is reverse complementary to the target nucleic acid, the backbone sequence interacts with the Cas12 protein;
[0120] Optionally, the Cas12 protein recognizes a PAM of 5'-VNWTV-3', 5'-VNWTC-3', or 5'-VNTTC-3';
[0121] Optionally, the direct repeat sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence set forth as any one of SEQ ID NOs: 101-114;
[0122] Optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16 or more amino acid residues corresponding to a position of SEQ ID NO: 22: M1, A2, S3, K4, H5, V6, V7, R8, P9, F10, N11, S12, V13, C14, T15, A16, K17, G18, D19, R20, L21, R22, Y23, E35, E62, L74, P76, G110, F111, D112, K115, Y116, N154, C155, D156, A157, G158, A159, G160, S161, N162, N163, A164, V165, S166, M167, L168, F169, G170, D171, G172, P173, K174, S175, D176, Y177, K180, Y222, G223, K224, T225, G226, S227, P228, S229, A230, M231, A232, R233, F234, S250, K251, K253, K254, F255, D258, K261, Q262, K265, L267, R275, E287, F288, Y289, A290, R291, A292, S294, A295, A298, N299, A301, S302, E303, I304, N305, A306, K307, F308, T309, H310, N311, C312, T313, F314, D317, Y345, M346, T348, V349, A350, E351, D352, C353, R354, Y355, V356, L357, A358, Y360, H361, E380, F383, N384, W387, E388, L391, I394, D395, F396, N397, Q398, K399, P400, P401, V402, R403, E404, L405, K407, S429, D432, R433, I434, D435, N436, V437, Y438, P439, H440, P441, F442, V443, Q444, G445, K446, Q447, G448, Y449, T450, F451, G452, P453, S454, N455, I456, E457, A459, N461, D462, M465, Q466, I467, K468, S469, I472, A473, E475, R476,P477, M478, M479, W480, V481, T482, T483, K484, D487, W491, I492, N493, H494, H495, L496, P497, F498, A499, N500, S501, R502, Y503, Y504, E505, E506, Y508, D522, G523, K524, F527, V528, L529, G530, K531, T532, I533, D534, A535, F536, A537, T538, G539, R540, I541, K542, T543, S544, V545, G546, R547, Q548, K549, A550, A551, K552, A553, I554, E555, R556, K558, D568, K570, T571, T572, F573, C574, R576, R577, K578, R581, V583, I584, A585, I586, N587, H588, R589, H590, D609, Q610, N611, E612, G613, A614, P615, S647, I648, Q649, S650, G651, K652, D653, V654, F655, Y657, S658, G659, V660, H661, D664, K665, A666, N667, G668, F669, D670, V671, L672, T685, E687, D688, A690, Y691, R694, S695, E697, W698, C699, L702, Y703, L708, R711, G714, K715, L716, I717, R718, K719, S739, P740, L741, S742, P743, V744, R745, L746, H747, S748, L749, S750, K752, S753, L754, E755, T757, K758, K759, I761, S762, C763, I764, S765, S766, Y767, F768, S769, V770, C771, N772, M773, K774, T775, V776, E777, E778, K779, Y787, W790, N791, K792, Y794, A795, S796, L797, V798, E799, R800, R801, K802, E803, R804, V805, K806, L807, S808, A809, G810, L811, I812, I813, R814, E827, G828, D829, L830, P831, T832, V833, A834, S835,G836, K837, S838, R839, Q840, N841, N842, S843, G844, K845, Q846, D847, W848, C849, A850, R851, E852, L853, K855, R856, E859, M860, A861, V863, V869, P870, V871, F872, P873, Q874, W875, T876, S877, H878, R894, S904, R906, D907, L909, A910, N913, T920, G921, T922, A923, Y925, Y926, M965, R966, G967, G968, R969, A996, D997, A1000, V1007;,
[0123] Optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the following positions corresponding to the sequence of SEQ ID NO: 22: D609, E827, D997; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0124] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence of SEQ ID NO: 23;
[0125] The Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves the target nucleic acid, or, the complex sequence-specifically binds the target nucleic acid, but does not cleave the target nucleic acid;
[0126] Optionally, the guide polynucleotide comprises a guide sequence and a direct repeat sequence; further optionally, the guide sequence is reverse-complementary to the target nucleic acid, the backbone sequence interacts with the Cas12 protein;
[0127] Optionally, the Cas12 protein recognizes a PAM of 5'-TTN-3';
[0128] Optionally, the direct repeat sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in any one of SEQ ID NOs: 115-116;
[0129] Optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16 or more amino acid residues corresponding to a position of SEQ ID NO: 23: M1, N2, N3, K4, N5, V6, K7, S8, Y9, N10, C11, Q12, I13, L14, T15, N16, R18, K19, F22, K60, A67, E68, E69, K70, N71, T72, K73, A74, S75, K76, K77, T78, N79, K80, I101, P103, N138, F139, N140, S141, E142, K143, Y144, K178, G181, L182, K183, F184, G185, E186, I187, W188, G189, I190, V191, S192, N193, L194, F195, G196, T197, G198, D199, K200, V201, P202, K203, K206, E225, L228, Q242, Y245, L246, Y248, F249, I250, S251, G252, R253, K254, P255, S256, E257, Y258, F259, Y260, K263, K269, I270, D271, K274, V275, K278, K281, N282, K285, Y286, L290, F309, N310, Q311, K312, S315, E316, F318, N319, A320, W322, P323, I324, I325, Q326, S327, K328, T329, T330, R331, N332, L333, N334, F335, E338, Q339, F361, S364, Y365, F366, K367, T368, D369, N370, K371, F372, I373, I374, K375, K377, H378, E382, E400, K401, E404, S407, I410, E411, D412, N413, S414, S415, K416, P417, D418, L419, M420, K423, Q445, Q446, F448, K449, I450, E451, N452, R453, F454, L455, N456, P457, I458, V459, D460, N461, S462, Y463, S464, Y465, N466, W467, G468, D469, K470,S471, K472, L473, N474, C476, I477, I478, S479, K482, K483, S484, K485, F486, N487, L488, K489, N490, N491, R492, P493, D494, Y495, D496, Y497, G498, I499, W500, M501, E502, L503, E504, W511, K513, H514, H515, F516, L517, V518, S519, N520, T521, R522, F523, M524, E525, E526, Y528, F545, T547, K548, R549, N550, F552, D553, N554, N555, V556, V557, L558, S559, D560, Q561, Q562, I563, Q564, N565, I566, R567, N568, A569, P570, K571, H572, R573, R574, R575, A576, I577, K578, R579, Q580, M581, R582, N599, D600, Y601, N602, I603, N604, I605, S606, K607, S608, N610, R613, A614, I615, I616, S617, K618, K619, F620, E621, I622, E623, I624, C625, K626, V628, D635, Q636, N637, Q638, S639, A640, N641, S675, K676, Q677, A678, V679, G680, K681, N682, E683, N684, K685, R686, E687, F688, D689, Q690, L691, S692, Y693, N694, G695, I696, K697, W698, G699, E700, F701, N702, D703, N717, V718, F719, K720, V721, N722, K723, F724, G725, V726, K727, S728, N729, V730, L732, L739, N742, N743, P744, V745, L746, Y747, Y748, M751, K752, N755, K758, N759, I760, L761, Y762, K763, K764, K784, F785, S786, V787, M788, K789, L790, S791, S792, L793, S794, G795, L796, S797, F798, S799, M800, I801, R802, S803, A804, K805,S806, L807, I808, S809, S810, Y811, F812, G813, N814, L815, L816, E817, G818, T819, T820, T821, D822, D823, Q824, K825, F833, R836, Q837, K838, E840, K841, K842, R843, K844, D845, K846, Q847, K848, S849, K850, K851, E852, L853, T854, A855, N856, K857, V859, S860, E877, D878, I879, G880, N881, M882, T883, S884, N885, S886, N887, K888, N889, S890, V891, N892, S893, A894, S895, M896, D897, W898, L899, A900, R901, G902, V903, A904, N905, K906, K908, Q909, L910, M913, H914, L918, Y919, Y920, S921, I922, N923, P924, F925, M926, T927, S928, H929, Q930, H935, N936, R940, F942, K943, A944, R945, Y953, L954, F955, E956, K957, D958, T972, R973, Q974, T975, T976, Y979, K1025, M1026, G1027, G1028, R1029, A1055, D1056, A1059, A1066, K1067, G1069, K1070, N1071, E1072, T1073, S1074, S1075, D1076;
[0130] Optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D635, E877, D1056 of the sequence set forth in SEQ ID NO: 23; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0131] In some embodiments of the disclosure, the Cas12 protein is a Cas12 inactivated variant. In some embodiments of the disclosure, the Cas12 protein is a variant inactivated for nuclease activity. In some embodiments of the disclosure, the Cas12 protein is a dead Cas12 inactivated variant or a nickase Cas12 inactivated variant. Optionally, the Ruvc domain of the Cas12 protein is inactivated.
[0132] In some embodiments of the disclosure, the Cas12 protein is selected from an active fragment of any Cas12 protein described in the disclosure.
[0133] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identical to any one of SEQ ID NOs: 18-20, 22, and 23;
[0134] Optionally, the Cas12 protein forms a complex with a guide polynucleotide; further, the complex specifically binds to a target nucleic acid; further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0135] Optionally, the Cas12 protein forms a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse-complementary to a target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that interacts with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the direct repeat sequence comprises or is a sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NOs: 84-91, 101-116, and 187-195;
[0136] Optionally, the backbone sequence does not comprise a tracrRNA sequence;
[0137] Optionally, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the following: WYR, BMCTTH, TTN, VNWTV, VNWTC, VNTTC, wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, V is A, C or G.
[0138] Optionally, the reverse complement is partially or fully complementary. In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid.
[0139] In some embodiments of the disclosure, the Cas12 protein is a mutant of a Cas protein set forth in any one of SEQ ID NOs: 1-35.
[0140] In some embodiments of the disclosure, the Cas12 protein is an inactivated variant of a Cas protein set forth in any one of SEQ ID NOs: 1-35.
[0141] In some embodiments of the disclosure, the Cas12 protein provided herein comprises one, two, or more mutations, e.g., a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof, compared to the Cas12 protein set forth in any of SEQ ID NOs: 1-35. In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid changes (e.g., insertions, deletions, or substitutions) compared to the Cas12 protein set forth in any of SEQ ID NOs: 1-35, but retains the ability to bind a target nucleic acid molecule that is complementary to a guide sequence of a guide polynucleotide, and / or retains the ability to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid changes (e.g., insertions, deletions, or substitutions) compared to the Cas12 protein set forth in any of SEQ ID NOs: 1-35, but retains the ability to bind a target nucleic acid molecule complementary to a guide sequence of a guide polynucleotide.
[0142] In another aspect, the present disclosure provides a technical solution: a guide polynucleotide comprising (i) a direct repeat sequence having at least 50% identity compared to any one of SEQ ID NOs: 36-170, 187-195, and (ii) a guide sequence engineered to hybridize to a target nucleic acid; the direct repeat sequence is linked to the guide sequence, the guide polynucleotide is capable of forming a complex with a Cas12 protein and directing the complex to bind to a sequence of the target nucleic acid in a sequence-specific manner.
[0143] In some embodiments of the present disclosure, the direct repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to any one of SEQ ID NOs: 36-170, 187-195.
[0144] In some embodiments of the disclosure, the direct repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 84.
[0145] In some embodiments of the disclosure, the direct repeat sequence has at least 60% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0146] In some embodiments of the disclosure, the direct repeat sequence has at least 65% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0147] In some embodiments of the disclosure, the direct repeat sequence has at least 70% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0148] In some embodiments of the disclosure, the direct repeat sequence has at least 75% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0149] In some embodiments of the disclosure, the direct repeat sequence has at least 80% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0150] In some embodiments of the disclosure, the direct repeat sequence has at least 85% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0151] In some embodiments of the disclosure, the direct repeat sequence has at least 90% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0152] In some embodiments of the disclosure, the direct repeat sequence has at least 95% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0153] In some embodiments of the disclosure, the direct repeat sequence has at least 96% sequence identity to any one of SEQ ID NOs: 36-170, 187-195.
[0154] In some embodiments of the disclosure, the direct repeat sequence has at least 97% sequence identity compared to any one of SEQ ID NOs: 36-170, 187-195.
[0155] In some embodiments of the disclosure, the direct repeat sequence has at least 98% sequence identity compared to any one of SEQ ID NOs: 36-170, 187-195.
[0156] In some embodiments of the disclosure, the direct repeat sequence has 100% sequence identity compared to any one of SEQ ID NOs: 36-170, 187-195.
[0157] In preferred embodiments, the Cas12 protein is a Cas12 protein described in the present disclosure.
[0158] In some embodiments of the disclosure, the guide sequence comprises 15-60 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-50 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-40 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-35 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-30 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-25 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 18-25 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-25 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 18-22 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-22 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides.
[0159] In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid, the guide sequence is 90-100% complementary to the target nucleic acid.
[0160] In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid.
[0161] In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid, the guide sequence mismatches the target nucleic acid by no more than one nucleotide.
[0162] In some embodiments of the disclosure, the direct repeat sequence comprises 15-100 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-90 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-80 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-70 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-60 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-50 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-40 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-40 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-30 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0163] In some embodiments of the disclosure, the guide sequence is located at the 3' end of the direct repeat sequence.
[0164] In some embodiments of the disclosure, the guide sequence is located at the 5' end of the direct repeat sequence.
[0165] In some embodiments of the disclosure, the guide polynucleotide further comprises a tracrRNA.
[0166] In some embodiments of the disclosure, the tracrRNA is complementary paired to the direct repeat sequence. In general, the complementary pairing is partial base complementary pairing. In some embodiments of the disclosure, the tracrRNA interacts with the direct repeat sequence.
[0167] In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 1-10 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 4 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a 5'-GAAA-3' sequence.
[0168] In some embodiments of the disclosure, the tracrRNA sequence is located at the 3' end of the direct repeat sequence.
[0169] In some embodiments of the disclosure, the tracrRNA sequence is located at the 5' end of the direct repeat sequence.
[0170] In some embodiments of the disclosure, the tracrRNA comprises 10-200 nucleotides. In some embodiments of the disclosure, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In some embodiments of the disclosure, the tracrRNA comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0171] Table 1 shows the amino acid sequences of Cas proteins.
[0172] Table 2 shows the direct repeat sequences (DRs) corresponding to the Cas proteins. When a certain Cas protein corresponds to multiple DR sequences, one of them can be optionally used.
[0173] In another aspect, the disclosure provides a technical solution: a Cas12 inactivated variant, characterized in that the Cas12 inactivated variant is a nuclease activity inactivated variant of the Cas12 protein as described in the disclosure.
[0174] In the context of the present disclosure, the reference to the Cas12 protein can encompass the Cas12 inactivation variant. However, given the importance of the Cas12 inactivation variant (non-limiting examples include the Cas12 inactivation variant fused to a deaminase for single base editing, fused to a transcriptional activation domain or a transcriptional repression domain for transcriptional regulation, etc.), the present disclosure will focus on the description of the Cas12 inactivation variant separately; this does not mean that the reference to the Cas12 protein does not include the Cas12 inactivation variant.
[0175] In some embodiments of the present disclosure, the Cas12 inactivation variant is optionally selected from the Cas12 proteins described in the present disclosure.
[0176] In some embodiments of the present disclosure, the Cas12 inactivation variant is a variant completely inactivated for nuclease activity, i.e., a dead Cas12 inactivation variant (dCas12). The dCas12 can only bind to a target nucleic acid under the mediation of a guide polynucleotide, and has no or almost no function of cutting the target nucleic acid. For example, the target nucleic acid cutting efficiency of the dCas12 is ≤20%, ≤15%, ≤10%, ≤5%, ≤4%, ≤3%, ≤2%, or ≤1% of the target nucleic acid cutting efficiency of the Cas12 protein before the inactivation mutation.
[0177] In some embodiments of the present disclosure, the Cas12 inactivation variant is a variant partially inactivated for nuclease activity. Further, the variant partially inactivated for nuclease activity is a Cas12 nickase (nickase Cas12, nCas12) which binds to a target nucleic acid under the mediation of a guide polynucleotide and then cuts one of the single strands in the double-stranded target nucleic acid, without cutting the other single strand.
[0178] In some embodiments of the present disclosure, the Cas12 inactivation variant is inactivated for the Ruvc domain of the Cas12 protein.
[0179] In some embodiments of the present disclosure, the Cas12 inactivation variant is inactivated for the Ruvc-I, Ruvc-II, or Ruvc-III domain of the Cas12 protein.
[0180] In some embodiments of the present disclosure, the Cas12 inactivation variant is obtained by introducing an inactivation mutation in the Ruvc-I, Ruvc-II, or Ruvc-III domain of the Cas12 protein.
[0181] In some embodiments of the disclosure, the amino acid sequence of the Cas12 inactivated variant comprises or is an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identical to any one of SEQ ID NOs: 18-20, 22, and 23.
[0182] In some embodiments of the disclosure, the Cas12 inactivated variant forms a complex with a guide polynucleotide; further, the complex specifically binds to a target nucleic acid; further, the complex cleaves, modifies, and / or modulates the expression of the target nucleic acid.
[0183] In some embodiments of the disclosure, the Cas12 inactivated variant forms a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence reverse complementary to a target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that interacts with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; further, the direct repeat sequence comprises or is a sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NOs: 84-91, 101-116, and 187-195; optionally, the backbone sequence does not include a tracrRNA sequence.
[0184] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 inactivated variant is optionally selected from any one or several of the following: WYR, BMCTTH, TTN, VNWTV, VNWT C, VNTTC, wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, V is A, C or G.
[0185] Optionally, the reverse complement is partially complementary or fully complementary. In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid.
[0186] In some embodiments of the disclosure, the amino acid sequence of the Cas12 inactivation variant comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 18;
[0187] Optionally, the Cas12 inactivation variant has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D480, E675, D757 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0188] In some embodiments of the disclosure, the amino acid sequence of the Cas12 inactivation variant comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 19;
[0189] Optionally, the Cas12 inactivation variant has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D619, E858, D1035 of the sequence set forth in SEQ ID NO: 19; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0190] In some embodiments of the disclosure, the amino acid sequence of the Cas12 inactivation variant comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 20;
[0191] Optionally, the Cas12 inactivation variant has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D380, E562, D647 of the sequence set forth in SEQ ID NO: 20; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0192] In some embodiments of the disclosure, the amino acid sequence of the Cas12 inactivation variant comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 22;
[0193] Optionally, the Cas12 inactivation variant has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to D609, E827, D997 of the sequence set forth in SEQ ID NO: 22; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0194] In some embodiments of the disclosure, the amino acid sequence of the Cas12 inactivation variant comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 23;
[0195] Optionally, the Cas12 inactivation variant has a mutation in the amino acid residue corresponding to any 1, any 2, or 3 of the following positions of the sequence set forth in SEQ ID NO: 23: D635, E877, D1056; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
[0196] In some embodiments of the disclosure, the PAM sequence recognized by the Cas12 inactivation variant is the same as the PAM sequence recognized by the Cas12 protein.
[0197] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 inactivation variant is optionally selected from any one or several of the following:
[0198] A, C, T, G,
[0199] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0200] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,
[0201] AAGG, AAGN, NTNN, TCGT, CNTG, NTGG, CCGN, ATAT, TGCA, NGGT, TGNT, NNTG, NCCG, ACAT, GNTG, CGCG, GACN, NTCG, TCNG, CTGC, TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG, CNTC, TGGN, NATC, GTCG, ACNC, GCNN, GACT, CTNT, NCTT, NAGG, NANC, CTTA, GTCT, ANAG, NGCN, CNNA, TCAG, ACAC, NCGG, TNNT, CAAG, ACCT, CCCA, GTNC, ANTC, GACC, AACG, TTAA, TCCG, CGCC, NCCN, TTNA, NCNT, NGCA, AGNN, AATC, GGGA, GNAN, NAGA, CGNA, GTAT, GTNA, ATNC, ACNA, GGAA, NTCC, GGCG, AATN, CNNT, AGGC, GCGN, GTGC, TTGA, AAGC, GAAG, ATNG, TGCT, TACT, CTAN, GGCT, GNGC, GTCN, CGAA, CNAC, GCCT, TAGG, ANGC, TNAA, GANT, NCNA, NCCT, AGAN, GTAA, TTTN, ATGA, TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG, NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNT,CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG, NCNG, AANG, GGTC, ACTC, TGAA, NAGN, NNCA, ACGG, TGAC, TCCN, ANNN, TCGN, TAAN, CAGG, TTAN, NGAN, NTGC, CCNC, TNTN, ATGN, GTGN, GCAT, NNGN, NNCC, CCNA, CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN, GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN, GAAC, NCCA, TTCC, NCNN, GNNN, ANGT, NTNA, CCCT, GNAA, TTNG, GTNN, GGNG, TCTA, NCAC, GANG, TTCG, CCTC, CNGG, ANNA, TCAN, ATCG, NTGA, CGTA, TTAC, GCTN, GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN,GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NNTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG, ACGN, GGCA, AAAG, TTGT, NGNA, NAAN, TATN, CGGG, CATA, ATGC, ACGC, ACCN, ATTT, TCNA, TNGC, NACA, NACC, CTCN, GGCC, TANG, AGAA, TNGA, TAGC, CAGN, GGCN, ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN, GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA, NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC, ATNN, NNNT, AGAC, AACC, ANNC, ANNG, ACAN, GTTC, TATA, GNTC, NCGT, NGNT, CGTC, CCGC, CGAC, GACG, ATTG, GNNC, CNAA, TATC, AGNA, CTNC, TTCA, ANCA, ACCC, AGTG, CCGT, ANAT, CTGT, GGGC, NTTA, NAAG, AANN, CNAN, NNCN, ANAA, ANAN, CTTG, NGNN, AGAG, TANA, TCNC, GCAA, NGNG, NAGC, NATA, ATCN, CGTT, CNGC, GATN, NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;
[0202] The N is A, T, C, or G.
[0203] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the Cas12 inactivation variant is optionally selected from any one, two or more of the following degenerate sequences or non-degenerate sequences (a non-degenerate sequence refers to any specific sequence encompassed by a degenerate sequence): WYR, BMCTTH, TTN, VNWTV, VNWTC, VNTTC, wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, and V is A, C or G.
[0204] In another aspect, the disclosure provides a fusion protein or conjugate comprising the following elements: (1) a Cas12 protein as described herein, or a Cas12 inactivation variant as described herein; and (2) a homologous or heterologous functional domain.
[0205] In this context, the reference to the Cas12 protein can encompass the Cas12 inactivation variant, depending on the contextual circumstances. However, given the importance of the Cas12 inactivation variant (non-limiting examples include the Cas12 inactivation variant fused to a deaminase for single base editing, fused to a transcriptional activation domain or a transcriptional repression domain for transcriptional regulation, etc.), the disclosure will focus on the Cas12 inactivation variant separately; this does not mean that the reference to the Cas12 protein does not necessarily include the Cas12 inactivation variant.
[0206] In some embodiments of the disclosure, a fusion protein is provided, comprising: (1) a Cas12 protein as described herein, or a Cas12 inactivation variant as described herein; and (2) a homologous or heterologous functional domain.
[0207] In some embodiments of the disclosure, a fusion protein is provided, comprising: (1) a Cas12 protein as described herein; and (2) a homologous or heterologous functional domain.
[0208] In some embodiments of the disclosure, a conjugate is provided, comprising: (1) a Cas12 protein as described herein, or a Cas12 inactivation variant as described herein; and (2) a homologous or heterologous functional domain.
[0209] In some embodiments of the disclosure, a conjugate is provided, comprising: (1) a Cas12 protein as described herein; and (2) a homologous or heterologous functional domain.
[0210] In some embodiments of the disclosure, the functional domain has an epigenetic modification activity. Including but not limited to DNA methylation, RNA methylation, RNA interference, nucleosome positioning, chromatin conformation alteration, chromatin remodeling, histone modification, long non-coding RNA sequence, etc.
[0211] In some embodiments of the disclosure, the functional domain has an enzymatic activity to modify a target nucleic acid sequence; such as nuclease activity, methyltransferase activity, demethylase activity, DNA nucleotide methyltransferase activity, DNA nucleotide demethylase activity, base deaminase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, histone acetylase activity, histone deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristoylation activity, and / or demyristoylation activity.
[0212] In some embodiments of the disclosure, the functional domain has a single base editing activity. In some embodiments of the disclosure, the functional domain is a single base editing functional domain. In some embodiments of the disclosure, the functional domain is a base converter. In some embodiments of the disclosure, the functional domain or base converter is a base deaminase. In some embodiments of the disclosure, the functional domain or base converter is an adenine deaminase or a cytosine deaminase.
[0213] In some embodiments of the disclosure, the functional domain is optionally one or more selected from the group consisting of a nuclease (e.g., Fokl), a methyltransferase, a demethylase, a DNA repair enzyme, a DNA damage enzyme, a deaminase, a dismutase, an alkylation enzyme, a depurination enzyme, an oxidation enzyme, a pyrimidine dimer formation enzyme, an integrase, a transposase, a recombinase, a polymerase, a ligase, a helicase, a photolyase, a glycosylase, a deglycosylase, an acetyltransferase, a deacetylase, a kinase, a phosphatase, a ubiquitin ligase, a deubiquitinating enzyme, an adenylation enzyme, a deadenylation enzyme, a SUMOylating enzyme, a deSUMOylating enzyme, a myristoylation enzyme, and / or a demyristoylation enzyme.
[0214] In some embodiments of the disclosure, the homologous or heterologous functional domain is optionally one, two, three, four or more selected from the group consisting of a subcellular localization signal, a DNA binding domain, a protease domain, a transcriptional activation domain, a transcriptional repression domain, a nuclease domain, a deaminase domain, a uracil DNA glycosylase domain (UDG), a uracil DNA glycosylase inhibitor domain (UGI), a methylase, a demethylase, a transcriptional release factor, a histone acetylase domain, a histone deacetylase domain, a DNA ligase, an affinity tag, a reporter tag, an affinity domain, and a reporter domain.
[0215] In some embodiments of the disclosure, the subcellular localization signal is optionally selected from the group consisting of a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, a chloroplast localization signal.
[0216] In some embodiments of the disclosure, the fusion protein or conjugate comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or more of the homologous or heterologous functional domains; the functional domains are the same or different.
[0217] In some embodiments of the disclosure, the fusion protein or conjugate optionally links 0, 1, 2, 3, 4, 5, 6, 7, 8 or more of the functional domains at the N-terminus and / or C-terminus of the Cas12 protein.
[0218] In some embodiments of the disclosure, the fusion protein comprises 1, 2, 3, 4 or more nuclear localization signals.
[0219] In some embodiments of the disclosure, the fusion protein is used to achieve base editing, e.g., in combination with a guide polynucleotide to achieve base editing. In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a deaminase domain.
[0220] In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and optionally 1 or 2 UGI domains. The fusion protein is used to achieve C→T base editing of a target nucleic acid.
[0221] In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, an adenosine deaminase domain. The fusion protein is used to achieve A→G base editing of a target nucleic acid.
[0222] In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, an adenosine deaminase domain. In some embodiments of the disclosure, the fusion protein comprises 1, 2, or 3 nuclear localization signals, and a deaminase domain. In some embodiments of the disclosure, the fusion protein comprises a UGI domain. In some embodiments of the disclosure, the fusion protein comprises 1, 2, or 3 nuclear localization signals, a deaminase domain, and 1 or 2 UGI domains.
[0223] In some embodiments of the disclosure, the fusion protein is used to effectuate transcriptional activation of a particular target gene, e.g., in combination with a guide polynucleotide to effectuate transcriptional activation of a particular target gene. In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a transcriptional activation domain.
[0224] In some embodiments of the disclosure, the fusion protein is used to effectuate transcriptional repression of a particular target gene, e.g., in combination with a guide polynucleotide to effectuate transcriptional repression of a particular target gene. In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a transcriptional repression domain.
[0225] In some embodiments of the disclosure, the fusion protein is used to effectuate methylation of a particular target sequence, e.g., in combination with a guide polynucleotide to effectuate methylation of a particular target sequence. In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a DNA methylation domain.
[0226] In some embodiments of the disclosure, the fusion protein is used to effectuate demethylation of a particular target sequence, e.g., in combination with a guide polynucleotide to effectuate demethylation of a particular target sequence. In some embodiments of the disclosure, the fusion protein comprises a nuclear localization signal, a DNA demethylation domain.
[0227] In some embodiments of the disclosure, the nuclease domain comprises a polypeptide having ssDNA cleavage activity and / or a polypeptide having dsDNA cleavage activity.
[0228] In some embodiments of the disclosure, the nuclease domain comprises a polypeptide having ssDNA cleavage activity.
[0229] In some embodiments of the disclosure, the nuclease domain comprises a polypeptide having dsDNA cleavage activity.
[0230] In some embodiments of the disclosure, the Cas12 protein or inactivated variant is directly or indirectly linked to the cognate or heterologous functional domain.
[0231] In preferred embodiments of the disclosure, the direct linkage is a covalent linkage and the indirect linkage is through an amino acid linker or a non-amino acid linker.
[0232] In more preferred embodiments of the disclosure, the cognate or heterologous functional domain is fused or conjugated to the N-terminus, C-terminus, or internally relative to the Cas12 protein or inactivated variant.
[0233] In the present disclosure, the fusion protein refers to the linkage between the element (1) and the element (2) through a peptide segment, or direct linkage; the conjugate refers to the linkage between the element (1) and the element (2) through a chemical bond of a non-peptide segment.
[0234] In some embodiments of the disclosure, the PAM sequence recognized by the fusion protein or conjugate is the same as the PAM sequence recognized by the Cas12 protein.
[0235] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the fusion protein or conjugate is optionally selected from any one or several of the following:
[0236] A, C, T, G,
[0237] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0238] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG,
[0239] AAGG, AAGN, NTNN, TCGT, CNTG, NTGG, CCGN, ATAT, TGCA, NGGT, TGNT, NNTG, NCCG, ACAT, GNTG, CGCG, GACN, NTCG, TCNG, CTGC, TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG, CNTC, TGGN, NATC, GTCG, ACNC, GCNN, GACT, CTNT, NCTT, NAGG, NANC, CTTA, GTCT, ANAG, NGCN, CNNA, TCAG, ACAC, NCGG, TNNT, CAAG, ACCT, CCCA, GTNC, ANTC, GACC, AACG, TTAA, TCCG, CGCC, NCCN, TTNA, NCNT, NGCA, AGNN, AATC, GGGA, GNAN, NAGA, CGNA, GTAT, GTNA, ATNC, ACNA, GGAA, NTCC, GGCG, AATN, CNNT, AGGC, GCGN, GTGC, TTGA, AAGC, GAAG, ATNG, TGCT, TACT, CTAN, GGCT, GNGC, GTCN, CGAA, CNAC, GCCT, TAGG, ANGC, TNAA, GANT, NCNA, NCCT, AGAN, GTAA, TTTN, ATGA, TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG, NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNT,CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG, NCNG, AANG, GGTC, ACTC, TGAA, NAGN, NNCA, ACGG, TGAC, TCCN, ANNN, TCGN, TAAN, CAGG, TTAN, NGAN, NTGC, CCNC, TNTN, ATGN, GTGN, GCAT, NNGN, NNCC, CCNA, CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN, GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN, GAAC, NCCA, TTCC, NCNN, GNNN, ANGT, NTNA, CCCT, GNAA, TTNG, GTNN, GGNG, TCTA, NCAC, GANG, TTCG, CCTC, CNGG, ANNA, TCAN, ATCG, NTGA, CGTA, TTAC, GCTN, GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN,GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NNTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG, ACGN, GGCA, AAAG, TTGT, NGNA, NAAN, TATN, CGGG, CATA, ATGC, ACGC, ACCN, ATTT, TCNA, TNGC, NACA, NACC, CTCN, GGCC, TANG, AGAA, TNGA, TAGC, CAGN, GGCN, ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN, GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA, NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC, ATNN, NNNT, AGAC, AACC, ANNC, ANNG, ACAN, GTTC, TATA, GNTC, NCGT, NGNT, CGTC, CCGC, CGAC, GACG, ATTG, GNNC, CNAA, TATC, AGNA, CTNC, TTCA, ANCA, ACCC, AGTG, CCGT, ANAT, CTGT, GGGC, NTTA, NAAG, AANN, CNAN, NNCN, ANAA, ANAN, CTTG, NGNN, AGAG, TANA, TCNC, GCAA, NGNG, NAGC, NATA, ATCN, CGTT, CNGC, GATN, NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TT TG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;
[0240] The N is A, T, C, or G.
[0241] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the fusion protein is optionally selected from any one, two or more degenerate sequences or non-degenerate sequences (non-degenerate sequences refer to any specific sequence encompassed by the degenerate sequences) of WYR, BMCTTH, TTN, VNWTV, VNWT C, VNTTC, wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, and V is A, C or G.
[0242] In some embodiments of the disclosure, the sequence recognized by the fusion protein is a PAM of 5'-WYR-3'.
[0243] In some embodiments of the disclosure, the sequence recognized by the conjugate is a PAM of 5'-WYR-3'.
[0244] In some embodiments of the disclosure, the PAM sequence (5'→3') recognized by the conjugate is optionally selected from any one, two or more degenerate sequences or non-degenerate sequences (non-degenerate sequences refer to any specific sequence encompassed by the degenerate sequences) of WYR, BMCTTH, TTN, VNWTV, VNWT C, VNTTC, wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, and V is A, C or G.
[0245] In some embodiments of the disclosure, a fusion protein is provided, the fusion protein comprising a Cas12 protein, and a homologous or heterologous functional domain;
[0246] the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 18;
[0247] Optionally, the Cas12 protein is non-natural, or, engineered;
[0248] Optionally, the Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves a target nucleic acid, or, the complex sequence-specifically binds a target nucleic acid, but does not cleave the target nucleic acid; optionally, the complex is non-natural, or, engineered;
[0249] Optionally, the guide polynucleotide comprises a guide sequence and a scaffold sequence; optionally, the guide sequence is reverse-complementary to a target nucleic acid, the scaffold sequence interacts with the Cas12 protein; optionally, the scaffold sequence is a direct repeat sequence; optionally, the guide sequence is located at the 5’ end or the 3’ end of the scaffold sequence; optionally, the guide polynucleotide is non-natural, or, engineered;
[0250] Optionally, the Cas12 protein recognition sequence is a PAM of 5’-WYR-3’, wherein W is A or T, Y is C or T, and R is A or G; further optionally, the Cas12 protein recognition sequence is a PAM of 5’-ACA-3’, 5’-TCA-3’, 5’-ATA-3’, 5’-TTA-3’, 5’-ACG-3’, 5’-TCG-3’, 5’-ATG-3’, 5’-TTG-3’, and / or 5’-TTN-3’;
[0251] Optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in any one of SEQ ID NOs: 84-86, 187-195; further optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to the sequence set forth in SEQ ID NO: 84;
[0252] Optionally, the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18, and wherein the Cas12 protein comprises an amino acid sequence that is at least 80% identical to the sequence set forth in SEQ ID NO: 18,293、294、295、296、297、298、299、300、301、302、303、305、306、308、309、310、313、315、316、317、318、319、320、321、322、323、324、325、327、328、329、330、331、332、333、334、335、336、337、339、340、341、342、343、344、345、346、347、348、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、431、432、433、435、436、437、439、440、441、442、443、444、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、467、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、496、497、499、500、501、502、503、504、506、507、508、509、510、511、512、513、514、515、516、517、518、519、520、521、522、523、524、525、526、527、528、529、531、532、533、534、535、536、537、538、539、540、541、542、543、544、545、546、547、548、549、550、552、553、555、556、557、558、559、560、561、562、563、564、565、566、567、568、569、570、571、572、573、574、575、576、577、578、579、580、581、582、583、584、585、586、587、589、590、592、593、594、595、596、597、598、599、601、602、603、604、605、606、607、608、609、610、611、612、613、614、615、616、618、619、620、621、622、623、624、625、626、627、628、630、631、632、633、634、635、636、637、638、639、640、641、642、643、644、645、646、647、648、649、650、651、652、653、654、655、656、657、658、659、660、661、662、663、664、665、666、667、668、669、670、671、672、673、674、675、676、678、679、680、681、683、684、685、686、688、689、691、692、693、694、695、696、697、698、699、700、701、702、703、704、705、706、707、708、709、710、711、712、713、715、716、717、719、720、721、722、723、724、725、727、728、729、730、731、732、733、734、736、737、738、739、740、741、742、743、744、745、746、747、748、749、751、752、753、754、755、756、758、759、760、761、762、764、765、766、767、768、769、771、772、773、774、775、776、779、780、781、782、783、784、785、786、787、789、790、791、792、794、795、797、798、800、801、802、804、805、806、807、808、809、810、811、812、813、814、815、817、818、819、821、822、823、824、825、826、827、828、829、830、831、832、833、834、835、836、837、838、839、at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, or at least 12 of amino acid residues at positions 840, 841, 842, 844, 845, 846, 847, 848, 849, 850, 851, 852, 853, 854, 855, 856, 857, 858, 859, 860, 862, 863, 864, 865, 866, 867, 868, 870, 872, 873, 874, 875, 876, 877, 879, 880, 881, 882, 883, 884, 885, 886, 887, 888, 890, and 891 have a mutation; optionally, the mutation is a mutation to any other natural amino acid residue; further optionally, the mutation is a mutation to residue R, H, K, or A; further optionally, the mutation is a mutation to residue R; in some embodiments of the disclosure, the mutation is a mutation to residue A;
[0253] optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, or more of the amino acid residues corresponding to positions N5, D9, E58, S100, N115, K142, C148, S147, K232, S245, I251, Y263, D279, A297, L300, E303, L337, M378, N394, T396, T443, K458, T468, K533, F537, F548, N550, D697, A706, I788 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to residue R;
[0254] optionally, the functional domain has epigenomic modification activity; optionally, the functional domain has epigenetic modification activity; optionally, the functional domain has epigenetic modification activity; optionally, the epigenomic modification, epigenetic modification, or epigenetic modification comprises, but is not limited to, DNA methylation, RNA methylation, RNA interference, nucleosome positioning, chromatin conformation alteration, chromatin remodeling, histone modification, modification of long non-coding RNA sequence;
[0255] optionally, the functional domain is an epigenomic modification functional domain; optionally, the functional domain is an epigenetic modification functional domain; optionally, the functional domain is an epigenetic modification functional domain;
[0256] Optionally, the functional domain is optionally selected from one or more of: a nuclease (e.g., Fokl), a DNA methylase, a DNA demethylase, a histone methylase, a histone demethylase, a DNA repair enzyme, a DNA damaging enzyme, a base deaminase (including but not limited to an adenine deaminase, a cytosine deaminase), a dismutase, an alkylating enzyme, a depurinase, an oxidase, a pyrimidine dimer forming enzyme, an integrase, a transposase, a recombinase, a polymerase, a ligase, a helicase, a photolyase, a glycosylase, a deglycosylase, an acetyltransferase, a deacetylase, a kinase, a phosphatase, a ubiquitin ligase, a deubiquitinating enzyme, an adenylylating enzyme, a deadenylase, a SUMOylating enzyme, a deSUMOylating enzyme, a myristoylating enzyme, and / or a demyristoylating enzyme; optionally, the functional domain is an adenine deaminase or a cytosine deaminase.
[0257] In some embodiments of the disclosure, the epigenome modification, epigenetic modification, and epigenetic modification can be used interchangeably.
[0258] Some sites that maintain or improve editing activity (e.g., editing efficiency is at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 110%, at least 120%, at least 130%, or at least 140% of that of the wild-type protein of SEQ ID NO: 18) after mutation and some sites that can significantly reduce editing activity (e.g., editing efficiency is reduced by at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% compared to the wild-type protein of SEQ ID NO: 18) after mutation have been listed in Tables 5 and 6 of the disclosure. These sites can be mutation sites of the Cas12 protein, and in some embodiments, the fusion protein of the disclosure is obtained by mutating the sites to other amino acid residues.
[0259] In another aspect, the disclosure provides a technical solution of an isolated nucleic acid encoding a Cas12 protein as described in the disclosure, a Cas12 inactivation variant as described in the disclosure, or a fusion protein or conjugate as described in the disclosure.
[0260] In some embodiments of the disclosure, the nucleic acid encodes a Cas12 protein as described in the disclosure or a fusion protein as described in the disclosure.
[0261] In some embodiments disclosed herein, the nucleic acid is a DNA or RNA sequence. In some embodiments disclosed herein, the nucleic acid contains modifications. In some embodiments disclosed herein, the nucleic acid contains modified nucleotides. In some embodiments disclosed herein, the nucleic acid is a DNA sequence containing RNA base modifications. In some embodiments disclosed herein, the nucleic acid is an RNA sequence containing DNA base modifications. In some embodiments disclosed herein, the nucleic acid is mRNA.
[0262] In some embodiments disclosed herein, the nucleic acid comprises biocompatible natural or non-natural nucleotide modifications.
[0263] In some embodiments disclosed herein, the nucleic acid optionally comprises one, two, or more nucleotide modifications selected from: 2'-O-methylation, pseudouridine (Ψ), N6-methyladenosine (N... 6 -methyladenosine,m 6 A) 5-methylcytidine (m 5 C), 7-methylguanosine (m) 7 G), 1-methyladenosine (m 1 A) 5-hydroxymethylcytidine (5hmC).
[0264] In some embodiments disclosed herein, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the uracil in the nucleic acid is replaced by pseudouracil. In some embodiments disclosed herein, the nucleic acid is mRNA, and at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the uracil in the mRNA is replaced by pseudouracil.
[0265] In some embodiments disclosed herein, the nucleic acid is mRNA, and the mRNA contains modifications. In some embodiments disclosed herein, the nucleic acid is mRNA, and the mRNA contains modified nucleotides. In some embodiments disclosed herein, the nucleic acid is mRNA, and the mRNA contains a cap structure at its 5' end. In some embodiments disclosed herein, the nucleic acid is mRNA, and the mRNA contains a Cap1 cap structure at its 5' end.
[0266] In some embodiments of the disclosure, the nucleic acid is mRNA comprising modified nucleotides located in the 5' untranslated region (5'-UTR), open reading frame (ORF), or 3' untranslated region (3'-UTR) of the mRNA. One skilled in the art would know that the distribution of the modified nucleotides can be optimized according to the desired expression of the protein of interest.
[0267] In some embodiments of the disclosure, the mRNA comprises a variable length 5' untranslated region (5'-UTR), wherein the 5'-UTR can include a cis-acting element that enhances translation initiation efficiency. In some embodiments of the disclosure, the mRNA comprises a variable length 3' untranslated region (3'-UTR), wherein the 3'-UTR can include an element that enhances mRNA stability or interacts with RNA binding proteins.
[0268] In some embodiments of the disclosure, the mRNA comprises a 5'-UTR comprising a fragment of the untranslated region derived from the human albumin gene, the a-globin gene, the b-globin gene, the g-globin gene, or a liver highly expressed gene.
[0269] In some embodiments of the disclosure, the mRNA comprises one or more 3'-UTR sequences derived from a human highly expressed gene to enhance stability and translation efficiency in human cells.
[0270] In some embodiments of the disclosure, the mRNA comprises a 3'-UTR comprising a fragment of the untranslated region derived from the human albumin gene, the a-globin gene, the b-globin gene, the g-globin gene, or a liver highly expressed gene.
[0271] In some embodiments of the disclosure, the mRNA comprises a 5'-UTR comprising a fragment of the 5'-UTR derived from the human g-globin gene and a 3'-UTR comprising a fragment of the 3'-UTR derived from the human g-globin gene.
[0272] In some embodiments of the disclosure, the mRNA comprises one or more poly(A) tails, which can be about 30-300 adenosine residues in length, to improve mRNA stability and its binding capacity to the translation initiation complex.
[0273] In some embodiments of the disclosure, the mRNA comprises an optimized codon usage pattern, in which codon combinations in the open reading frame are optimized for high frequency codons commonly found in the target host of expression (e.g., human cells, mammalian cells, or HEK293 cells) to improve translation efficiency.
[0274] In some embodiments of the disclosure, the mRNA comprises one, two or more structural features selected from the group consisting of: an optimized 5'-UTR sequence, a Cap1 cap structure, nucleotides containing 2'-O-methylation modification, a pseudouridine substitution, a 3'-UTR enhancer element, a polyadenylate tail, a multiply modified open reading frame.
[0275] In some embodiments of the disclosure, the mRNA is introduced with a cap structure after in vitro transcription by enzymatic or co-transcriptional capping, which includes the use of 7-methylguanine (m 7 G) as the cap group, and the introduction of a 2'-O-methyl modification at the first nucleotide to form a Cap1 structure.
[0276] In some embodiments of the disclosure, the mRNA further comprises one or more RNA stabilization structural elements, which are located in the 3'-UTR region of the mRNA and can include a low GC content region, an AU-rich element, or an exogenous RNA stabilization sequence.
[0277] In some embodiments of the disclosure, the cap structure and modification of the mRNA can be combined to reduce the probability of activation of pattern recognition receptors associated with innate immunity (e.g., RIG-I, MDA5, TLR7 / 8) to attenuate the interferon response.
[0278] In some embodiments of the disclosure, the mRNA is obtained from cell-free transcription and double-stranded RNA impurities are removed by high pressure liquid chromatography (HPLC) or other purification means to reduce non-specific immune stimulation and improve safety in therapeutic use.
[0279] In preferred embodiments of the disclosure, the nucleic acid is codon-optimized for expression in a cell.
[0280] In more preferred embodiments of the disclosure, the nucleic acid is codon-optimized for expression in a eukaryote, a mammal such as a human or a non-human mammal, a plant, an insect, a bird, a reptile, a rodent (e.g., a mouse, a rat), a fish, a worm / nematode, or a yeast.
[0281] In another aspect, the disclosure provides a technical solution of a CRISPR-Cas12 system, which comprises:
[0282] a) a Cas12 protein as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, or an isolated nucleic acid as described herein; and
[0283] b) a guide polynucleotide, or a polynucleotide sequence encoding the guide polynucleotide;
[0284] the Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate forms a complex with the guide polynucleotide; the guide polynucleotide comprises a guide sequence engineered to direct sequence-specific binding of the complex to a target nucleic acid.
[0285] the isolated nucleic acid encodes a Cas12 protein as described herein, a Cas12 inactivated variant as described herein, or a fusion protein or conjugate as described herein.
[0286] In some embodiments of the disclosure, the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence.
[0287] In some embodiments of the disclosure, the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence set forth in any one of SEQ ID NOs: 1-35.
[0288] In some embodiments of the disclosure, the Cas12 protein belongs to Cas12h subtype (V-H), the Cas12 protein is capable of forming a complex with a guide polynucleotide, the complex targets binding to a target nucleic acid within a eukaryotic cell. In some embodiments of the disclosure, the Cas12 protein is capable of forming a complex with a guide polynucleotide, the complex targets binding to and cleaving a target nucleic acid within a eukaryotic cell.
[0289] Optionally, the fusion protein comprises an amino acid sequence of the Cas12 protein.
[0290] In some embodiments of the disclosure, the direct repeat sequence has at least 50% identity to any one of SEQ ID NOs: 36-170, 187-195.
[0291] In some embodiments of the disclosure, the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence. Further, in some specific embodiments, the direct repeat sequence has at least 50% identity to any one of SEQ ID NOs: 36-170, 187-195. In some specific embodiments, the direct repeat sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 36-170, 187-195. Further, in some specific embodiments, the direct repeat sequence comprises or is the sequence set forth in any one of SEQ ID NOs: 36-170, 187-195.
[0292] In some embodiments of the disclosure, the guide sequence comprises 15-60 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-50 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-40 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-35 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-30 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-25 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 18-25 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-25 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 18-22 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-22 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides.
[0293] In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid, the guide sequence is 90-100% complementary to the target nucleic acid.
[0294] In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid.
[0295] In some embodiments of the disclosure, the guide sequence hybridizes to the target nucleic acid with no more than one nucleotide mismatch.
[0296] In some embodiments of the disclosure, the direct repeat sequence comprises 15-100 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-90 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-80 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-70 nucleotides. In some embodiments of the disclosure, the direct repeat sequence comprises 15-60 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-50 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15-40 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-40 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 20-30 nucleotides. In some embodiments of the disclosure, the guide sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0297] In some embodiments of the disclosure, the guide sequence is located at the 3' end of the direct repeat sequence.
[0298] In some embodiments of the disclosure, the guide sequence is located at the 5' end of the direct repeat sequence.
[0299] In some embodiments of the disclosure, the guide polynucleotide does not comprise a tracrRNA sequence.
[0300] In some embodiments of the disclosure, the guide polynucleotide further comprises a tracrRNA sequence.
[0301] In some embodiments of the disclosure, the tracrRNA base pairs with the direct repeat sequence. In general, the base pairing is partial. In some embodiments of the disclosure, the tracrRNA interacts with the direct repeat sequence.
[0302] In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 1-10 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 4 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a 5'-GAAA-3' sequence.
[0303] In some embodiments of the disclosure, the tracrRNA sequence is located at the 3' end of the direct repeat sequence.
[0304] In some embodiments of the disclosure, the tracrRNA sequence is located at the 5' end of the direct repeat sequence.
[0305] In some embodiments of the disclosure, the tracrRNA comprises 10-200 nucleotides. In some embodiments of the disclosure, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In some embodiments of the disclosure, the tracrRNA comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0306] In preferred embodiments of the disclosure, the guide polynucleotide is a guide polynucleotide as described herein.
[0307] In some embodiments of the disclosure, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA.
[0308] In preferred embodiments of the disclosure, the DNA is eukaryotic DNA; preferably the eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, avian DNA, reptilian DNA, rodent DNA, fish DNA, worm / nematode DNA, or yeast DNA.
[0309] In some embodiments of the disclosure, the target nucleic acid is a disease or disorder associated gene or a signaling biochemical pathway associated gene, or the target nucleic acid is a reporter gene; for example, the disease or disorder is a hematological disease or disorder, an ophthalmic disease or disorder, a neurological disease or disorder, a respiratory disease or disorder, a liver disease or disorder, a metabolic system disease or disorder, a cancer, or an infectious disease.
[0310] In some embodiments of the disclosure, the target nucleic acid is optionally selected from a list of target nucleic acids / target genes described in patent application publication number WO2025061113A1.
[0311] In another aspect, the disclosure provides a technical solution of a vector system comprising one or more recombinant vectors, the recombinant vector comprising an isolated nucleic acid as described in the disclosure, or a CRISPR-Cas12 system as described in the disclosure.
[0312] In some embodiments of the disclosure, the recombinant vector further comprises a regulatory sequence.
[0313] In some embodiments of the disclosure, the vector system comprises one or more recombinant vectors, the recombinant vector comprising a polynucleotide sequence encoding a Cas12 protein, a Cas12 inactivated variant or a fusion protein or conjugate as described in the disclosure, and a polynucleotide sequence encoding the guide polynucleotide.
[0314] In some embodiments of the disclosure, the polynucleotide sequence encoding the Cas12 protein, a Cas12 inactivated variant or a fusion protein or conjugate is operably linked to a regulatory sequence 1.
[0315] In some embodiments of the disclosure, the polynucleotide sequence encoding the guide polynucleotide is operably linked to a regulatory sequence 2.
[0316] Further, in some embodiments of the disclosure, the regulatory sequence 1 and the regulatory sequence 2 are the same or different sequences.
[0317] In preferred embodiments of the disclosure, the regulatory sequence is optionally selected from one or more of a promoter, an enhancer, an internal ribosome entry site, and a transcription termination signal; the promoter is for example a constitutive promoter, an inducible promoter, a broad-spectrum promoter, or a tissue-specific promoter, and / or the transcription termination signal is for example a polyadenylation signal or a poly-U sequence.
[0318] In some embodiments of the disclosure, the backbone of the recombinant vector is an adeno-associated viral vector, a lentiviral vector, a ribonucleoprotein complex, or a virus-like particle.
[0319] In preferred embodiments of the disclosure:
[0320] when the backbone is an adeno-associated viral vector, the adeno-associated viral vector is a recombinant adeno-associated viral vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShHlO, AAVrh10, or AAVrh74;
[0321] when the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; preferably, the isolated nucleic acid is linked to an aptamer sequence;
[0322] when the backbone is a virus-like particle, the isolated nucleic acid is linked to a gene encoding a gag protein.
[0323] In another aspect, the present disclosure provides a technical solution: a delivery system, comprising: (1) a delivery tool, and (2) a Cas12 protein as described in the present disclosure, a guide polynucleotide as described in the present disclosure, a Cas12 inactivated variant as described in the present disclosure, a fusion protein or conjugate as described in the present disclosure, a nucleic acid as described in the present disclosure, a CRISPR-Cas12 system as described in the present disclosure, or a vector system as described in the present disclosure.
[0324] In preferred embodiments of the present disclosure, the delivery tool is a virus, a lipid nanoparticle, a nanoparticle, a liposome, an exosome, a microvesicle, or a gene gun.
[0325] In more preferred embodiments of the present disclosure, the delivery tool is a lipid nanoparticle comprising the guide polynucleotide and an mRNA encoding the Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate.
[0326] In another aspect, the present disclosure provides a technical solution: a cell comprising a Cas12 protein as described in the present disclosure, a guide polynucleotide as described in the present disclosure, a Cas12 inactivated variant as described in the present disclosure, a fusion protein or conjugate as described in the present disclosure, a nucleic acid as described in the present disclosure, a CRISPR-Cas12 system as described in the present disclosure, or a vector system as described in the present disclosure.
[0327] In some embodiments of the present disclosure, the cell is a prokaryotic cell.
[0328] In some embodiments of the disclosure, the cell is a eukaryotic cell.
[0329] In some embodiments of the disclosure, the eukaryotic cell is a mammalian cell.
[0330] In another aspect, the disclosure provides a technical solution of: a pharmaceutical composition comprising a Cas12 protein as described in the disclosure, a guide polynucleotide as described in the disclosure, a Cas12 inactivated variant as described in the disclosure, a fusion protein or conjugate as described in the disclosure, a nucleic acid as described in the disclosure, a CRISPR-Cas12 system as described in the disclosure, a vector system as described in the disclosure, a delivery system as described in the disclosure, or a cell as described in the disclosure.
[0331] Preferably, the pharmaceutical composition comprises a pharmaceutically acceptable excipient.
[0332] In another aspect, the disclosure provides a technical solution of: a kit comprising a Cas12 protein as described in the disclosure, a guide polynucleotide as described in the disclosure, a Cas12 inactivated variant as described in the disclosure, a fusion protein or conjugate as described in the disclosure, a nucleic acid as described in the disclosure, a CRISPR-Cas12 system as described in the disclosure, a vector system as described in the disclosure, a delivery system as described in the disclosure, or a cell as described in the disclosure.
[0333] In preferred embodiments of the disclosure, the kit further comprises a Cut Buffer. The Cut Buffer can be any buffer known in the art suitable for cleaving a target nucleic acid by a Cas12 protein.
[0334] In another aspect, the disclosure provides a technical solution of: use of a Cas12 protein as described in the disclosure, a guide polynucleotide as described in the disclosure, a Cas12 inactivated variant as described in the disclosure, a fusion protein or conjugate as described in the disclosure, a nucleic acid as described in the disclosure, a CRISPR-Cas12 system as described in the disclosure, a vector system as described in the disclosure, a delivery system as described in the disclosure, a cell as described in the disclosure, a pharmaceutical composition as described in the disclosure, or a kit as described in the disclosure in the preparation of a reagent or a drug for diagnosis, treatment and / or prevention of a disease or disorder associated with a target nucleic acid.
[0335] In some embodiments of the disclosure, the disease or condition is a hematological disease or condition, an ophthalmological disease or condition, a neurological disease or condition, a respiratory disease or condition, a liver disease or condition, a metabolic disease or condition, a cancer, or an infectious disease; and / or, the agent or drug is for cleaving or nicking one or more target nucleic acid molecules, activating or upregulating expression of one or more target nucleic acid molecules, activating or repressing transcription of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling or detecting one or more target nucleic acid molecules, binding to one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules.
[0336] In some embodiments of the disclosure, the target nucleic acid is optionally selected from the list of target nucleic acids / target genes as described in Table 27 of patent application publication no. WO2025061113A1, and the disease or condition is the corresponding disease or condition as listed in Table 27 of patent application publication no. WO2025061113A1.
[0337] In another aspect, the disclosure provides a technical solution of a method of detecting, binding or cleaving a target nucleic acid, the method comprising contacting the target nucleic acid with a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein.
[0338] In preferred embodiments of the disclosure, the method is a method for non-diagnostic and / or therapeutic purposes; and / or the fusion protein or conjugate comprises a detectable label, for example a label detectable by fluorescence, Southern blotting or FISH.
[0339] In more preferred embodiments of the disclosure, when the method is cleaving a target nucleic acid, the method further comprises performing a cleavage reaction using a cleavage buffer. The cleavage buffer can be any buffer known in the art suitable for cleaving a target nucleic acid by a Cas12 protein.
[0340] In another aspect, one technical solution provided by the disclosure is a method of altering a state of a cell, comprising contacting a cell with a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein, thereby altering the state of the cell.
[0341] In some embodiments of the disclosure, the method results in one or more of: an increase or decrease in expression of a particular gene, induction of cell senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, cell growth promotion and / or cell growth inhibition in vitro or in vivo, induction of anergy in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo.
[0342] In more preferred embodiments of the disclosure, the method is a method for non-diagnostic and / or therapeutic purposes.
[0343] In another aspect, one technical solution provided by the disclosure is a method of diagnosing, treating or preventing a disease or disorder associated with a target nucleic acid, comprising administering to a sample of a subject in need thereof or to a subject in need thereof a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein.
[0344] In some embodiments of the disclosure, the target nucleic acid is optionally selected from the list of target nucleic acids / target genes described in Table 27 of patent application publication no. WO2025061113A1, and the disease or disorder is the corresponding disease or disorder as listed in Table 27 of patent application publication no. WO2025061113A1.
[0345] In some embodiments of the disclosure, the disease or disorder is a hematological disease or disorder, an ophthalmological disease or disorder, a neurological disease or disorder, a respiratory disease or disorder, a liver disease or disorder, a metabolic disease or disorder, a cancer, or an infectious disease.
[0346] In another aspect, a technical solution provided by the disclosure is: the Cas12 protein as described in the disclosure, the guide polynucleotide as described in the disclosure, the Cas12 inactivation variant as described in the disclosure, the fusion protein or conjugate as described in the disclosure, the nucleic acid as described in the disclosure, the CRISPR-Cas12 system as described in the disclosure, the vector system as described in the disclosure, the delivery system as described in the disclosure, the cell as described in the disclosure, the pharmaceutical composition as described in the disclosure or the kit as described in the disclosure, which is used for diagnosing, treating or preventing a disease or disorder related to a target nucleic acid.
[0347] In some embodiments of the disclosure, the target nucleic acid is optionally selected from a series of target nucleic acids / target genes described in Table 27 of the patent application with publication number WO2025061113A1, and the disease or disorder is the corresponding disease or disorder as listed in Table 27 of the patent application with publication number WO2025061113A1.
[0348] In some embodiments of the disclosure, the disease or disorder is a hematological disease or disorder, an ophthalmic disease or disorder, a nervous system disease or disorder, a respiratory system disease or disorder, a liver disease or disorder, a metabolic system disease or disorder, a cancer or an infectious disease.
[0349] In some embodiments of the disclosure, the disease or disorder is optionally selected from the group consisting of: hemophilia A, Best vitelliform macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency disorder, CLN2 disease, Niemann Pick type C, Dravet syndrome, FOXG1 syndrome, GM1 gangliosidosis, GM2 gangliosidoses, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IIA, Mucopolysaccharidosis type IIIA, Mucopolysaccharidosis type IIIB, Gaucher disease type III, Mucopolysaccharidosis type II, Type II diabetes, Mucopolysaccharidosis type IV, Gaucher disease type I, Mucopolysaccharidosis type I, Type I diabetes, Usher syndrome type I, KCNQ2 encephalopathy, Leber hereditary optic neuropathy, Leigh syndrome, Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, alpha 1-antitrypsin deficiency, alpha-mannosidosis, alpha-thalassemia, beta-thalassemia, Alzheimer's disease, Bardet-Biedl syndrome, Best disease, Leukocyte adhesion deficiency type I, Galactosemia, Bladder cancer, Overactive bladder, Phenylketonuria, Nasopharyngeal carcinoma, Bietti crystalline dystrophy, Pyruvate kinase deficiency, Erectile dysfunction, Autosomal recessive congenital ichthyosis, Adult polyglucosan body disease, Traumatic arthritis, Homozygous familial hypercholesterolemia, Fragile X syndrome, Thalassemia, Hypophosphatasia, Epilepsy, Multiple myeloma, Multiple system atrophy, Frontotemporal dementia, Catecholaminergic polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, Aromatic amino acid decarboxylase deficiency, Radiation-induced xerostomia, Non-Hodgkin lymphoma, Non-muscle invasive bladder cancer, Non-alcoholic fatty liver disease, Non-small cell lung cancer, Hypertrophic cardiomyopathy, Hypertrophic scarring, Obesity, Charcot-Marie-Tooth disease type 1A, Charcot-Marie-Tooth disease type 2A, Pulmonary hypertension, Friedreich's ataxia, Peritoneal cancer, Liver cancer, Hepatocellular carcinoma, Dry age-related macular degeneration, Sjogren's syndrome, Hyperuricemia, Hyperlipidemia, Gaucher disease, Autism spectrum disorder, Osteoarthritis, Bone marrow failure syndromes, Cysteine I, Melanoma, Huntington's disease, Amyotrophic lateral sclerosis, Urge urinary incontinence, Acute intermittent porphyria, Acute lymphoblastic leukemia, Spinocerebellar ataxia, Spinal muscular atrophy with respiratory distress type 1, Spinal muscular atrophy, Familial black-out dementia, Methylmalonic acidemia, Thyroid cancer, Pseudohypertrophic muscular dystrophy, Anaplastic astrocytoma, Intermittent claudication, Junctional epidermolysis bullosa, Glioma, Glioblastoma, Corneal graft rejection, Colorectal cancer, Progressive multifocal leukoencephalopathy, Progressive familial intrahepatic cholestasis, Giant axonal neuropathy, Canavan disease, Cocaine addiction, Krabbe disease,Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Ritter syndrome. Triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignant osteosclerosis. Nutritional bullous epidermolysis, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0350] In some embodiments disclosed herein, the genes associated with thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR.
[0351] The Leber's hereditary optic neuropathy-associated gene includes but is not limited to MT-ND4;
[0352] The AATD liver disease-associated gene includes but is not limited to AATD;
[0353] The AATD lung disease-associated gene includes but is not limited to AATD;
[0354] The graft versus host disease-associated gene includes but is not limited to thymidine kinase gene;
[0355] The inherited retinal dystrophy-associated gene includes but is not limited to RPE65;
[0356] The spinal muscular atrophy-associated gene includes but is not limited to SMN1;
[0357] The osteoarthritis-associated gene includes but is not limited to TGF-β1;
[0358] The hemophilia A-associated gene includes but is not limited to factor VIII;
[0359] The hemophilia B-associated gene includes but is not limited to factor IX;
[0360] The cystic fibrosis-associated gene includes but is not limited to CFTR;
[0361] The Parkinson's disease-associated gene includes but is not limited to Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB and REST;
[0362] The Usher syndrome-associated gene includes but is not limited to USH2A;
[0363] The alpha-thalassemia, beta-thalassemia, sickle cell disease-associated gene includes but is not limited to BCL11A, HBG, HBA and HBB;
[0364] The pulmonary hypertension-associated gene includes but is not limited to eNOS;
[0365] The Stargardt disease-associated gene includes but is not limited to ABCA4;
[0366] The genes related to age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang 2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0367] The genes related to glaucoma include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Rh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12;
[0368] The genes related to idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0369] The genes related to hyperlipidemia include, but are not limited to, PCSK9;
[0370] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0371] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0372] The genes related to anemia of chronic kidney disease include, but are not limited to, EPO;
[0373] The genes related to congenital blindness include, but are not limited to, RPE65;
[0374] The genes related to retinitis pigmentosa include, but are not limited to, PDE6B;
[0375] The genes related to phenylketonuria include, but are not limited to, PAH; and / or
[0376] The genes related to epilepsy include, but are not limited to, GAT1.
[0377] On the basis of common general knowledge in the art, the above-mentioned preferred conditions can be combined in any manner, i.e., to obtain each preferred embodiment of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0378] FIG. 1 shows a map of vector P15A-C12-334-HDV.
[0379] Figure 2 shows a schematic diagram of the site targeted for cleavage in the plasmid elimination experiment, with a 7 nt random sequence at the 5’-end of the target site. The sequences in the figure and their reverse complements are set forth in SEQ ID NOs: 259 and 287, respectively.
[0380] Figure 3 shows that the PAM motif grabbed by C12-334 in the plasmid elimination experiment is 5’-WYR-3’.
[0381] Figure 4 shows that the PAM motif grabbed by C12-335 in the plasmid elimination experiment is 5’-BMCTTH-3’.
[0382] Figure 5 shows that the PAM motif grabbed by C12-336 in the plasmid elimination experiment is 5’-TTN-3’.
[0383] Figure 6 shows that the PAM motif grabbed by C12-340 in the plasmid elimination experiment is 5’-VNWTV-3’, 5’-VNWTC-3’, or 5’-VNTTC-3’.
[0384] Figure 7 shows that the PAM motif grabbed by C12-341 in the plasmid elimination experiment is 5’-TTN-3’.
[0385] Figure 8 shows a display of the major indel results generated by C12-334 targeting TTR gene editing. The sequences in the figure are set forth in SEQ ID NOs: 260-268.
[0386] Figures 9A, 9B show a phylogenetic relationship diagram of the Cas proteins of the present disclosure with known Cas12 subtype proteins (phylogenetic tree constructed with FastTree after multiple sequence alignment). In the phylogenetic tree, the partial proteins of the present disclosure form a separate, clearly separated branch (different cluster [CLUSTER]) compared to the known Cas12 proteins, i.e., are not mixed with the known Cas12 proteins.
[0387] Figures 10A, 10B show the editing efficiency based on cleavage activity of each mutant tested in the SSA reporter cell line.
[0388] Figure 11 shows the Alphafold3 structure prediction results based on the C12-334+gRNA+target DNA ternary complex, demonstrating that many of the advantage mutants with improved gene editing efficiency screened in the experiment are related to the binding mechanism of the Cas protein and the nucleic acid (sgRNA / dsDNA).
[0389] Figure 12 shows the editing efficiency of C12-334 different mutants in combination with C12-334-HPRT1-sgRNA05 targeting the HPRT1 gene.
[0390] FIG. 13A-FIG. 13C show the indel frequencies and distribution detected by NGS after the mRNA of C12-334 wild type or mutants were targeted the HPRT1 gene in combination with the modified gRNA (C12-334-dmHPRT1-sgRNA05-01), respectively. The editing efficiency of wild type and 2 mutants reached 48.80%, 90.88% and 92.77%, respectively. The sequences in FIG. 13A are set forth in SEQ ID NOs: 269-275; the sequences in FIG. 13B are set forth in SEQ ID NOs: 276-284; the sequences in FIG. 13C are set forth in SEQ ID NOs: 276-286. DETAILED DESCRIPTION
[0391] In the present disclosure, the scientific and technical terms used herein have the meanings commonly understood by a person of ordinary skill in the art, unless otherwise indicated. Also, the molecular genetic, nucleic acid chemical, chemical, molecular biological, biochemical, cell culture, microbiological, cell biological, genomic, and recombinant DNA, among other procedural steps used herein are in accordance with conventional methods well established and commonly employed in the respective fields. Also, for better understanding of the present disclosure, the definitions and explanations of relevant terms are provided below.
[0392] In the present disclosure, “a plurality of” can refer to ≥2; “plurality” can refer to greater than or equal to two. Depending on the context, “a plurality of” can refer to ≥3; “plurality” can refer to greater than or equal to 3.
[0393] In the present disclosure, depending on the context, “cleave” can refer to cleaving the backbone of a polynucleotide chain; non-limiting examples include completely breaking a single-stranded DNA, breaking one of the single strands of a double-stranded DNA, or breaking both of the single strands of a double-stranded DNA.
[0394] In the present disclosure, depending on the context, “modification” can refer to other forms of chemical reactions on nucleic acid strands other than “cleavage”; including but not limited to base substitution, addition and / or deletion, and methylation, demethylation of nucleic acid strands. Non-limiting examples include substitution of a base on a target nucleic acid strand, such as A→G, C→T, T→C or G→A nucleotide mutation, and other types of nucleotide mutations (such as A→T, C→G, T→A, G→C, etc.), for example, by single base editing (Cas12 of the present disclosure is fused with a deaminase domain, in combination with a gRNA); addition or deletion of a base, for example, by Prime editing technology (Cas12 of the present disclosure is fused with a reverse transcriptase, in combination with a PegRNA); or by HDR homologous recombination (Cas12 of the present disclosure in combination with a gRNA and a donor template); Cas12 of the present disclosure can also be fused with a DNA methylase or a DNA demethylase, in combination with a gRNA for targeting, to modulate the methylation level of a target nucleic acid.
[0395] In the present disclosure, depending on the context, “modulating expression of a target nucleic acid” can refer to modulation of transcription of a target nucleic acid; non-limiting examples include enhancement or inhibition of transcription of a target nucleic acid by CRISPRa, CRISPRi technology, with the aid of a transcription activation or inhibition domain fused to Cas12.
[0396] In the present disclosure, the letters in an amino acid sequence represent the one-letter abbreviations of amino acids well known in the art, as described in, for example, J. Biol. Chem, 243, p3558 (1968): alanine: Ala-A, arginine: Arg-R, aspartic acid: Asp-D, cysteine: Cys-C, glutamine: Gln-Q, glutamic acid: Glu-E, histidine: His-H, glycine: Gly-G, asparagine: Asn-N, tyrosine: Tyr-Y, proline: Pro-P, serine: Ser-S, methionine: Met-M, lysine: Lys-K, valine: Val-V, isoleucine: Ile-I, phenylalanine: Phe-F, leucine: Leu-L, tryptophan: Trp-W, threonine: Thr-T.
[0397] In the present disclosure, “amino acid difference” refers to the difference of an amino acid residue at a specific position in the amino acid sequence of a protein, including substitution, addition or deletion.
[0398] As is known to those skilled in the art, in a protein or peptide, two adjacent amino acids each lose one OH or H and condense to form a peptide bond, and each amino acid is actually present as an amino acid residue. Thus, in the present disclosure, the terms "amino acid" and "amino acid residue" generally represent the same meaning. In addition, in order to simplify the expression, in the present disclosure, the amino acid residue before substitution is retained before the site of the amino acid residue, and the letter before the site represents the original amino acid residue, and the letter after the site represents the amino acid residue after substitution. For example, S211 represents that the original amino acid residue at the 211 site is S, and when it is replaced by R, it can be represented as S211R.
[0399] Herein, sometimes the symbol "+" is used to connect one amino acid mutation before and after it, which means that the two point mutations exist simultaneously in one mutant; if multiple point mutations are connected by two or more symbols "+", it is intended to mean that the multiple point mutations exist simultaneously.
[0400] As used herein, a mutation can refer to a mutation to any other natural amino acid residue. Alternatively, the mutation is to residue R, H, K or A; alternatively, the mutation is to residue R; alternatively, the mutation is to residue A.
[0401] In the present disclosure, if an amino acid is substituted, it means that it is replaced by another amino acid residue different from the original amino acid residue. If the original amino acid belongs to a positively charged amino acid, it is replaced by a positively charged amino acid, which means that it is replaced by another positively charged amino acid residue different from the original amino acid residue. For example, the original amino acid residue is R, which is replaced by a positively charged amino acid, which means that it is replaced by H or K.
[0402] In the present disclosure, when referring to an RNA sequence, "T" in the sequence can be used interchangeably with "U". When referring to a "guide sequence", "T" in the sequence can be used interchangeably with "U". When referring to a "direct repeat sequence", "T" in the sequence can be used interchangeably with "U". When referring to a "tracrRNA sequence", "T" in the sequence can be used interchangeably with "U".
[0403] Identity of sequences
[0404] As used herein, the terms "identity" or "percent identity" are used in reference to the matching of sequences between two polypeptides or between two nucleic acids. When a position in both of the sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., each position in both of the DNA molecules is occupied by adenine, or each position in both of the polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent sequence identity" between two sequences is a function of the number of matching positions shared by the sequences, divided by the number of positions compared x 100%. For example, if 6 of 10 positions in two sequences are matched then the sequences have 60% sequence identity. Generally, the comparison is made over the full length of the two sequences being compared. Such a comparison can be made using publicly available and commercially available alignment algorithms and programs, such as, but not limited to, CLUSTAL Omega, MAFFT, Probcons, T-Coffee, Probalign, BLAST, as can be reasonably selected by one of ordinary skill in the art. One of ordinary skill in the art can determine appropriate parameters for aligning sequences, including, for example, any algorithm needed to achieve optimal alignment or best fit over the full length of the sequences being compared, and any algorithm needed to achieve optimal alignment or best fit over a portion of the sequences being compared.
[0405] CRISPR-Cas12 system
[0406] As used herein, the terms "Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-CRISPR-Associated (Cas) (CRISPR-Cas) system," "CRISPR-Cas12 system," and "CRISPR system" are used interchangeably. The CRISPR-Cas12 system generally comprises a Cas12 protein sequence or a nucleic acid encoding the same, and a guide polynucleotide or a nucleic acid encoding the same.
[0407] The Zhang Feng group discovered Cas12a in 2015, which is classified as V-type in Class II CRISPR-Cas system. After a detailed study of V-A subtype (Cas12a), the Zhang Feng group reported Cas12b (C2C1) in 2015. In 2017, Burstein et al. reported Cas12e (CasX) nuclease. In 2019, Winston X. Yan et al. reported the newly discovered V-type Cas effector proteins Cas12c, Cas12h, Cas12i, and Cas12g through bioinformatics analysis.
[0408] In some embodiments of the disclosure, a Cas12 protein described herein refers to a protein whose amino acid sequence comprises or is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to any one of SEQ ID NOs: 1-35. When the CRISPR-Cas12 system comprises a fusion protein or a conjugate comprising the Cas12 protein and a functional domain, the percent sequence identity between the Cas12 portion of the fusion protein or conjugate and the reference sequence is calculated.
[0409] In the present disclosure, a CRISPR-Cas12 system comprises a Cas12 protein having at least 50% sequence identity compared to any one of SEQ ID NOs: 1-35, or a nucleic acid encoding the Cas12 protein; and a guide polynucleotide or a nucleic acid encoding the guide polynucleotide; the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence engineered to hybridize to a target nucleic acid, the guide polynucleotide is capable of forming a complex with the Cas12 protein and directing the complex to sequence-specific binding to the target nucleic acid.
[0410] Guide polynucleotide
[0411] As used herein, the term "guide polynucleotide" is used to refer to a molecule in a CRISPR-Cas system that forms a complex with a Cas protein and directs the complex to a target sequence. Typically, a guide polynucleotide comprises a backbone sequence linked to a guide sequence, which can hybridize to a target sequence. The backbone sequence typically comprises a direct repeat sequence, and sometimes a tracrRNA sequence. In some embodiments of the disclosure, the guide polynucleotide does not comprise a tracrRNA sequence. In some embodiments of the disclosure, the guide polynucleotide comprises a tracrRNA sequence.
[0412] In some embodiments of the disclosure, the guide polynucleotide of the CRISPR-Cas12 system is a guide RNA. In some embodiments of the disclosure, the guide polynucleotide is a chemically modified guide polynucleotide. In some embodiments of the disclosure, the guide polynucleotide comprises at least one chemically modified nucleotide.
[0413] In some embodiments of the disclosure, the guide polynucleotide comprises at least one guide sequence (also referred to as spacer sequence) linked to at least one direct repeat sequence (DR). In some embodiments of the disclosure, the guide sequence is located at the 3' end of the direct repeat sequence. In some embodiments of the disclosure, the guide sequence is located at the 5' end of the direct repeat sequence.
[0414] In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence.
[0415] In some embodiments of the disclosure, the tracrRNA sequence is located at the 5' or 3' end of the direct repeat sequence. In some embodiments of the disclosure, the tracrRNA sequence is located at the 5' end of the direct repeat sequence. In some embodiments of the disclosure, the tracrRNA sequence is located at the 3' end of the direct repeat sequence.
[0416] In some embodiments of the disclosure, the nucleotide sequence of the guide polynucleotide comprises, in order from 5' to 3', a direct repeat sequence, a guide sequence.
[0417] In some embodiments of the disclosure, the nucleotide sequence of the guide polynucleotide comprises, in order from 5' to 3', a guide sequence, a direct repeat sequence.
[0418] In some embodiments of the disclosure, the nucleotide sequence of the guide polynucleotide comprises, in order from 5' to 3', a tracrRNA, a direct repeat sequence, a guide sequence.
[0419] In some embodiments of the disclosure, the nucleotide sequence of the guide polynucleotide comprises, in order from 5' to 3', a tracrRNA, a linker sequence, a direct repeat sequence, a guide sequence.
[0420] In some embodiments of the disclosure, the nucleotide sequence of the guide polynucleotide comprises, in order from 5' to 3', a tracrRNA, a loop sequence, a direct repeat sequence, a guide sequence.
[0421] In some embodiments of the disclosure, the structure of the guide polynucleotide is 5'-tracrRNA-loop-direct repeat sequence-guide sequence-3'.
[0422] In some embodiments of the disclosure, the tracrRNA and the direct repeat sequence of the guide polynucleotide are linked by a nucleotide sequence.
[0423] In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a nucleotide sequence consisting of 4 nucleotides. In some embodiments of the disclosure, the tracrRNA sequence is linked to the direct repeat sequence by a 5'-GAAA-3' sequence.
[0424] In some embodiments of the disclosure, the guide sequence has sufficient complementarity to the target nucleic acid sequence to hybridize to the target nucleic acid and direct sequence-specific binding of the CRISPR-Cas12 complex to the target nucleic acid. In some embodiments of the disclosure, the guide sequence has 100% complementarity to the target nucleic acid, but the guide sequence can have less than 100% complementarity to the target nucleic acid, for example at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% complementarity.
[0425] In some embodiments of the disclosure, the guide sequence is engineered to hybridize to the target nucleic acid with no more than two nucleotide mismatches. In some embodiments of the disclosure, the guide sequence is engineered to hybridize to the target nucleic acid with no more than one nucleotide mismatch. In some embodiments of the disclosure, the guide sequence is engineered to hybridize to the target nucleic acid with or without mismatches.
[0426] In some embodiments of the disclosure, the CRISPR-Cas12 system comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different guide polynucleotides. In some embodiments of the disclosure, the guide polynucleotides target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different target nucleic acid molecules, or target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different regions of one or more target nucleic acid molecules.
[0427] In some embodiments of the disclosure, the guide polynucleotide comprises a constant direct repeat sequence upstream of the variable guide sequence. In some embodiments of the disclosure, a plurality of guide polynucleotides are part of an array (which can be part of a vector, such as a viral vector or a plasmid). For example, a guide array comprising the sequence DR-spacer-DR-spacer-DR-spacer-...-DR-spacer can comprise a plurality of unique unprocessed guide polynucleotides (one per DR-spacer or spacer-DR sequence). Upon introduction into a cell or cell-free system, the array is processed by the Cas12 protein into a number of individual mature guide polynucleotides. This allows multiplexing, for example, to deliver multiple guide polynucleotides to a cell or system to target multiple target nucleic acids or multiple regions within a single target nucleic acid.
[0428] The ability of a guide polynucleotide to direct sequence-specific binding of a CRISPR complex to a target nucleic acid can be assessed by any suitable assay. For example, components of a CRISPR system sufficient to form a complex (CRISPR complex), including a guide polynucleotide to be tested, can be provided to a host cell having a corresponding target nucleic acid molecule, for example by transfection of a vector encoding components of the CRISPR complex, and then the preferential cleavage within the target sequence is assessed. Similarly, cleavage of a target nucleic acid sequence can be assessed in vitro by providing a target nucleic acid, components of a CRISPR complex, including a guide polynucleotide to be tested and a control guide polynucleotide different from the guide polynucleotide to be tested, and comparing the ability of the guide polynucleotide to be tested and the control guide polynucleotide to bind to the target nucleic acid or the rate of cleavage of the target nucleic acid. The ability of a CRISPR complex to cleave a target nucleic acid or a target nucleic acid can also be assessed by the assays described above.
[0429] Cas12 mutants
[0430] As described herein, when reference is made to a "corresponding position to the sequence set forth in SEQ ID NO:XX" or using similar language, the corresponding position can be determined by amino acid sequence alignment (XX is a positive integer). Typically, the comparison is made when two sequences are aligned to yield maximum sequence identity. Such alignment can be made by using publicly available and commercially available alignment algorithms and programs, such as, but not limited to, Clustal Omega, MAFFT, Probcons, T-Coffee, Probalign, BLAST, which can be reasonably selected for use by one of ordinary skill in the art. One of skill in the art can determine appropriate parameters for aligning sequences, for example, including any algorithm required to achieve an optimal alignment or best fit over the entire length of the sequences being compared, and any algorithm required to achieve an optimal alignment or best fit over a portion of the sequences being compared.
[0431] In some embodiments of the disclosure, the Cas12 protein provided herein comprises one or more mutations, e.g., a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof, compared to a Cas12 protein set forth in any of SEQ ID NOs: 1-35. In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid changes (e.g., insertions, deletions, or substitutions) compared to a Cas12 protein set forth in any of SEQ ID NOs: 1-35, but retains the ability to bind a target nucleic acid molecule that is complementary to a guide sequence of a guide polynucleotide, and / or retains the ability to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.In some examples, the Cas12 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid changes (e.g., insertions, deletions, or substitutions) compared to the Cas12 protein set forth in any of SEQ ID NOs: 1-35, but retains the ability to bind a target nucleic acid molecule that is complementary to a guide sequence of a guide polynucleotide.
[0432] One type of modification or mutation includes substitution of an amino acid residue with an amino acid having similar biochemical properties, i.e., a conservative substitution. In general, conservative substitutions will have little or no effect on the activity of the resulting protein or peptide. For example, a conservative substitution is a substitution of an amino acid in a Cas12 protein that does not substantially affect the binding of the Cas12 protein to a target nucleic acid molecule that is complementary to a guide sequence of a gRNA molecule, and / or the process of processing a guide array RNA transcript into a gRNA molecule.
[0433] More substantial changes can be made by using less conservative substitutions, e.g., selecting residues that differ more significantly in their effect on maintaining (a) the structure of the polypeptide backbone in the area of the substitution, e.g., as a helical conformation or a folded conformation; (b) the charge or hydrophobicity of the region at the site of the substitution; or (c) the bulk of the side chain. Often, the substitutions that have minimal effects on the function of the polypeptide are those in which (a) the residue is substituted by one of its structurally similar or conservative counterparts; (b) a hydrophilic residue (e.g., serine or threonine) is substituted for a hydrophilic residue (e.g., leucine, isoleucine, phenylalanine, valine, or alanine); (c) a cysteine or proline is substituted for any other residue; (d) a residue with a positive side chain (e.g., lysine, arginine, or histidine) is substituted for a residue with a negative side chain (e.g., glutamic acid or aspartic acid); or (e) a residue with a bulky side chain (e.g., phenylalanine) is substituted for a residue with no side chain (e.g., glycine).
[0434] Cas12 active fragment
[0435] In the present disclosure, the Cas12 protein can comprise only a WED-I domain, a Helical-I1 domain, a PI domain, a Helical-I2 domain, a Helical-II domain, a WED-II domain, a Ruvc-I domain, a Helical-III domain, a BH domain, a Ruvc-II domain, a Nuc domain, and / or a Ruvc-III domain.
[0436] The Cas12 protein of the present disclosure, in addition to comprising the domains, can comprise other domains of the prior art Cas12 proteins, which are collectively combined into the complete structure of the Cas12 protein to achieve the functions of the Cas12 protein of the present disclosure, including but not limited to retaining the ability of the Cas12 protein to form a complex with a gRNA, retaining the ability of the Cas12 protein to form a complex with a gRNA and target a target nucleic acid, retaining the ability of the complex of the Cas12 protein with a gRNA to target the expression of a target nucleic acid, retaining the ability of the complex of the Cas12 protein with a gRNA to target the cleavage of a single strand or a double strand of a target nucleic acid, retaining the ability of the Cas12 protein to bind a target nucleic acid molecule that is complementary to a guide sequence of a guide polynucleotide, and / or retaining the ability of the Cas12 protein to process an RNA transcript comprising a guide sequence into a guide polynucleotide molecule.
[0437] C12-334 protein comprises the following domain fragments: aa1-24 WED, aa25-109 Helical 1, aa110-182 PI, aa183-340 Helical 1, aa341-447 WED, aa448-522 RuvC, aa523-644 Helical 2, aa645-720 RuvC, aa721-756 Nuc, aa757-769 RuvC, aa770-891 Nuc.
[0438] wherein “aa” refers to amino acid, for example, “aa1-24 WED” refers to the first to the 24th amino acid of C12-334 protein is WED domain.
[0439] In some embodiments of the disclosure, the Cas12 protein, Cas12 inactive variant, conjugate or fusion protein comprises:
[0440] (a) any of the domain fragments set forth in aa1-24 WED, aa25-109 Helical 1, aa110-182 PI, aa183-340 Helical 1, aa341-447 WED, aa448-522 RuvC, aa523-644 Helical 2, aa645-720 RuvC, aa721-756 Nuc, aa757-769 RuvC and aa770-891 Nuc of C12-334 protein; or,
[0441] (b) an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any of the fragments in (a). Non-limiting examples include a new protein obtained by fusing the PI domain of C12-334 protein or a similar sequence fragment to other proteins (such as Cas12 protein, Cas9 protein or IscB protein); the new protein has the function of recognizing PAM with WYR.
[0442] Cas12 inactive variant
[0443] By means of point mutation, the RuvC domain of Cas12 is inactivated, the Cas12 protein will lose the endonuclease activity, and the dCas12 (dead Cas12) formed can only bind to the target gene under the mediation of the guide polynucleotide, and does not have the function of cutting DNA.
[0444] The RuvC domain of Cas12 can also be partially inactivated by means of point mutations to form a Cas12 nickase (nCas12) that, under the mediation of a guide polynucleotide, binds to a target gene and cleaves one of the single strands in the double-stranded nucleic acid without cleaving the other single strand.
[0445] Accordingly, dCas12 or nCas12 can be fused or conjugated with other domains, including but not limited to deaminase domains, transcription activation domains, transcription repression domains, methylation domains, demethylation domains, histone acetylation domains, histone deacetylation domains, to exert corresponding functions by virtue of the other domains after being guided by a guide polynucleotide to a target sequence of a target nucleic acid; for example, to effect C→T conversion of a target nucleic acid by deamination of cytosine bases, A→G conversion of a target nucleic acid by deamination of adenine bases, transcription repression of a target nucleic acid by a transcription repression domain KRAB, transcription promotion of a target nucleic acid by a transcription activation domain VP64, DNA methylation or expression repression by a DNMT3A / 3B / 3L domain.
[0446] Functional domain
[0447] In some embodiments of the disclosure, the Cas12 protein or Cas12 inactivated variant is covalently linked or fused to a homologous or heterologous functional domain.
[0448] In some embodiments of the disclosure, the functional domain has an enzymatic activity that modifies a target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deaminase activity, dismutase activity, alkylating activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristoylation activity, and / or demyristoylation activity.
[0449] In some embodiments of the disclosure, the functional domain is optionally one or more of: a nuclease (e.g., Fokl), a methyltransferase, a demethylase, a DNA repair enzyme, a DNA damaging enzyme, a deaminase, a dismutase, an alkylating enzyme, a depurinase, an oxidase, a pyrimidine dimer forming enzyme, an integrase, a transposase, a recombinase, a polymerase, a ligase, a helicase, a photolyase, a glycosylase, a deglycosylase, an acetyltransferase, a deacetylase, a kinase, a phosphatase, a ubiquitin ligase, a deubiquitinating enzyme, an adenylylating enzyme, a deadenylating enzyme, a SUMOylating enzyme, a deSUMOylating enzyme, a myristoylating enzyme, and / or a demyristoylating enzyme.
[0450] In some embodiments of the disclosure, the functional domain is optionally one, two, three, four or more of: a subcellular localization signal, a DNA binding domain, a protease domain, a transcription activation domain, a transcription inhibition domain, a nuclease domain, a deaminase domain, a uracil DNA glycosylase domain (UDG), a uracil DNA glycosylase inhibitor domain (UGI), a methylase, a demethylase, a transcription release factor, a histone acetylase domain, a histone deacetylase domain, a DNA ligase, an epitope tag, and / or a reporter domain.
[0451] In some embodiments of the disclosure, the deaminase is an adenine deaminase or a cytosine deaminase.
[0452] In some embodiments of the disclosure, the deaminase domain is optionally selected from: APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, activation-induced cytidine deaminase (AID), CDA from Lampetra, mutants of adenosine deaminases engineered to act on DNA (TadA).
[0453] In some embodiments of the disclosure, the functional domain has epigenome modification activity. In some embodiments of the disclosure, the functional domain has epigenetic modification activity. In some embodiments of the disclosure, the functional domain has epigenetic modification activity. Optionally, the epigenome modification, epigenetic modification, or epigenetic modification includes, but is not limited to, DNA methylation, RNA methylation, RNA interference, nucleosome positioning, chromatin conformational change, chromatin remodeling, histone modification, modification of long non-coding RNA sequences.
[0454] In some embodiments of the disclosure, the functional domain is an epigenome modification functional domain. In some embodiments of the disclosure, the functional domain is an epigenetic modification functional domain. In some embodiments of the disclosure, the functional domain is an epigenetic modification functional domain.
[0455] In some embodiments of the disclosure, the functional domain has single base editing activity. In some embodiments of the disclosure, the functional domain is a single base editing functional domain. In some embodiments of the disclosure, the functional domain is a base conversion enzyme. In some embodiments of the disclosure, the functional domain or base conversion enzyme is a base deaminase. In some embodiments of the disclosure, the functional domain or base conversion enzyme is an adenine deaminase or a cytosine deaminase. In some embodiments of the disclosure, the functional domain is an adenine deaminase or a cytosine deaminase. In some embodiments of the disclosure, the base conversion enzyme is an adenine deaminase or a cytosine deaminase.
[0456] In some embodiments of the disclosure, the transcriptional activation domain is optionally selected from the group consisting of: P65, VPR, VP16, VP64, VTR1, VTR2, VTR3, p65, MyoD1, HSF1, RTA, SET7 / 9, and histone acetyltransferases. In some embodiments of the disclosure, the transcriptional activation domain is optionally selected from the group consisting of: the sequence ETFSDLWKL (SEQ ID NO: 230) from p53 TAD1, the sequence DDIEQWFTE (SEQ ID NO: 231) from p53 TAD2, the sequence SDIMDFVLK (SEQ ID NO: 232) from MLL, the sequence DLLDFSMMF (SEQ ID NO: 233) from E2A, the sequence ETLDFSLVT (SEQ ID NO: 234) from Rtg3, the sequence RKILNDLSS (SEQ ID NO: 235) from CREB, the sequence EAILAELKK (SEQ ID NO: 236) from CREBaB6, the sequence DDVVQYLNS (SEQ ID NO: 237) from Gli3, the sequence DDVYNYLFD (SEQ ID NO: 238) from Gal4, the sequence DLFDYDFLV (SEQ ID NO: 239) from Oaf1, the sequence DFFDYDLLF (SEQ ID NO: 240) from Pip2, the sequence EDLYSILWS (SEQ ID NO: 241) from Pdr1, the sequence TDLYHTLWN (SEQ ID NO: 242) from Pdr3.
[0457] In some embodiments of the disclosure, the transcription repression domain is optionally selected from KOX1 (Krüppel-associated box protein 1), KAP-1 (KRAB-associated protein 1), MAD (MAX dimerization protein), FKHR (Forkhead box protein O1), EGR-1 (Early growth response protein 1), ERD (Estrogen response element-binding domain), SID (mSin3 interaction domain), a concatemer of SID (e.g., SID4X, i.e., a four-fold repeat of mSin3 interaction domain), TIEG (TGF-beta-inducible early growth response protein), v-ERB-A (viral erythroblastosis oncogene A), MBD2 (Methyl-CpG-binding domain protein 2), MBD3 (Methyl-CpG-binding domain protein 3), TRa (Thyroid hormone receptor alpha), Histone methyltransferase, Histone deacetylase (HDAC), Nuclear hormone receptor (e.g., Estrogen receptor or Thyroid hormone receptor), DNA methyltransferase family members (e.g., DNMT1, DNMT3A, and DNMT3B), KRAB domain of MeCP2 (Methyl-CpG-binding protein 2), ROM2 (Ral guanine nucleotide dissociation stimulator-like 2), and AtHD2A (Arabidopsis thaliana Histone deacetylase 2A).
[0458] In some embodiments of the disclosure, the transcription repression domain is a KRAB (Krüppel-associated box) domain from a KOX1 protein.
[0459] In some embodiments of the disclosure, the nuclease domain is optionally selected from Fokl (restriction nuclease derived from Flavobacterium okeanokoites), a polypeptide having single-stranded DNA (ssDNA) cleavage activity, or a polypeptide having double-stranded DNA (dsDNA) cleavage activity.
[0460] In some embodiments of the disclosure, the methylase domain is selected from DNA methyltransferases, including but not limited to DNMT1 (DNA methyltransferase 1), DNMT3A (DNA methyltransferase 3 alpha), and DNMT3B (DNA methyltransferase 3 beta).
[0461] In some embodiments of the disclosure, the demethylase is selected from TET1CD (Ten-Eleven Translocation methylcytosine dioxygenase 1 catalytic domain), TET1 (Ten-Eleven Translocation methylcytosine dioxygenase 1), ROS1 (Repressor of Silencing 1), DME (Demeter), DML2 (Demeter-like protein 2), and DML3 (Demeter-like protein 3).
[0462] Methylation and demethylation are recognized in the art as important ways of epigenetic gene regulation.
[0463] In some embodiments of the disclosure, the homologous or heterologous functional domain is a sequence tag useful for solubilization, purification, or detection of the fusion protein or conjugate. Suitable protein tag sequences are provided herein, including but not limited to a Biotin Carboxyl Carrier Protein (BCCP tag), a myc tag (a small epitope tag derived from the c-Myc protein), a Calmodulin tag, a FLAG tag (DYKDDDDK sequence tag), a Hemagglutinin tag (HA tag), a Polyhistidine tag (also known as a His tag), a Maltose-binding protein tag (MBP tag), a nus tag (N-utilization substance protein A tag), a Glutathione S-transferase tag (GST tag), a Green Fluorescent Protein tag (GFP tag), a Thioredoxin tag, an S-tag (a short peptide tag derived from RNase A), a Softag tag (e.g., Softag 1 and Softag 3), a Strep tag (Streptavidin-binding peptide tag), a Biotin ligase tag, a FlAsH tag (Fluorescein Arsenical Hairpin binder tag), a V5 tag (an epitope tag derived from Simian Virus 5), and a SBP tag (Streptavidin-binding peptide tag). Additional suitable sequences will be apparent to those skilled in the art.
[0464] Subcellular localization signal
[0465] In some embodiments of the disclosure, the Cas12 protein is fused to at least one homologous or heterologous subcellular localization signal. In some embodiments of the disclosure, the Cas12 protein is fused to at least one homologous or heterologous subcellular localization signal. Exemplary subcellular localization signals include organelle localization signals, such as a nuclear localization signal (NLS), a nuclear export signal (NES), or a mitochondrial localization signal.
[0466] Non-limiting examples of NLSs include NLS sequences derived from: the NLS of the SV40 virus large T antigen, which has the amino acid sequence PKKKRKV (SEQ ID NO:243); the NLS from nucleoplasmin (e.g., the sequence KRPAATKKAGQAKKKK (SEQ ID NO:244)); the c-myc NLS, which has the amino acid sequence PAAKRVKLD (SEQ ID NO:245) or RQRRNELKRSP (SEQ ID NO:246); the hRNPA1 M9 NLS, which has the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:247); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:248) from the IBB domain; the sequences VSRKRPRP (SEQ ID NO:249) and PPKKARED (SEQ ID NO:250) of the T protein of myotubes; the sequence PQPKKKPL (SEQ ID NO:251) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO:252) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO:253) and PKQKKRK (SEQ ID NO:254) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO:255) of hepatitis delta antigen; the sequence REKKKFLKRR (SEQ ID NO:256) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:257) of human PARP enzyme; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO:258) of steroid hormone receptors. In some embodiments, the nuclear localization sequence is of sufficient strength to drive the fusion protein or conjugate of the disclosure to accumulate in a detectable amount in the nucleus of a eukaryotic cell. In general, the strength of nuclear localization activity can arise from the number of NLSs, the particular NLS(s) used, or a combination of these factors. Detection of nuclear accumulation can be performed by any suitable technique. For example, a detectable label can be fused to the Cas protein, such that the location within the cell can be visualized, such as in combination with a means of detecting the location of the nucleus (e.g., a dye specific for the nucleus, such as DAPI). The nucleus can also be isolated from the cell, and its contents can then be analyzed by any suitable method for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay.Accumulation in the nucleus can also be determined indirectly, such as by assaying the effect of nucleic acid targeting complex formation (e.g., assaying DNA or RNA cleavage or mutation at a target sequence, or assaying altered gene expression activity due to the effect of DNA or RNA targeting complex formation and / or DNA or RNA targeting Cas protein activity), in comparison to a control that is not exposed to a nucleic acid targeting Cas protein or nucleic acid targeting complex, or is exposed to a nucleic acid targeting Cas protein that lacks one or more NLS.
[0467] Nucleic acid / polynucleotide
[0468] Nucleic acid or polynucleotide refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in single- or double-stranded form. The term “nucleic acid” includes, but is not limited to, genes, cDNA, and mRNA. In one embodiment, a nucleic acid molecule is synthetic (e.g., chemically synthesized) or recombinant. Unless specifically limited, the term encompasses nucleic acids containing analogues or derivatives of natural nucleotides that have similar binding properties and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated.
[0469] Vector system
[0470] Another aspect of the disclosure relates to a vector system comprising a CRISPR-Cas12 system described herein, the vector system comprising one or more recombinant vectors comprising a polynucleotide sequence encoding the Cas12 protein and a polynucleotide sequence encoding the guide polynucleotide.
[0471] In some embodiments of the disclosure, the vector system comprises at least one plasmid or viral recombinant vector (e.g., a retrovirus, a lentivirus, an adenovirus, an adeno-associated virus, or a herpes simplex virus). In some embodiments of the disclosure, the polynucleotide sequence encoding the Cas12 protein and the polynucleotide sequence encoding the guide polynucleotide are located on the same recombinant vector. In some embodiments of the disclosure, the polynucleotide sequence encoding the Cas12 protein and the polynucleotide sequence encoding the guide polynucleotide are located on multiple recombinant vectors.
[0472] In some embodiments of the disclosure, the polynucleotide sequence encoding a Cas12 protein and / or the polynucleotide sequence encoding a guide polynucleotide is operably linked to a regulatory sequence (also referred to as a regulatory element). The regulatory elements include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Regulatory elements include those that result in constitutive expression of a nucleotide sequence in many types of host cells, as well as those that result in expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression directly in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes), for example. Regulatory elements can also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which can or can not also be tissue- or cell type-specific. In some embodiments, the regulatory element is an enhancer element, such as the WPRE, the CMV enhancer, the R-U5 segment in the LTR of HTLV-1, the SV40 enhancer, or the intron sequence between exons 2 and 3 of rabbit beta-globin.
[0473] In some embodiments of the disclosure, the recombinant vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., the Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the beta-actin promoter, the phosphoglycerol kinase (PGK) promoter, or the EF1a promoter), or a pol III promoter and a pol II promoter.
[0474] In some embodiments of the disclosure, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1a, CAG, and beta-actin promoters. In some embodiments of the disclosure, the promoter is an inducible promoter, which is regulated by external signals or molecules (e.g., transcription factors).
[0475] In some embodiments of the disclosure, the promoter is a tissue-specific promoter that is used to drive tissue-specific expression of the Cas12 protein. Suitable muscle-specific promoters include, but are not limited to, CK8, MHCK7, Myoglobin promoter (Mb), Desmin promoter, muscle creatine kinase promoter (MCK) and variants thereof, and SPc5-12 synthetic promoter. Suitable immune cell-specific promoters include, but are not limited to, B29 promoter (B cells), CD14 promoter (monocytes), CD43 promoter (leukocytes and platelets), CD68 (macrophages), and SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include, but are not limited to, CD43 promoter (leukocytes and platelets), CD45 promoter (hematopoietic cells), INF-beta (hematopoietic cells), WASP promoter (hematopoietic cells), SV40 / CD43 promoter (leukocytes and platelets), and SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include, but are not limited to, elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, Fit-1 promoter and ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, GFAP promoter (astrocytes), SYN1 promoter (neurons), and NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, Nphsl promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, alpha-MHC.
[0476] AAV vector
[0477] Another aspect of the disclosure relates to an adeno-associated virus (AAV) vector comprising the CRISPR-Cas12 system described herein, wherein the adeno-associated virus (AAV) vector comprises DNA encoding the Cas12 protein and / or guide polynucleotide described herein.
[0478] In some embodiments of the disclosure, the AAV vector comprises a DNA sequence encoding the Cas12 protein described herein. In some embodiments of the disclosure, the AAV vector comprises a DNA sequence encoding the fusion protein described herein. In some embodiments of the disclosure, the AAV vector comprises a DNA sequence encoding the guide polynucleotide described herein.
[0479] Delivery of CRISPR-Cas systems by AAV vectors is described in Maeder et al., Nature Medicine 25:229-233 (2019), which is incorporated by reference herein in its entirety. In some embodiments of the disclosure, the AAV vector comprises a ssDNA genome comprising a coding sequence for an RNA-guided nuclease and a guide RNA flanked by ITRs.
[0480] In some embodiments of the disclosure, the Cas12 protein, guide polynucleotide, Cas12 inactivated variant, fusion protein or conjugate comprising the Cas12 protein, isolated nucleic acid, and / or CRISPR-Cas12 system described herein is packaged in an AAV vector, e.g., packaged into an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74 capsid.
[0481] In some embodiments of the disclosure, the Cas12 protein, guide polynucleotide, Cas12 inactivated variant, fusion protein or conjugate comprising the Cas12 protein, isolated nucleic acid, and / or CRISPR-Cas12 system described herein is packaged into an AAV2, AAV5, AAV6, AAV8, AAV9, or AAV PHP.eB capsid.
[0482] In some embodiments of the disclosure, the AAV vector described herein is optionally selected from the group consisting of: AAV2 / 2, AAV2 / 3, AAV2 / 4, AAV2 / 5, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2 / 10, AAV2 / 11, AAV2 / 12, AAV2 / 13, AAV2 / PHP.B, AAV2 / PHP.B2, AAV2 / PHP.B3, AAV2 / PHP.A, AAV2 / PHP.eB, AAV2 / PHP.eS, AAV2 / 2.7m8, AAV2 / 8.7m8, AAV2 / ShH10, AAV2 / rh10, and AAV2 / rh74.
[0483] In some embodiments of the disclosure, the AAV vector described herein is optionally selected from the group consisting of: AAV2 / 2, AAV2 / 5, AAV2 / 6, AAV2 / 8, AAV2 / 9, AAV2 / PHP.eB.
[0484] In some embodiments of the disclosure, the CRISPR-Cas12 system described herein is packaged in an AAV vector comprising an engineered capsid with tissue tropism, for example, an engineered ocular tissue-tropic capsid.
[0485] Lipid nanoparticle
[0486] Another aspect of the disclosure relates to a lipid nanoparticle (LNP) comprising the CRISPR-Cas12 system described herein, wherein the LNP comprises the guide polynucleotide described herein and the mRNA encoding the Cas12 protein described herein.
[0487] LNP delivery of CRISPR-Cas systems is described in Gillmore et al., N. Engl. J. Med., 385:493-502 (2021), the lipid nanoparticle (LNP) consists of 4 lipids including proprietary ionizable lipid LP000001; DSPC; cholesterol and DMG-PEG2k, the LNP suspension is formulated in an aqueous buffer of Tris, NaCl and sucrose, pH 7.4, which is incorporated herein by reference in its entirety. In some embodiments of the disclosure, in addition to the RNA payload (Cas12 mRNA and guide polynucleotide), the lipid nanoparticle (LNP) comprises four components: a cationic or ionizable lipid, cholesterol, a helper lipid, and a PEG-lipid. In some embodiments of the disclosure, the cationic or ionizable lipid comprises cKK-E12, C12-200, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA, FTT5, Moderna SM-102, and Intellia LP01. In some embodiments of the disclosure, the PEG-lipid comprises PEG-2000-C-DMG, PEG-2000-DMG, or ALC-0159. In some embodiments of the disclosure, the helper lipid comprises DSPC. Components of LNPs are described in Paunovska et al., Nature Reviews Genetics 23:265-280 (2022), which is incorporated by reference herein in its entirety, FDA-approved LNPs contain variants of the four basic ingredients: a cationic or ionizable lipid, cholesterol, a helper lipid, and a polyethylene glycol (PEG) lipid.
[0488] Lentiviral vector
[0489] Another aspect of the disclosure relates to a lentiviral vector comprising a CRISPR-Cas12 system described herein, wherein the lentiviral vector comprises a guide polynucleotide described herein and an mRNA encoding a Cas12 protein described herein. In some embodiments of the disclosure, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein, such as VSV-G. In some embodiments of the disclosure, the mRNA encoding the Cas12 protein is linked to an aptamer sequence.
[0490] RNP complex
[0491] Another aspect of the disclosure relates to a ribonucleoprotein complex comprising a CRISPR-Cas12 system described herein, wherein the ribonucleoprotein complex is formed from a guide polynucleotide described herein and a Cas12 protein. In some embodiments of the disclosure, the ribonucleoprotein complex is delivered to a eukaryotic cell, a mammalian cell, or a human cell by microinjection or electroporation. In some embodiments of the disclosure, the ribonucleoprotein complex is packaged in a virus-like particle and delivered in vivo to a mammal or a human subject.
[0492] Virus-like particle
[0493] Another aspect of the disclosure relates to a virus-like particle (VLP) comprising a CRISPR-Cas12 system described herein, wherein the virus-like particle comprises a guide polynucleotide described herein and a Cas12 protein or a ribonucleoprotein complex consisting of the guide polynucleotide and the Cas12 protein.
[0494] Banskota et al. Cell 185(2):250-265 (2022) report the development and application of DNA-free virus-like particles (eVLPs) that efficiently package and deliver base editors or Cas9 ribonucleoproteins; Mangeot et al., Nature Communications 10(1):1-15 (2019) use engineered mouse leukemia virus-like particles (Nanoblades) loaded with Cas9-sgRNA ribonucleoproteins to induce highly efficient genome editing in cell lines and primary cells, including human induced pluripotent stem cells, human hematopoietic stem cells, and mouse bone marrow cells; Campbell, et al., Molecular Therapy 27:151-163 (2019) utilize a specialized extracellular vesicle called a “gesicle” to efficiently but transiently deliver Cas9 targeting HIV long terminal repeat (LTR) in ribonucleoprotein form, Gesicles are produced by expressing a vesicular stomatitis virus glycoprotein and packaging protein as their cargo, thus obviating the need for transgene delivery, allowing for more fine-tuned control of Cas9 expression and Mangeot et al. Molecular Therapy, 19(9):1656-1666 (2011) report that overexpression of the spike glycoprotein of vesicular stomatitis oral virus (VSV-G) in human cells induced the release of fusogenic vesicles named gesicles, biochemical and functional studies showed that glial cells incorporated proteins from producer cells and could deliver them to recipient cells, this protein transduction approach allowed for the direct transport of cytoplasmic, nuclear, or surface proteins in target cells. Each of these references describe engineered VLPs, which are incorporated herein by reference in their entirety.
[0495] In some embodiments of the disclosure, the engineered virus-like particle (VLP) is pseudotyped with a homologous or heterologous envelope protein, e.g., VSV-G. In some embodiments of the disclosure, the Cas12 protein is fused to a gag protein (e.g., MLV gag) via a cleavable linker, wherein cleavage of the linker in the target cell exposes an NLS located between the linker and the Cas12 protein. In some embodiments of the disclosure, the fusion protein or conjugate comprises (e.g., from 5’ to 3’) a gag protein (e.g., MLV gag), one or more NES, a cleavable linker, one or more NLS, and a Cas12, as described in Banskota et al. Cell 185(2):250-265 (2022).
[0496] In some embodiments of the disclosure, the Cas12 protein is fused to a first dimerization domain that is capable of dimerizing or heterodimerizing with a second dimerization domain fused to a membrane protein, wherein the presence of a ligand promotes the dimerization and enriches the Cas12 protein or fusion protein or conjugate into a VLP, as described in Campbell, et al., Molecular Therapy 27: 151-163 (2019).
[0497] Cell
[0498] Another aspect of the disclosure relates to a cell comprising a CRISPR-Cas12 system described herein. The cell (e.g., which can be used to generate a cell-free system) can be eukaryotic or prokaryotic. Examples of such cells include, but are not limited to, bacterial, archaeal, plant, fungal, yeast, insect, and mammalian cells, e.g., lactobacillus, lactococcus, bacillus (e.g., B. subtilis), escherichia (e.g., E. coli), clostridium, saccharomyces, or pichia (e.g., S. cerevisiae or P. pastoris), K. lactis, S. typhimurium, Drosophila cells, C. elegans cells, Xenopus cells, SF9 cells, C129 cells, 293 cells, Neurospora, and immortalized mammalian cell lines (e.g., HeLa cells, myeloid cell lines, and lymphoid cell lines).
[0499] In some embodiments of the disclosure, the cell is a prokaryotic cell, e.g., a bacterial cell, e.g., an E. coli cell. In some embodiments of the disclosure, the cell is a eukaryotic cell, e.g., a mammalian cell or a human cell. In some embodiments of the disclosure, the cell is a primary eukaryotic cell, a stem cell, a tumor / cancer cell, a circulating tumor cell (CTC), a blood cell (e.g., a T cell, a B cell, a NK cell, a Treg, etc.), a hematopoietic stem cell, a specialized immune cell (e.g., a tumor infiltrating lymphocyte or a tumor inhibiting lymphocyte), a stromal cell in the tumor microenvironment (e.g., a cancer associated fibroblast, etc.). In some embodiments of the disclosure, the cell is a brain or neuronal cell of the central or peripheral nervous system (e.g., a neuron, an astrocyte, a microglia, a retinal ganglion cell, a rod / cone cell, etc.).
[0500] Target nucleic acid or target DNA
[0501] In some embodiments of the disclosure, the target nucleic acid is a target DNA.
[0502] The CRISPR-Cas12 systems described herein can be used to target one or more target nucleic acid molecules, e.g., target nucleic acid molecules present in a biological sample, an environmental sample (e.g., a soil, air, or water sample), etc.
[0503] In some embodiments of the disclosure, the target nucleic acid is a disease or disorder associated gene. In some embodiments of the disclosure, the target nucleic acid is a disease associated gene. In some embodiments of the disclosure, the disease associated gene is a disease-causing gene that directly causes the disease. In some embodiments of the disclosure, the disease associated gene is an abnormal gene or a gene whose expression is abnormal that directly causes the disease. For example, the gene has an adverse mutation that leads to the occurrence of the disease. For another example, the gene is overexpressed or underexpressed that leads to the occurrence of the disease. In some embodiments of the disclosure, the gene is overexpressed that leads to the occurrence of the disease. In some embodiments of the disclosure, the gene is underexpressed that leads to the occurrence of the disease. In some embodiments of the disclosure, the gene is overexpressed that is associated with the occurrence of the disease. In some embodiments of the disclosure, the gene is underexpressed that is associated with the occurrence of the disease.
[0504] In some embodiments of the disclosure, the disease or disorder is a hematological disease or disorder, an ophthalmic disease or disorder, a neurological disease or disorder, a respiratory disease or disorder, a liver disease or disorder, a metabolic disease or disorder, a cancer, or an infectious disease.
[0505] A list of target nucleic acids / target genes and corresponding diseases or disorders is disclosed in Table 27 of patent application WO2025061113A1, which is incorporated herein by reference. These target nucleic acids / target genes can be edited (including but not limited to modulating the expression level of the target nucleic acid or target gene or other related nucleic acid or gene, changing the nucleotide sequence, or changing the epigenetic modification, by introducing indels, enabling HDR, single base editing, epigenetic editing, or Prime Editing) by the Cas12 protein, guide polynucleotide, Cas12 inactivation variant, fusion protein or conjugate comprising the Cas12 protein, isolated nucleic acid, CRISPR-Cas12 system, vector system, delivery system, cell, pharmaceutical composition, and / or kit of the disclosure, thereby preventing, diagnosing, or treating the corresponding disease or disorder.
[0506] In some embodiments of the disclosure, the target nucleic acid is a reporter gene. Examples of reporter genes include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP).
[0507] Use for treating or preventing a disease
[0508] Another aspect of the disclosure relates to a pharmaceutical composition comprising a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, or a cell as described herein. The pharmaceutical composition can comprise, for example, an AAV vector encoding a Cas12 protein or Cas12 inactivated variant and a guide polynucleotide described herein. The pharmaceutical composition can comprise, for example, a lipid nanoparticle comprising a guide polynucleotide and an mRNA encoding a Cas12 protein described herein. The pharmaceutical composition can comprise, for example, a lentiviral vector comprising a guide polynucleotide and an mRNA encoding a Cas12 protein described herein. The pharmaceutical composition can comprise, for example, a virus-like particle comprising or formed from a guide polynucleotide and a Cas12 protein described herein.
[0509] Another aspect of the disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein in cleaving or editing a target nucleic acid in a mammalian cell.
[0510] Another aspect of the disclosure relates to the use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein in any of the following: cleaving or making a nick in one or more target nucleic acid molecules, activating or upregulating expression of one or more target nucleic acid molecules, activating or repressing transcription of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling, or detecting one or more target nucleic acid molecules, binding to one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules.
[0511] Another aspect of the disclosure relates to use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein for modifying one or more target nucleic acid molecules, the modifying comprising one or more of: nucleic acid base substitution, nucleic acid base deletion, nucleic acid base insertion, cleavage of a target nucleic acid, nucleic acid methylation, and nucleic acid demethylation.
[0512] Another aspect of the disclosure relates to use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein for the diagnosis, treatment or prevention of a disease or disorder associated with a target nucleic acid.
[0513] Another aspect of the disclosure relates to use of a Cas12 protein as described herein, a guide polynucleotide as described herein, a Cas12 inactivated variant as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, a CRISPR-Cas12 system as described herein, a vector system as described herein, a delivery system as described herein, a cell as described herein, a pharmaceutical composition as described herein, or a kit as described herein for the manufacture of a medicament for the diagnosis, treatment or prevention of a disease or disorder associated with a target nucleic acid.
[0514] In some embodiments of the disclosure, the targeted editing by the CRISPR-Cas12 system described herein is directed to a target nucleic acid / target gene selected from the group consisting of the target nucleic acids / target genes described in Table 27 of patent application publication number WO2025061113A1, thereby preventing, diagnosing or treating the corresponding disease or disorder.
[0515] In some embodiments of the disclosure, the pharmaceutical composition is delivered in vivo to a human subject. The pharmaceutical composition can be delivered by any effective route. Exemplary routes of administration include, but are not limited to, intravenous infusion, intravenous injection, intraperitoneal injection, intramuscular injection, intratumoral injection, subcutaneous injection, intradermal injection, intraventricular injection, intravascular injection, intracerebellar injection, intraocular injection, subretinal injection, intravitreal injection, intracameral injection, intratympanic injection, intranasal administration, and inhalation.
[0516] Diagnostic applications
[0517] Another aspect of the disclosure relates to an in vitro composition comprising a CRISPR-Cas12 system described herein and a labeled detector DNA that is not hybridizable to a guide polynucleotide described herein.
[0518] Another aspect of the disclosure relates to use of a CRISPR-Cas12 system described herein in detecting a target nucleic acid in a nucleic acid sample suspected of comprising the target nucleic acid.
[0519] Another aspect of the disclosure relates to use of a CRISPR-Cas12 system described herein in detecting a target nucleic acid in a nucleic acid sample comprising the target nucleic acid.
[0520] In some embodiments of the disclosure, the target nucleic acid detected is a target RNA.
[0521] In some embodiments of the disclosure, the target nucleic acid detected is a target DNA. In some embodiments of the disclosure, the method of detecting a target DNA comprises a Cas12 protein fused to a fluorescent protein or other detectable label and a guide polynucleotide comprising a guide sequence specific to the target DNA. Binding of Cas12 to the target DNA can be visualized by microscopy or other imaging methods.
[0522] In some embodiments of the disclosure, the method of detecting a target nucleic acid in a cell-free system results in the production of a detectable label or enzymatic activity. For example, by using a Cas12 protein, a guide polynucleotide comprising a guide sequence specific to the target nucleic acid, and a detectable label, the target nucleic acid will be recognized by Cas12. Binding of Cas12 to the target nucleic acid triggers its DNase activity, which results in cleavage of the target nucleic acid and the detectable label.
[0523] In some embodiments of the disclosure, the detectable label is DNA linked to a fluorescent probe and a quencher. The intact detectable DNA links the fluorescent probe and quencher, inhibiting fluorescence. Upon cleavage of the detectable DNA by Cas12, the fluorescent probe is released from the quencher and shows fluorescent activity. This method can be used to determine whether a target DNA is present in a lysed cell sample, a lysed tissue sample, a blood sample, a saliva sample, an environmental sample (e.g., a water, soil, or air sample), or other lysed cell or cell-free sample. This method can also be used to detect a pathogen, such as a virus or bacteria, or to diagnose a disease state, such as cancer.
[0524] In some embodiments of the disclosure, detection of a target nucleic acid aids in diagnosing a disease and / or pathological state, or the presence of a viral or bacterial infection.
[0525] The C12-334, C12-335, C12-336, C12-340 and C12-341 proteins of the present disclosure have DNA cleavage activity, the length of the amino acid sequence thereof is from 760 aa to 1080 aa, which is relatively short compared with the commonly used SpCas9 protein (1368 aa) and AsCpf1 protein (1307 aa), is easier to be packaged by a small volume gene therapy vector (such as AAV), and the PAM is different from the NGG PAM of SpCas9, which expands the range of gene editing.
[0526] EMBODIMENTS
[0527] The present disclosure is further illustrated by way of examples below, but the present disclosure is not limited to the scope of the examples described. The experimental methods in the following examples are not specified, and the methods are selected according to the conventional methods and conditions, or according to the instructions of the products.
[0528] Example 1, Screening of Cas12 Proteins
[0529] Through complex bioinformatics methods, a plurality of Cas proteins with the sequences of SEQ ID NO: 1-35 were obtained, as shown in Table 1.
[0530] Table 1. Cas proteins screened
[0531] The DR sequences of the gRNA corresponding to the Cas proteins are shown in Table 2. When there are multiple DR sequences in a single cell in Table 2, the corresponding Cas protein can be optionally combined with one of them for gene editing (i.e., the Cas protein in the table is optionally combined with a guide polynucleotide comprising any corresponding DR sequence for gene editing).
[0532] Table 2. Direct repeat sequences (DR sequences) corresponding to Cas proteins
[0533] The characteristics of the C12-334, C12-335, C12-336, C12-340, C12-341 and other proteins are shown in Table 3.
[0534] Table 3. Characteristics of some Cas proteins
[0535] The enzyme active center of the C12-334 protein comprises the D480, E675, D757 residues.
[0536] The enzyme active center of the C12-335 protein comprises the D619, E858, D1035 residues.
[0537] The enzyme active center of C12-336 protein comprises D380, E562, D647 residues.
[0538] The enzyme active center of C12-340 protein comprises D609, E827, D997 residues.
[0539] The enzyme active center of C12-341 protein comprises D635, E877, D1056 residues.
[0540] Example II, bacterial in vivo editing test, identification of PAM sequence
[0541] In this example, first, a plasmid library containing 7 nt random sequence is constructed, and then expression plasmids of different Cas proteins are constructed. After the expression plasmid is transformed into bacteria to prepare competent cells, the 7 nt random sequence plasmid library is electroporated. If the plasmid in the 7 nt random sequence library can be recognized and targeted by the Cas protein, it will be removed from the library, and the bacteria containing the corresponding sequence plasmid will not be able to grow in a specific antibiotic environment. The specific steps are as follows:
[0542] (1) Construction of 7 nt random sequence plasmid library
[0543] The pLVX-EF1a-BSD vector plasmid (SEQ ID NO: 171) is digested with EcoRV and XhoI, and then agarose gel electrophoresis is performed to recover the linearized vector. The prepared pCDH-CMV-EGFP-reporter3-EF1-Puro plasmid is used as a template, and the primer Puro-PF1 (SEQ ID NO: 172) and Puro-PR1 (SEQ ID NO: 173) are used to amplify the DNA fragment containing the coding sequence of the Puro resistance gene by PCR, and the DNA fragment is inserted into the digested pLVX-EF1a-BSD vector by homologous recombination (NEB, Gibson Master Mix) to construct the recombinant vector pLVX-7NN-Puro library plasmid (SEQ ID NO: 174) containing 7 nt random sequence. The reaction solution is transformed into Stbl3 competent cells, and the LB plate containing ampicillin is coated and incubated at 37°C overnight. After all the colonies are scraped, plasmid extraction is performed to obtain the library plasmid.
[0544] (2) Synthesis of bacterial expression plasmids of different Cas proteins and preparation of competent cells containing the plasmids
[0545] Each bacterial expression plasmid can express different Cas protein (optimized codon) and crRNA (containing DR sequence corresponding to each Cas protein, and guide sequence targeting 7nt random sequence plasmid library) respectively.
[0546] The sequence of each bacterial expression plasmid is as follows:
[0547] P15A-C12-335-HDV plasmid (SEQ ID NO: 177), P15A-C12-336-HDV plasmid (SEQ ID NO: 178), P15A-C12-340-HDV plasmid (SEQ ID NO: 179), P15A-C12-341-HDV plasmid (SEQ ID NO: 180)
[0548] The sequence underlined and not bolded is the coding sequence of Cas protein, the bold italic is the coding sequence of crRNA (gRNA), and the bold italic and underlined is the coding sequence of guide sequence (guide sequence).
[0549] The vector map of P15A-C12-334-HDV is exemplarily shown in FIG. 1.
[0550] (b) Preparation of competent cells containing different bacterial expression plasmids
[0551] Each bacterial expression plasmid of Cas protein was transformed into DH5a competent cells, and single colonies were inoculated into LB medium containing chloramphenicol and cultured at 37°C overnight.
[0552] The electrotransformation competent cells were prepared according to the following steps:
[0553] The culture was inoculated into 100 ml of fresh LB medium containing chloramphenicol at a ratio of 1:100 and cultured at 37°C and 220 rpm;
[0554] The culture was transferred to a 50 ml centrifuge tube and pre-cooled on ice for 30 min when the OD600 reached 0.5;
[0555] Centrifugation at 4000 rpm and 4°C for 10 min, and the cells were resuspended with an equal volume of pre-cooled sterile water;
[0556] The above steps were repeated;
[0557] The cells were resuspended with 1 / 10 volume of pre-cooled sterile water containing 10% glycerol, and aliquoted at 50 uL / tube to obtain competent cells containing each Cas protein bacterial expression plasmid, which was stored at -80°C.
[0558] (c) Plasmid elimination to identify PAM sequence of Cas protein
[0559] pLVX-7NN-Puro library plasmid 100 ng, respectively, electrotransform the competent cells of each bacterial expression plasmid containing Cas protein, and DH5a competent cells, respectively, labeled as Lib1 (electrotransform the competent cells containing Cas protein bacterial expression plasmid) and Lib2 (electrotransform DH5a competent cells);
[0560] After electrotransformation, add 10 ml of LB medium, incubate at 37°C and 220 rpm for 2 h;
[0561] Centrifuge the recovered bacterial solution at 4000 rpm for 2 min, collect the bacterial cells, resuspend in 400 ul of LB, and then spread on LB plates. The bacterial solution electrotransformed with DH5a is spread on LB plates containing ampicillin, and the bacterial solution electrotransformed with each bacterial expression plasmid containing Cas protein is spread on LB plates containing chloramphenicol and ampicillin, and incubated at 37°C overnight;
[0562] Scrape the bacterial cells from the culture plates, respectively, and extract the plasmid DNA by alkaline lysis method;
[0563] Take 100 ng of each of the two extracted plasmid DNAs as a PCR template, and perform PCR amplification using primers SiteSeq-PF1 (SEQ ID NO: 181) and SiteSeqPuro-PR (SEQ ID NO: 182) (as shown in FIG. 2), obtain the fragment, and use the NGS library preparation kit (SynplSeq DNA Library Prep Kit for Illumina) to perform amplicon library preparation. The specific library preparation process is described in the kit instructions. The library after library preparation is subjected to NGS sequencing.
[0564] Compare the NGS sequencing differences of Lib1 and Lib2 cells, analyze, and identify the PAM motif according to the captured sequences, as shown in FIGS. 3-7.
[0565] The results show that C12-334 recognizes the PAM motif of 5'-WYR-3'; C12-335 recognizes the PAM motif of 5'-BMCTTH-3'; C12-336 recognizes the PAM motif of 5'-TTN-3'; C12-340 recognizes the PAM motif of 5'-VNWTV-3', 5'-VNWTC-3' or 5'-VNTTC-3'; C12-341 recognizes the PAM motif of 5'-TTN-3'; wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, and V is A, C or G.
[0566] Example Three, Cleavage activity of C12-334 protein on target nucleic acid in 293T cells
[0567] In this example, an sgRNA targeting the TTR gene in HEK293T cells was first designed, and the guide sequence was SEQ ID NO: 183. At the same time, the C12-334 protein expression framework and the sgRNA expression framework were constructed into the commonly used mammalian expression vector pCDNA3.1(+), obtaining the expression vector plasmid C12-334-TTR-sgRNA05 (SEQ ID NO: 184). After transfecting HEK293T cells, the cleavage activity of C12-334 protein in 293T cells was verified by NGS sequencing.
[0568] TTR gene editing efficiency detection:
[0569] Plating: when the 293T cell line fusion degree is 70-80%, plating is performed, and the number of cells inoculated in a 24-well plate is 5*10^5 cells / well.
[0570] Transfection: 12-14h after plating, transfection was performed, 100ul Opti-MEM, 1.5ul PEI (Yikai Biological, MW 25000), 500ng C12-334-TTR-sgRNA05 plasmid were added to each well of a 24-well plate, mixed, and incubated at room temperature for 20 minutes, then added to 293T cells for cell transfection, and after transfection overnight, fresh culture medium was replaced for continued culture.
[0571] DNA extraction, PCR amplification, NGS library construction: after 72h of culture, the cells were washed with PBS, then 100ul cell lysis solution (Viagen, Lysis Reagent (Cell) ), and a lysate containing genomic DNA was obtained. The region near the target sequence was amplified from the genomic DNA using primers TTR-NGS-PF1 (SEQ ID NO: 185) and TTR-NGS-PR1 (SEQ ID NO: 186), and the PCR product was subjected to NGS library construction, sequencing, and analysis of the sequencing results. The indel efficiency reached 23.58% very unexpectedly, and the results are shown in Figure 8.
[0572] Example Four, Construction of Different Reporter Cell Lines
[0573] A fluorescent reporter system is an important tool for evaluating gene editing efficiency. It mainly reflects the occurrence of editing events by monitoring changes in fluorescent signals. Common fluorescent reporter systems include a reporter system that restores GFP coding to produce a fluorescent signal by introducing indels (insertions / deletions), and a fluorescent system that restores GFP expression through SSA (single-strand annealing). The SSA (Single-Strand Annealing) fluorescent reporter system is designed based on the "single-strand annealing repair pathway" in the DNA double-strand break repair mechanism.
[0574] For the above two cases, the inventors constructed: a pCDH-CMV-GFP-Reporter3-EF1a-Puro cell line (referred to as Reporter3 cell line) based on indel to restore GFP coding and produce a fluorescent signal; and a pCDH-CMV-SSA-GXXFP3 cell line (referred to as SSA cell line) based on SSA to restore GFP coding and produce a fluorescent signal.
[0575] Construction strategy of two different cell lines: different GFP expression modules were inserted into the lentiviral plasmid backbone, then different modules were integrated into the chromosomes of HEK293 cells by lentivirus packaging and infection, and different cell lines with stable genetic fluorescent systems were obtained by drug screening.
[0576] Core sequence of the GFP expression module for constructing the Reporter3 cell line:
[0577] The uppercase base sequence is the GFP expression framework, which is not an integer multiple of 3 (frameshift). The uppercase bold sequence is the target region, which is edited to produce indels, which may restore the normal reading frame of GFP and allow GFP to be normally expressed, producing fluorescence. The designed spacer sequence is the uppercase bold and underlined sequence, i.e., GAACGGCTCGGAGATCATCATTG (SEQ ID NO: 196), and the corresponding PAM is ATG.
[0578] Core sequence of GFP expression module of SSA cell line:
[0579] Wherein the capital base sequence is GFP expression frame, non-3 integer multiple (frame shift), and contains SSA (single strand annealing) homologous fragment (capital, not bold and underlined sequence), the capital bold part is the target region, which is edited to cause the homologous fragment to recombine, so as to restore the normal reading frame of GFP and make GFP normal expression, produce fluorescence; The capital bold and underlined sequence is the designed spacer sequence, that is, GAACGGCTCGGAGATCATCATTG (SEQ ID NO: 198), and the corresponding PAM is ATG.
[0580] The lentivirus backbone plasmids of two different cell lines are constructed as follows:
[0581] pCDH-CMV-GFP-Reporter3-EF1a-Puro (SEQ ID NO: 199) and
[0582] pCDH-CMV-SSA-GXXFP3 (SEQ ID NO: 200).
[0583] Example five, site-directed mutagenesis of C12-334 protein to construct different mutants
[0584] The amino acid sequence of C12-334 protein is as follows:
[0585] 1) Mutant template plasmid information
[0586] First, the sequence of C12-334 and the target sequence of the corresponding reporter system are constructed into pCDNA3.1(+) plasmid, and the resistance gene is changed from Amp to Kan, obtaining the template plasmid C12-334-pCDHPAM02 (SEQ ID NO: 201). It encodes C12-334 protein and gRNA, and the gRNA guide sequence is CAATGATGATCTCCGAGCCGTTC (SEQ ID NO: 202).
[0587] Different mutant clones were obtained: Specifically, two mutant primers F / R were designed at the mutation site to introduce the desired mutation sequence by primers. Combined with universal primers at both ends of the vector, PCR amplification was performed to obtain two mutant fragments F1 and F2. The mutant fragments F1 and F2 were recombined with the linearized vector to obtain the mutant plasmid. Taking the two sites of amino acid 115 and amino acid 697 as examples, how to construct single-point mutant (N115R, D697R) and multi-point mutant (N115R&D697R) plasmids was described. The specific steps are as follows:
[0588] The primers for constructing site 115 and 697 site-directed mutations are shown in Table 4.
[0589] Table 4. Primers
[0590] Using the C12-334-pCDHPAM02 plasmid as a template, the primer ChkCas12-PF1+334_115_R was used for PCR amplification (UltraHiPFTMDNA Polymerase Kit) to obtain the mutant fragment N115R-F1, and the primer ChkCas12-PR4+334_115_F was used for PCR amplification to obtain the mutant fragment N115R-F2. The plasmid C12-334-pCDHPAM02 was digested with HindIII+KpnI, and the 5665bp vector fragment was recovered from the gel, and the fragments N115R-F1 and N115R-F2 were recombined in vitro (NEB, Gibson Mix), and the 115 site mutant plasmid C12-334-pCDHPAM02-115 was obtained by heat shock transformation of E. coli.
[0591] Using the same method, the primer 334_115_F was replaced by 334_697_F, and the primer 334_115_R was replaced by 334_697_R to obtain the 697 site mutant plasmid C12-334-pCDHPAM02-697. The construction method of other different site Cas protein mutant plasmids is the same as the construction method of the 115 and 697 site mutant plasmids described above, except that mutant primers for each different site are designed.
[0592] 2) Verification of Cas protein mutant editing activity
[0593] Reporter3 or SSA cell culture and plating: The cell line was cultured to a fusion degree of 70-80% for plating, and the cell number for inoculation in a 24-well plate was 5*10^5 cells / well.
[0594] Transfection: Transfection was performed 12-14h after plating. 1.5ul PEI (Promega) + 500ng mutant plasmid was added into 100ul Opti-MEM per well of 24-well plate, mixed and incubated at room temperature for 20min, then added into corresponding cell line for cell transfection. After overnight transfection, fresh medium was replaced and the cells were cultured for another 72h. Flow cytometry was used to detect the editing efficiency of different mutant clones, and the fold change of editing efficiency of each mutant compared with wild type was calculated. The test was repeated and the results were averaged.
[0595] 3) Construction of mutant clone combination and verification of editing activity
[0596] According to the editing efficiency results of different mutation sites described above, different mutation sites were selected for multi-mutation combination. For example, the plasmid construction of 115 and 697 double-site mutant.
[0597] Using C12-334-pCDHPAM02 plasmid as template, primer ChkCas12-PF1+334_115_R was used for PCR amplification to obtain mutant fragment N115R-F1, primer 334_115_F+334_697_R was used for PCR amplification to obtain mutant fragment N115R-D697R-F1, and primer ChkCas12-PR4+334_697_F was used for PCR amplification to obtain mutant fragment D697R-F2. The plasmid C12-334-pCDHPAM02 was digested with HindIII+KpnI, and the 5665bp vector fragment was recovered by gel recovery, and the fragments N115R-F1, N115R-D697R-F1 and D697R-F2 were recombined in vitro, and the E. coli was transformed by heat shock to obtain the variant plasmid C12-334-pCDH-115-697. The construction method of other multi-point mutant plasmids is consistent with C12-334-pCDH-115-697.
[0598] Using the same method as described above, the editing activity of the multi-point mutant was verified using the SSA cell line. The editing efficiency of different mutant clones was represented by the proportion of GFP positive cells, and the fold change of editing efficiency of each mutant compared with wild type was calculated. The test was repeated and the results were averaged.
[0599] The editing efficiency results of each mutant (single-point mutant or multi-point mutant) are shown in Table 5 and Table 6. The data in the table is the average of repeated test results.
[0600] The editing efficiency of multi-point mutant in SSA cell line is shown in Figure 10A and Figure 10B.
[0601] Table 5. Editing efficiency of each mutant in Reporter3 cell line
[0602] Table 6. Editing efficiency of each mutant in SSA cell line
[0603] Based on the analysis of the Alphafold3 structure prediction results of the C12-334+gRNA+target DNA ternary complex, it can be seen that many of the gene editing efficiency improved advantage mutants screened in the experiment are related to the binding mechanism of Cas protein and nucleic acid (sgRNA / dsDNA): after the corresponding amino acids are mutated to Arg, they form salt bridges with the phosphate backbone of sgRNA or dsDNA, and form cation-pi or hydrogen bonds with the bases, thereby improving the binding ability of Cas protein and sgRNA / dsDNA, thereby improving the editing efficiency. As shown in Table 7 and FIG. 11.
[0604] In addition, D480, E675, D757 belong to the enzyme active center, and after the mutation of these residues or nearby residues, the cleavage activity is completely lost or partially lost.
[0605] The C12-334 protein comprises domains (expressed in amino acid residue position intervals): aa1-24 WED, aa25-109 Helical 1, aa110-182 PI, aa183-340 Helical 1, aa341-447 WED, aa448-522 RuvC, aa523-644 Helical 2, aa645-720 RuvC, aa721-756 Nuc, aa757-769 RuvC, and aa770-891 Nuc.
[0606] Table 7. Structure analysis of mutants with improved cleavage activity
[0607] Example Six, Editing Efficiency Detection of Different Mutants for HPRT1
[0608] In this example, first, the target point with PAM sequence WYR for the HPRT1 gene (GeneBank: NG_012329.2) is selected, and sgRNA targeting different positions (as shown in Table 8) is designed and constructed, which is used for combination with different mutants to detect editing efficiency.
[0609] Table 8. sgRNAs targeting HPRT1
[0610] Based on the sgRNA sequences in the table above, sgRNA plasmids were constructed respectively: the vector plasmid SpCas9-gRNA-pUC57Kan (SEQ ID NO:219) was linearized by restriction enzyme digestion with BbsI (Thermofisher) and XhoI (Thermofisher), primers were synthesized for different sgRNA sequences, and after annealing, the plasmids were ligated into the linearized vector and transformed into E. coli to obtain the final sgRNA expression vector plasmid.
[0611] 1) Editing efficiency of C12-334 wild-type and different sgRNA combinations targeting HPRT1
[0612] First, the wild-type C12-334 plasmid, C12-334-pCDHPAM02, was combined with five different sgRNA plasmids, and HEK293T cells were transfected with PEI. Cells were collected 48 hours later and then... Lysis Reagent (Cell) (VIAGEN) cells were lysed, and different primers were selected for PCR amplification based on different target sites. Sanger sequencing was performed, and TIDE analysis was used to determine the editing efficiency.
[0613] The specific amplification and sequencing primers are shown in Table 9. The editing efficiency results are shown in Table 10. Among them, C12-334-HPRT1-sgRNA05 showed the highest editing efficiency.
[0614] Table 9. Primers for amplification and sequencing of different targets
[0615] Table 10. Editing efficiency of wild-type C12-334 targeting HPRT1.
[0616] 2) Editing efficiency of different C12-334 mutants combined with C12-334-HPRT1-sgRNA05
[0617] Similar to the editing efficiency test method in "1), different mutants were combined with C12-334-HPRT1-sgRNA05, and HEK293T cells were transfected with PEI. Cells were collected 48 hours later and then... Lysis Reagent (Cell) (VIAGEN) cells were lysed, amplified by PCR, and sequenced using Sanger sequencing. Editing efficiency was analyzed using TIDE. The editing results produced by different mutants are shown in Figure 12, which illustrates the editing efficiency of single-point and three-point mutants.
[0618] Example Seven
[0619] This example tests editing activity by delivering a modified gRNA (C12-334-dmHPRT1-sgRNA05-01) and mRNA.
[0620] The sequence of the C12-334-dmHPRT1-sgRNA05-01 gRNA targeting the HPRT1 gene is: dG*dT*dT*dGdCdAdAdTdCdCdCdAdAdGrGrUrGrCrGrArArArCrGrGrUrCrUrCrGrUrUrArGrArGrGrCrUrGrGrUrUrCrArArGrCrArCrGrArArArGrGrGrUrGrUrUrUrArUrUrCrCrUrCrA*mU*mG*mG (SEQ ID NO: 224), with the first 14 nt at the 5’ end being a DNA base-modified sequence, and the 5’ end 3 bases of the gRNA being phosphorothioate-modified, and the last 3 bases being phosphorothioate- and 2’-methoxy-modified.
[0621] C12-334 wild-type and N115R, D697R mutant encoding mRNA were obtained by in vitro transcription and combined with chemically synthesized C12-334-dmHPRT1-sgRNA05-01, respectively, to transfect HEK293 cells by electroporation. After 72 h, the cells were collected, lysed with Lysis Reagent (Cell) (VIAGEN), and subjected to PCR amplification using primers ChkHPRT1-NGS-PF1 (SEQ ID NO: 225) and ChkHPRT1-NGS-PR1 (SEQ ID NO: 226) to construct a library and perform NGS sequencing to detect editing efficiency. The editing efficiency of the wild-type, N115R, and D697R mutant combined with C12-334-dmHPRT1-sgRNA05-01 reached 48.80%, 90.88%, and 92.77%, respectively. As shown in FIGS. 13A-13C.
[0622] Example Eight, Activity Testing of C12-314 to C12-333, C12-337, C12-342 to C12-354 in Table 1
[0623] 1. Vector Construction
[0624] The pET28a vector plasmid was double-digested with BamHI and XhoI, and the linearized vector was recovered by agarose gel electrophoresis. The DNA fragment encoding the sequence of the Cas protein of the present disclosure was obtained, and homologous recombination was used (NEB, Gibson The reaction solution was transformed into Stbl3 competent cells, and coated on LB plates resistant to kanamycin sulfate. After overnight culture at 37°C, the clones were picked and sequenced for identification.
[0625] The positive clones with correct sequences were cultured overnight, and the plasmids were extracted and transformed into the expression strain Rosetta (DE3). The LB plates containing kanamycin sulfate were coated and incubated overnight at 37°C.
[0626] 2. Expression of recombinant protein
[0627] The single clones were inoculated into 5 ml LB culture solution containing kanamycin sulfate and incubated overnight at 37°C.
[0628] The culture was transferred into 500 ml LB culture solution containing kanamycin sulfate at a ratio of 1:100, and incubated at 37°C at a speed of 220 rpm until the OD reached 0.6. IPTG was added to a final concentration of 0.2 mM, and the culture was induced at 16°C for 24 hours.
[0629] The bacteria were washed with 15 ml PBS and centrifuged to collect the bacterial cells. The cells were broken by ultrasonic treatment in lysis buffer, and centrifuged at 10,000 g for 30 min to obtain the supernatant containing the recombinant protein. The supernatant was filtered through a 0.45 μm filter and then purified by column chromatography.
[0630] 3. Purification of recombinant protein
[0631] The N-terminal 6 His were used as a purification tag for IMAC purification. The recombinant protein of Cas protein was purified by chromatography, and the structure of the recombinant protein was His tag-NLS-Cas-NLS-NLS. The purified recombinant protein was detected by SDS-PAGE electrophoresis.
[0632] 4. Determination of PAM sequence recognized by Cas protein
[0633] The sgRNA (single guide RNA) containing the specific guide sequence was mixed with the purified recombinant protein, and the in vitro cleavage substrate (containing a spacer sequence and a 7 nt random sequence) was cleaved. After incubation at 37°C, the product was purified, library construction was performed, and NGS sequencing was carried out to determine the PAM sequence recognized by the Cas protein.
[0634] The designed in vitro cleavage substrate sequence is as follows:
[0635] In the sequence, N represents any one of A, T, C, or G.
[0636] The cleavage substrate was taken to the sequencing company for PCR-Free library construction and NGS sequencing.
[0637] 5. Preparation of sgRNA
[0638] sgRNA (5'-DR-guide sequence-3') containing any corresponding DR sequence of Cas12 protein in Table 2 was transcribed in vitro, and the transcription product was precipitated and purified with LiCl. The guide sequence was GUGAGCAAGGGCGAGGAGCUGUUC (SEQ ID NO: 228) or CGCAAUGAUGAUCUCCGAGCCGUUCC (SEQ ID NO: 229).
[0639] PAM library cleavage, NGS sequencing, and analysis of NGS results: The captured 7 nt random sequence was analyzed using WebLogo software according to the method in the reference (A compact Cas9 ortholog from Staphylococcus Auricularis (SauriCas9) expands the DNA targeting scope. PLoS biology, 2020, 18(3), e3000686.) to identify the PAM sequence.
[0640] 6. Test of in vitro cleavage activity of Cas protein
[0641] The foregoing sgRNA and recombinant protein were mixed to cleave the target DNA (dsDNA or ssDNA) in vitro, and the cleavage product exhibited the cleavage effect of the Cas protein by gel electrophoresis.
[0642] 7. Test of editing activity in eukaryotic cells
[0643] The foregoing sgRNA (the guide sequence was replaced with a guide sequence targeting human cells, i.e., any one of SEQ ID NOs: 209-213) and recombinant protein were mixed to obtain an RNP, which was transfected into 293T cells. After overnight transfection, fresh culture medium was added for continued culture.
[0644] DNA extraction, PCR amplification, and Sanger sequencing: After 72 h of culture, the cells were washed with PBS, and then 100 μl of cell lysis solution was added for lysis to obtain a lysis solution containing genomic DNA. The genomic DNA was amplified near the target sequence region, and the PCR product was sent to a sequencing company for Sanger sequencing. Sequencing data analysis: The sequencing peak map and gRNA guide sequence related information were analyzed by TIDE to obtain the editing efficiency of the target nucleic acid.
Claims
1. A Casl2 protein, characterized in that, The Cas12 protein meets any one of the following (a)-(c): (a) the Cas12 protein is a CLUSTER1 protein, a CLUSTER2 protein, a CLUSTER3 protein, a CLUSTER4 protein, a CLUSTER5 protein, a CLUSTER6 protein, a CLUSTER7 protein, a CLUSTER8 protein, a CLUSTER9 protein, a CLUSTER10 protein, a CLUSTER11 protein, or a CLUSTER12 protein; (b) the amino acid sequence of the Cas12 protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity compared to any one of SEQ ID NOs: 1-35; and (c) the Cas12 protein belongs to a Cas12h subtype, the Cas12 protein is capable of forming a complex with a guide polynucleotide, the complex is targeted to bind to a target nucleic acid within a eukaryotic cell; Optionally, the Cas12 protein retains the function of the protein as shown in the sequence of any one of SEQ ID NOs: 1-35; Optionally, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or several of the following: A, C, T, G, TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA, NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG, AAGG, AAGN, NTNN, TCGT, CNTG, NTGG, CCGN, ATAT, TGCA, NGGT, TGNT, NNTG, NCCG, ACAT, GNTG, CGCG, GACN, NTCG, TCNG, CTGC, TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG, CNTC, TGGN, NATC, GTCG, ACNC, GCNN, GACT, CTNT, NCTT, NAGG, NANC, CTTA, GTCT, ANAG, NGCN, CNNA, TCAG, ACAC, NCGG, TNNT, CAAG, ACCT, CCCA, GTNC, ANTC, GACC, AACG, TTAA, TCCG, CGCC, NCCN, TTNA, NCNT, NGCA, AGNN, AATC, GGGA, GNAN, NAGA, CGNA, GTAT, GTNA, ATNC, ACNA, GGAA, NTCC, GGCG, AATN, CNNT, AGGC, GCGN, GTGC, TTGA, AAGC, GAAG, ATNG, TGCT, TACT, CTAN, GGCT, GNGC, GTCN, CGAA, CNAC, GCCT, TAGG, ANGC, TNAA, GANT, NCNA, NCCT, AGAN, GTAA, TTTN, ATGA, TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN, TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG, NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNT,CCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG, NCNG, AANG, GGTC, ACTC, TGAA, NAGN, NNCA, ACGG, TGAC, TCCN, ANNN, TCGN, TAAN, CAGG, TTAN, NGAN, NTGC, CCNC, TNTN, ATGN, GTGN, GCAT, NNGN, NNCC, CCNA, CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN, GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC, CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN, GAAC, NCCA, TTCC, NCNN, GNNN, ANGT, NTNA, CCCT, GNAA, TTNG, GTNN, GGNG, TCTA, NCAC, GANG, TTCG, CCTC, CNGG, ANNA, TCAN, ATCG, NTGA, CGTA, TTAC, GCTN, GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC, TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTAN,GCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NNTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG, ACGN, GGCA, AAAG, TTGT, NGNA, NAAN, TATN, CGGG, CATA, ATGC, ACGC, ACCN, ATTT, TCNA, TNGC, NACA, NACC, CTCN, GGCC, TANG, AGAA, TNGA, TAGC, CAGN, GGCN, ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN, GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA, NCTC, NNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC, ATNN, NNNT, AGAC, AACC, ANNC, ANNG, ACAN, GTTC, TATA, GNTC, NCGT, NGNT, CGTC, CCGC, CGAC, GACG, ATTG, GNNC, CNAA, TATC, AGNA, CTNC, TTCA, ANCA, ACCC, AGTG, CCGT, ANAT, CTGT, GGGC, NTTA, NAAG, AANN, CNAN, NNCN, ANAA, ANAN, CTTG, NGNN, AGAG, TANA, TCNC, GCAA, NGNG, NAGC, NATA, ATCN, CGTT, CNGC, GATN, NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATT,NNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTC, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC, Optionally, the PAM sequence (5'→3') recognized by the Cas12 protein is optionally selected from any one or more of the following: WYR, BMCTTH, TTN, VNWTV, VNWTC, VNTTC, wherein W is A or T, Y is C or T, R is A or G, B is C, G or T, M is A or C, H is A, T or C, N is A, T, C or G, V is A, C or G; Optionally, the Cas12 protein forms a complex with a guide polynucleotide; further, the complex specifically binds to a target nucleic acid; further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate expression of the target nucleic acid; Optionally, the Cas12 protein forms a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence reverse-complementary to a target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that interacts with the Cas12 protein; further, the backbone sequence comprises or is a direct repeat sequence; Optionally, the backbone sequence does not comprise a tracrRNA sequence; Optionally, the Cas12 protein is a nuclease-inactive variant; optionally, the Cas12 protein is a dead Cas12 inactive variant or a nickase Cas12 inactive variant; optionally, the Ruvc domain of the Cas12 protein is inactivated. Optionally, the Cas12 protein is capable of forming a complex with a guide polynucleotide and targetably binding to a target nucleic acid within a eukaryotic cell; further optionally, the Cas12 protein is capable of forming a complex with a guide polynucleotide and targetably binding to and cleaving a target nucleic acid within a eukaryotic cell; Optionally, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 18; optionally, the Cas12 protein is non-natural, or, engineered; optionally, the Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves a target nucleic acid, or, the complex sequence-specifically binds a target nucleic acid, but does not cleave the target nucleic acid; optionally, the complex is non-natural, or, engineered; optionally, the guide polynucleotide comprises a guide sequence and a scaffold sequence; optionally, the guide sequence is reverse-complementary to a target nucleic acid, the scaffold sequence interacts with the Cas12 protein; optionally, the scaffold sequence is a direct repeat sequence; optionally, the guide sequence is located at the 5’ end or the 3’ end of the scaffold sequence; optionally, the guide polynucleotide is non-natural, or, engineered; optionally, the Cas12 protein recognition sequence is a PAM of 5'-WYR-3', wherein W = A or T, Y = C or T, R = A or G; optionally, the Cas12 protein recognition sequence is a PAM of 5'-ACA-3', 5'-TCA-3', 5'-ATA-3', 5'-TTA-3', 5'-ACG-3', 5'-TCG-3', 5'-ATG-3', 5'-TTG-3', and / or 5'-TTN-3'; optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 84-86, 187-195.Optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 84; optionally, the Cas12 protein has a mutation at at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, or at least 12 of amino acid residues corresponding to positions 1-891 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue R, H, K, or A; optionally, the mutation is a mutation to residue R; optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, or more of the amino acid residues corresponding to positions N5, D9, E58, S100, N115, K142, C148, S147, K232, S245, I251, Y263, D279, A297, L300, E303, L337, M378, N394, T396, T443, K458, T468, K533, F537, F548, N550, D697, A706, I788 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to residue R; optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to positions D480, E675, D757 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A.
2. A guide polynucleotide, comprising, comprising (i) a direct repeat sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one of SEQ ID NOs: 36-170, 187-195, (ii) a guide sequence engineered to hybridize to a target nucleic acid; the direct repeat sequence is linked to the guide sequence, the guide polynucleotide is capable of forming a complex with a Cas12 protein and directing the complex to sequence-specific binding to the target nucleic acid; Preferably, the Cas12 protein is as defined in claim 1. Optionally, the guide sequence comprises 15-60 nucleotides, Optionally, the guide sequence hybridizes to the target nucleic acid with no more than one nucleotide mismatch. Optionally, the guide polynucleotide does not comprise or comprises a tracrRNA. Optionally, the tracrRNA sequence is linked to the direct repeat sequence. Optionally, the tracrRNA comprises 10-200 nucleotides. Optionally, the guide sequence is located at the 3' end of the direct repeat sequence. Optionally, the guide sequence is located at the 5' end of the direct repeat sequence.
3. A Casl2 inactivated variant, characterized in that, The Cas12 inactivated variant is a nuclease-activity-inactivated variant of the Cas12 protein as defined in claim 1. Optionally, the Cas12 inactivated variant is a dead Cas12 inactivated variant or a nickase Cas12 inactivated variant. Optionally, the Cas12 inactivated variant is a Ruvc domain-inactivated of the Cas12 protein.
4. A fusion protein or conjugate, characterized in that, The fusion protein or conjugate comprises the following elements: (1) the Cas12 protein as defined in claim 1, or the Cas12 inactivated variant as defined in claim 3; and (2) a homologous or heterologous functional domain. Optionally, the functional domain has an enzymatic activity that modifies a target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deaminase activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristoylation activity, and / or demyristoylation activity; Optionally, the homologous or heterologous functional domain is optionally selected from one or more of the following: a subcellular localization signal, a DNA binding domain, a protease domain, a transcriptional activation domain, a transcriptional repression domain, a nuclease domain, a deaminase domain, a uracil DNA glycosylase domain (UDG), a uracil DNA glycosylase inhibitor domain (UGI), a methylase, a demethylase, a transcriptional release factor, a histone acetylase domain, a histone deacetylase domain, a DNA ligase, an affinity tag, a reporter tag, an affinity domain, and a reporter domain; Optionally, the nuclease domain comprises a polypeptide having ssDNA cleavage activity and / or a polypeptide having dsDNA cleavage activity; Optionally, the Cas12 protein or inactivated variant is directly or indirectly linked to the homologous or heterologous functional domain; preferably, the direct linkage is a covalent linkage and the indirect linkage is via an amino acid linker or a non-amino acid linker; Optionally, the homologous or heterologous functional domain is fused or conjugated to the N-terminus, C-terminus, or internally relative to the Cas12 protein or inactivated variant; Optionally, the amino acid sequence of the Cas12 protein comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in SEQ ID NO: 18; optionally, the Cas12 protein is non-natural, or, engineered; optionally, the Cas12 protein forms a complex with a guide polynucleotide, the complex sequence-specifically binds to a target nucleic acid; optionally, the complex sequence-specifically binds and cleaves a target nucleic acid, or, the complex sequence-specifically binds a target nucleic acid, but does not cleave the target nucleic acid; optionally, the complex is non-natural, or, engineered; optionally, the guide polynucleotide comprises a guide sequence and a scaffold sequence; optionally, the guide sequence is reverse-complementary to a target nucleic acid, the scaffold sequence interacts with the Cas12 protein; optionally, the scaffold sequence is a direct repeat sequence; optionally, the guide sequence is located at the 5’ end or the 3’ end of the scaffold sequence; optionally, the guide polynucleotide is non-natural, or, engineered; optionally, the Cas12 protein recognition sequence is a PAM of 5'-WYR-3', wherein W = A or T, Y = C or T, R = A or G; optionally, the Cas12 protein recognition sequence is a PAM of 5'-ACA-3', 5'-TCA-3', 5'-ATA-3', 5'-TTA-3', 5'-ACG-3', 5'-TCG-3', 5'-ATG-3', 5'-TTG-3', and / or 5'-TTN-3'; optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 84-86, 187-195.Optionally, the scaffold sequence comprises a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 84; optionally, the Cas12 protein has a mutation at at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, or at least 12 of amino acid residues corresponding to positions 1-891 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue R, H, K, or A; optionally, the mutation is a mutation to residue R; optionally, the Cas12 protein has a mutation at any 1, any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, or more of the amino acid residues corresponding to positions N5, D9, E58, S100, N115, K142, C148, S147, K232, S245, I251, Y263, D279, A297, L300, E303, L337, M378, N394, T396, T443, K458, T468, K533, F537, F548, N550, D697, A706, I788 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to residue R; optionally, the Cas12 protein has a mutation at any 1, any 2, or 3 of the amino acid residues corresponding to positions D480, E675, D757 of the sequence set forth in SEQ ID NO: 18; optionally, the mutation is a mutation to any other natural amino acid residue; optionally, the mutation is a mutation to residue A; optionally, the functional domain has an epigenome modification activity, an epigenetic modification activity, or an epigenetic modification activity; optionally, the epigenome modification, epigenetic modification, or epigenetic modification comprises, but is not limited to, DNA methylation, RNA methylation, RNA interference, nucleosome positioning, chromatin conformation alteration, chromatin remodeling, histone modification, modification of long non-coding RNA sequence; optionally, the functional domain is an epigenome modification functional domain, an epigenetic modification functional domain, or an epigenetic modification functional domain; optionally, the functional domain has a single base editing activity.Optionally, the functional domain is optionally selected from one or more of: a nuclease (e.g., Fokl), a DNA methylase, a DNA demethylase, a histone methylase, a histone demethylase, a DNA repair enzyme, a DNA damaging enzyme, a base deaminase (including but not limited to an adenine deaminase, a cytosine deaminase), a dismutase, an alkylating enzyme, a depurinase, an oxidizing enzyme, a pyrimidine dimer forming enzyme, an integrase, a transposase, a recombinase, a polymerase, a ligase, a helicase, a photolyase, a glycosylating enzyme, a deglycosylating enzyme, an acetyltransferase, a deacetylase, a kinase, a phosphatase, a ubiquitin ligase, a deubiquitinating enzyme, an adenylylating enzyme, a deadenylase, a SUMOylating enzyme, a deSUMOylating enzyme, a myristoylating enzyme, and / or a demyristoylating enzyme; optionally, the functional domain is an adenine deaminase or a cytosine deaminase.
5. An isolated nucleic acid, comprising, the nucleic acid encodes the Cas12 protein of claim 1, the Cas12 inactivated variant of claim 3, or the fusion protein or conjugate of claim 4; Optionally, the nucleic acid is codon-optimized for expression in a cell; Optionally, the nucleic acid is codon-optimized for expression in a prokaryotic cell; Optionally, the nucleic acid is codon-optimized for expression in a eukaryotic cell; Optionally, the nucleic acid is codon-optimized for expression in a eukaryote, a mammal such as a human or a non-human mammal, a plant, an insect, a bird, a reptile, a rodent (e.g., a mouse, a rat), a fish, a worm / nematode, or a yeast.
6. A CRISPR-Cas12 system, characterized in that, The CRISPR-Cas12 system comprises: a) the Cas12 protein of claim 1, the Cas12 inactivated variant of claim 3, the fusion protein or conjugate of claim 4, or the nucleic acid of claim 5; and b) a guide polynucleotide, or a polynucleotide sequence encoding the guide polynucleotide; Optionally, the Cas12 protein, Cas12 inactivated variant, fusion protein or conjugate forms a complex with the guide polynucleotide; the guide polynucleotide comprises a guide sequence engineered to direct sequence-specific binding of the complex to a target nucleic acid; Optionally, the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, preferably the direct repeat sequence has at least 50% identity compared to any one of SEQ ID NOs: 36-170, 187-195; Optionally, the guide sequence comprises 15-35 nucleotides, and / or the guide sequence is 90-100% complementary to the target nucleic acid, preferably with no more than one nucleotide mismatch; Optionally, the guide sequence comprises 15-60 nucleotides; Optionally, the guide sequence hybridizes to the target nucleic acid; Optionally, the guide sequence has no more than one nucleotide mismatch with the target nucleic acid; Optionally, the guide sequence is located at the 3' end of the direct repeat sequence; Optionally, the guide sequence is located at the 5' end of the direct repeat sequence; Optionally, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA; Optionally, the DNA is eukaryotic DNA; preferably the eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, avian DNA, reptilian DNA, rodent DNA, fish DNA, worm / nematode DNA or yeast DNA; Optionally, the target nucleic acid is a disease- or disorder-associated gene or a signaling biochemical pathway-associated gene, or the target nucleic acid is a reporter gene.
7. A vector system characterized by comprising The vector system comprises one or more recombinant vectors, the recombinant vector comprising the isolated nucleic acid of claim 5, or the CRISPR-Cas12 system of claim 6; Optionally, the recombinant vector further comprises a regulatory sequence; Optionally, the polynucleotide sequence encoding the Cas12 protein, Cas12 inactivated variant or fusion protein or conjugate is operably linked to a regulatory sequence, and / or the polynucleotide sequence encoding the guide polynucleotide is operably linked to a regulatory sequence; more preferably, the regulatory sequence is optionally selected from one or more of a promoter, an enhancer, an internal ribosome entry site and a transcription termination signal, the promoter being for example a constitutive promoter, an inducible promoter, a broad-spectrum promoter or a tissue-specific promoter, and / or the transcription termination signal being for example a polyadenylation signal or a poly-U sequence; Optionally, the backbone of the recombinant vector is an adeno-associated viral vector, a lentiviral vector, a ribonucleoprotein complex, or a virus-like particle; preferably, when the backbone is an adeno-associated viral vector, the adeno-associated viral vector is a recombinant adeno-associated viral vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74; when the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; Optionally, the isolated nucleic acid is linked to an aptamer sequence; when the backbone of the recombinant vector is a virus-like particle, the isolated nucleic acid is linked to a gene encoding a gag protein.
8. A delivery system characterized by, The delivery system comprises: (1) a delivery tool, and (2) the Cas12 protein of claim 1, the guide polynucleotide of claim 2, the Cas12 inactivated variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, or the vector system of claim 7; Optionally, the delivery tool is a virus, a lipid nanoparticle, a nanoparticle, a liposome, an exosome, a microvesicle, or a gene gun; Optionally, the delivery tool is a lipid nanoparticle, the lipid nanoparticle comprising the guide polynucleotide and an mRNA encoding the Cas12 protein, the Cas12 inactivated variant, or the fusion protein or conjugate.
9. A cell, characterized in that, The cell comprises the Cas12 protein of claim 1, the guide polynucleotide of claim 2, the Cas12 inactivated variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, or the delivery system of claim 8; Optionally, the cell is a prokaryotic cell; Optionally, the cell is a eukaryotic cell; Optionally, the eukaryotic cell is a mammalian cell.
10. A pharmaceutical composition, characterized by, The pharmaceutical composition comprises the Cas12 protein of claim 1, the guide polynucleotide of claim 2, the Cas12 inactivated variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the CRISPR-Cas12 system of claim 6, the vector system of claim 7, the delivery system of claim 8, or the cell of claim 9; Preferably, the pharmaceutical composition comprises a pharmaceutically acceptable excipient.
11. A kit characterized in that, The kit comprises a Cas12 protein of claim 1, a guide polynucleotide of claim 2, a Cas12 inactivated variant of claim 3, a fusion protein or conjugate of claim 4, a nucleic acid of claim 5, a CRISPR-Cas12 system of claim 6, a vector system of claim 7, a delivery system of claim 8, a cell of claim 9, or a pharmaceutical composition of claim 10.
12. Use of a Cas12 protein of claim 1, a guide polynucleotide of claim 2, a Cas12 inactivated variant of claim 3, a fusion protein or conjugate of claim 4, a nucleic acid of claim 5, a CRISPR-Cas12 system of claim 6, a vector system of claim 7, a delivery system of claim 8, a cell of claim 9, a pharmaceutical composition of claim 10, or a kit of claim 11 in the manufacture of a medicament or a drug for the diagnosis, treatment and / or prevention of a disease or disorder associated with a target nucleic acid; Optionally, the disease or disorder is a hematological disease or disorder, an ophthalmological disease or disorder, a neurological disease or disorder, a respiratory disease or disorder, a liver disease or disorder, a metabolic disease or disorder, a cancer, or an infectious disease; and / or the medicament or drug is for cleaving or nicking one or more target nucleic acid molecules, activating or upregulating the expression of one or more target nucleic acid molecules, activating or repressing the transcription of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling or detecting one or more target nucleic acid molecules, binding to one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules.
13. A method of detecting, binding or cleaving a target nucleic acid, characterized in that, The method comprises contacting a target nucleic acid with a Cas12 protein of claim 1, a guide polynucleotide of claim 2, a Cas12 inactivated variant of claim 3, a fusion protein or conjugate of claim 4, a nucleic acid of claim 5, a CRISPR-Cas12 system of claim 6, a vector system of claim 7, a delivery system of claim 8, a cell of claim 9, a pharmaceutical composition of claim 10, or a kit of claim 11; Optionally, the method is a method for non-diagnostic and / or therapeutic purposes; and / or the fusion protein or conjugate comprises a detectable label, such as a label detectable by fluorescence, Southern blotting, or FISH.
14. A method of changing the state of a cell, characterized by, The method comprises contacting a Cas12 protein of claim 1, a guide polynucleotide of claim 2, a Cas12 inactivated variant of claim 3, a fusion protein or conjugate of claim 4, a nucleic acid of claim 5, a CRISPR-Cas12 system of claim 6, a vector system of claim 7, a delivery system of claim 8, a cell of claim 9, a pharmaceutical composition of claim 10, or a kit of claim 11 with a cell, thereby altering the state of the cell; Optionally, the method results in one or more of: an increase or decrease in expression of a particular gene, induction of cellular senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, cell growth promotion and / or cell growth inhibition in vitro or in vivo, induction of anergy in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo; Optionally, the method is a method for non-diagnostic and / or therapeutic purposes.
15. A method of diagnosing, treating or preventing a disease or disorder associated with a target nucleic acid, characterized in that, A sample of a subject in need thereof or a subject in need thereof is administered a Cas12 protein of claim 1, a guide polynucleotide of claim 2, a Cas12 inactivated variant of claim 3, a fusion protein or conjugate of claim 4, a nucleic acid of claim 5, a CRISPR-Cas12 system of claim 6, a vector system of claim 7, a delivery system of claim 8, a cell of claim 9, a pharmaceutical composition of claim 10, or a kit of claim 11; Optionally, the disease or disorder is a hematological disease or disorder, an ophthalmological disease or disorder, a neurological disease or disorder, a respiratory disease or disorder, a liver disease or disorder, a metabolic disease or disorder, a cancer, or an infectious disease.
16. A Cas12 protein of claim 1, a guide polynucleotide of claim 2, a Cas12 inactivated variant of claim 3, a fusion protein or conjugate of claim 4, a nucleic acid of claim 5, a CRISPR-Cas12 system of claim 6, a vector system of claim 7, a delivery system of claim 8, a cell of claim 9, a pharmaceutical composition of claim 10, or a kit of claim 11 for use in the diagnosis, treatment or prevention of a disease or disorder associated with a target nucleic acid; Optionally, the disease or disorder is a hematological disease or disorder, an ophthalmological disease or disorder, a neurological disease or disorder, a respiratory disease or disorder, a liver disease or disorder, a metabolic disease or disorder, a cancer, or an infectious disease.
Citation Information
Patent Citations
Cas12 protein and use thereof
WO2025061113A1
An application of a Cas protein, and a method and kit for detecting a target nucleic acid molecule
CN107488710A
CRISPR-Cas system and application thereof
CN117384883A
Cas12 protein, gene editing system containing cas12 protein, and application
WO2022253185A1
Cas12 protein, crispr-cas system and uses thereof
WO2024042479A1