Engineered cas12 protein and use thereof
By engineering the structural domains of the Cas12 protein, its gene editing efficiency in mammalian cells has been improved, solving the problem of insufficient efficiency in existing technologies and achieving more efficient gene editing results.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ACCUREDIT THERAPEUTICS (SUZHOU) CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-07-30
AI Technical Summary
Existing CRISPR-Cas systems have limited gene editing efficiency in mammalian cells, making it difficult to achieve effective genome editing across multiple loci.
The Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II domains of the Cas12 protein were engineered and replaced with corresponding variant module sequences to form an engineered Cas12 protein, thereby improving its targeting efficiency in mammalian cells.
It achieves higher gene editing efficiency, lower off-target efficiency, and higher specificity in mammalian cells.
Smart Images

Figure PCTCN2026075196-FTAPPB-I100001 
Figure PCTCN2026075196-FTAPPB-I100002 
Figure PCTCN2026075196-FTAPPB-I100003
Abstract
Description
Engineered Cas12 protein and its applications Technical Field
[0001] This disclosure pertains to the field of biotechnology, and more specifically to the field of gene editing technology. More specifically, this disclosure relates to engineered Cas12 proteins that exhibit enhanced targeting efficiency in mammalian cells, engineered guide RNAs (gRNAs) capable of cooperating with said engineered Cas12 proteins, and their uses. Background Technology
[0002] Genome editing is an important and useful technique in genome research. Several systems are available for genome editing, including the CRISPR-Cas system, the TALEN (transcription activator-like effector nuclease) system, and the zinc finger nuclease (ZFN) system.
[0003] The CRISPR / Cas system is an acquired immune system evolved by bacteria and archaea to defend against invasion by exogenous viruses or plasmids. It is a highly efficient and cost-effective genome editing technology with wide applications in a range of eukaryotes, from yeast and plants to zebrafish and humans (see Van der Oost 2013, Science 339: 768-770, and Charpentier and Doudna, 2013, Nature 495: 50-51). Besides basic research, the CRISPR / Cas gene editing system also has broad clinical application prospects. To date, based on the system's outstanding functional and evolutionary modularity, two classes of CRISPR-Cas systems (class 1 and class 2) have been characterized, including six types (I-VI). The type V CRISPR-Cas system, also known as the Cas12 family, differs from other CRISPR-Cas systems in that it is an RNA-mediated single-effects ribozyme driven by a single RuvC active site. As more and more Cas12s are discovered and identified, this family now includes multiple categories such as VA, VB, VC, VD, VE, VF, VG, VH, VI, VJ, and VK. Among them, Cas12a and Cas12b, due to their excellent reactivity, are used in various nucleic acid detection technologies and clinical testing products combined with amplification techniques.
[0004] However, current CRISPR-Cas systems have several limitations, including limited gene editing efficiency. Therefore, there is a need to improve methods and systems to increase targeting efficiency in mammalian cells, thereby enabling efficient genome editing across multiple loci. Summary of the Invention
[0005] One object of this disclosure is to provide an engineered Cas12 protein that has improved targeting efficiency in mammalian cells.
[0006] The wild-type Cas12 protein contains multiple domains. The inventors have discovered that by engineering six of these domains (Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II), engineered Cas12 proteins with superior editing efficiency compared to the wild-type Cas12 protein. Specifically, the inventors designed multiple variant sequences as module sequences for each of these six domains, replacing their respective wild-type domains with one or more module sequences to assemble a series of engineered Cas12 proteins (hereinafter also referred to as Cas12 variants). In other words, relative to the wild-type Cas12 protein, one or more of the Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II domains in the engineered Cas12 protein described in this disclosure are replaced with the corresponding variant module sequences.
[0007] Based on the above findings, the inventors completed this invention.
[0008] In a first aspect, this disclosure provides an engineered Cas12 protein, which, relative to the wild-type Cas12 protein, comprises:
[0009] (i) A Helical-I domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:17-26, or
[0010] (ii) A Helical-II domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:27-36, or
[0011] (iii) A Helical-III domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:37-46, or
[0012] (iv) A NUC domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO:47-56, 79-81, and 343-356, or
[0013] (v) A PI domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO:57-66, 77-78, and 327-342, or
[0014] (vi) A WED-II domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:67-76, or
[0015] (vii) Any combination of two or more of the above (i)-(vi).
[0016] The engineered Cas12 protein described herein is obtained by replacing one or more of the Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II domains of the wild-type Cas12 protein with the corresponding domains of (i) to (vi).
[0017] In some embodiments, the wild-type Cas12 protein is selected from any one of wild-type Cas12-1, Cas12-2, Cas12-3, and Cas12-4. The amino acid sequences of wild-type Cas12-1, Cas12-2, Cas12-3, and Cas12-4 are shown in SEQ ID NO:2-5, respectively. In a preferred embodiment, the wild-type Cas12 protein is the wild-type Cas12-1 protein. For example, the amino acid sequence of the wild-type Cas12-1 protein is shown in SEQ ID NO:2, wherein the amino acid sequences of the Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II domains are shown in SEQ ID NO:11-16, respectively. Furthermore, one or more of the Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II domains in the wild-type Cas12-1 protein can be replaced by corresponding domains from wild-type Cas12-2, Cas12-3, or Cas12-4. For example, the NUC domain in the wild-type Cas12-1 protein can be replaced by the NUC domain (SEQ ID NO:2) from wild-type Cas12-2, Cas12-3, or Cas12-4. The PI domain in the wild-type Cas12-1 protein can be replaced by a PI domain from wild-type Cas12-2, Cas12-3, or Cas12-4 (e.g., any one of SEQ ID NO: 77-78) or other homologous sequences (e.g., any one of SEQ ID NO: 327-342). Similarly, one or more of the above-mentioned domains in the wild-type Cas12-2 protein can also be replaced by the corresponding domains from wild-type Cas12-1, Cas12-3, or Cas12-4.
[0018] In one embodiment, the Helical-I domain comprises or consists of any of the sequences shown in SEQ ID NO:17-26. In one embodiment, the Helical-II domain comprises or consists of any of the sequences shown in SEQ ID NO:27-36. In one embodiment, the Helical-III domain comprises or consists of any of the sequences shown in SEQ ID NO:37-46. In one embodiment, the NUC domain comprises or consists of any of the sequences shown in SEQ ID NO:47-56, 79-81, and 343-356. In one embodiment, the PI domain comprises or consists of any of the sequences shown in SEQ ID NO:57-66, 77-78, and 327-342.
[0019] In some embodiments, the engineered Cas12 protein comprises, or is composed of, any of the amino acid sequences in SEQ ID NO:83-116 having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity. In a preferred embodiment, the engineered Cas12 protein comprises, or is composed of, any of the amino acid sequences in SEQ ID NO:83-116.
[0020] In some embodiments, the engineered Cas12 protein comprises, or is composed of, an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any of SEQ ID NO:121-144 or 370-404. In a preferred embodiment, the engineered Cas12 protein comprises, or is composed of, an amino acid sequence of any of SEQ ID NO:121-144 or 370-404.
[0021] In some embodiments, the engineered Cas12 protein comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO:133-144 having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity. In a preferred embodiment, the engineered Cas12 protein comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO:133-144. In a more preferred embodiment, the engineered Cas12 protein comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO:141, 133, 134, 135, 136, and 137. In another preferred embodiment, the engineered Cas12 protein comprises, or is composed of, the amino acid sequence of SEQ ID NO: 141 or 137. In yet another preferred embodiment, the engineered Cas12 protein comprises, or is composed of, the amino acid sequence of SEQ ID NO: 136 or 133.
[0022] In some embodiments, the engineered Cas12 protein described herein comprises, or is composed of, any of the amino acid sequences in SEQ ID NO:370-404.
[0023] In some embodiments, the engineered Cas12 protein of this disclosure comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO:157-180. In a preferred embodiment, the engineered Cas12 protein of this disclosure comprises an amino acid sequence of any one of SEQ ID NO:157-180.
[0024] In some embodiments, the engineered Cas12 protein described herein includes any one or more variations of the following positions relative to SEQ ID NO:141: E462X, D926X, T505X, E494X, D293X, T505X, D709X, E350X, A691X, E736X, E179X, D165X, D704X, E399X, Y292X, K931X, E262X, K703X, E1035X, V842X, E675X, E875X, E601X, E504X, C567X, K277X, D356X, D904X, K563X, E285X, E388X, G478X, E271X, H520X, K809X, L967 X, E815X, V842X, V602X, T354X, K138X, I657X, Y468X, N879X, Q186X, S318 X, D356X, I265X, D591X, K404X, V842X, K270X, H858X, E321X, W1046X, K18 7X, G232X, F891X, L97X, L13X, L74X, M618X, I1048X, S582X, D120X, V1047 X, S243X, H858X, N168X, E681X, Q620X, K401X, T943X, W927X, E253X, A854 X, G402X, K774X, W1046X, E793X, E252X, R302X, E447X, V1039X, R330X, D 906X, S240X, A624X, T235X, R902X, V1014X, R902X, D851X, F583X, Y537X, 443_insert_X, 1046_insert_X, 221_insert_X, D907X, D684X, V1047X, L274X, E157X, W1046X, E479X, R302X, N456X, 1046_drop, I440X, K181X, P2 75X, 1048_drop, S547X, R902X, R233X, A130R, R902X, R433X, V189X, R902X, V1047X, A239X, H901X, V532X, 1045_insert_X, R267X, K1052X, E434X, S548X, S470X, 265_insert_X, R433X, S664R, E244X, R369X, Q209X, R302X, R263X, S625X, R902X, E422X, R267X, L192X, 1047_drop, E328X, V1022X,I199X, D864X, F891X, K281X, T238X, W648X, 609_insert_X, E518X, A239X, V173X, E658X, I414X, K886X, or R433X, where X is any amino acid different from the original amino acid at the indicated position, _insert_ indicates the insertion of any amino acid at the position indicated by the preceding number, and _drop_ indicates the deletion of the amino acid at the position indicated by the preceding number. Preferably, it includes any one or more changes selected from the following positions: E46 2X, D926X, T505X, E494X, D293X, T505X, D709X, E350X, A691X, E736X, E179X, D165X, D704X, E399X, Y292X, K931X, E262X, K703X, E1035X, V842X, D356X, D904X, E388X, K809X, V842X, N879X, S318X, D356X, V842X, H858X, I1048X, H858X, or A854X. ,
[0025] In a preferred embodiment, relative to the number in SEQ ID NO:141, the engineered Cas12 protein comprises a mutation selected from one or more of the following: E462R, D926R, T505K, E494R, D293R, T505R, D709R, E350R, A691R, E736R, E179R, D165R, D704R, E399R, Y292R, K931R, E262R, K703R, E1035R, V842M, E675R, E875R, E601R, E504R, C567R, K277R, D356N, D904R, K563R, E285R, E388R, G478R, E271R, H52 0Y, K809R, L967K, E815R, V842A, V602R, T354S, K138R, I657L, Y468R, N879 T, Q186R, S318F, D356H, I265R, D591R, K404R, V842L, K270R, H858M, E321R , W1046R, K187R, G232D, F891Q, L97F, L13K, L74W, M618L, I1048V, S582R, D 120R, V1047K, S243R, H858R, N168F, E681R, Q620L, K401A, T943R, W927R, E 253R, A854Y, G402W, K774R, W1046S, E793R, E252R, R302G, E447R, V1039H, R330P, D906R, S240R, A624R, T235R, R902G, V1014I, R902E, D851R, F583L, Y537H, 443_insert_A, 1046_insert_D, 221_insert_V, D907R, D684F, V1047D, L274R, E157R, W1046E, E479R, R302A, N456G, 1046_drop, I440M, K181 R, P275K, 1048_drop, S547D, R902H, R233W, A130R, R902D, R433V, V189R, R902I, V1047E, A239R, H901M, V532F, 1045_insert_N, R267L, K1052N, E434 R, S548V, S470R, 265_insert_G, R433M, S664R, E244K, R369Y, Q209H, R302V, R263A, S625P, R902W, E422R, R267G, L192R, 1047_drop, E328R, V1022A,I199L, D864R, F891E, K281R, T238R, W648D, 609_insert_R, E518R, A239C, V173C, E658R, I414L, K886N, or R433I, where _insert_ indicates the insertion of an amino acid at the position indicated by the preceding number, and _drop_ indicates the deletion of an amino acid at the position indicated by the preceding number; preferably, it contains mutations selected from one or more of the following: E462R, D926R, T505K, E494R, D293R, T505R, D709R, E350R, A691R, E736R, E179R, D165R, D704R, E399R, Y292R, K931R, E262R, K703R, E1035R, V842M, D356N, D904R, E388R, K809R, V842A, N879T, S318F, D356H, V842L, H858M, I1048V, H858R, or A854Y.
[0026] In another preferred embodiment, relative to SEQ ID NO:141, the engineered Cas12 protein comprises a combination of mutations selected from one or more of the following:
[0027] E736R, E388R, E462R;
[0028] E462R, A854Y;
[0029] A691R, E388R, T505K;
[0030] E462R, A691R;
[0031] E462R, V842M;
[0032] D926R, A691R, E736R;
[0033] E462R, A691R, E736R, D926R;
[0034] D926R, A691R;
[0035] D704R, I1048V;
[0036] E388R, A854Y;
[0037] D704R, D356H;
[0038] D704R, T505K;
[0039] E388R, D926R;
[0040] A854Y,T505K,A691R,E736R;
[0041] E388R,A691R;
[0042] E388R,I1048V;
[0043] D704R,E388R;
[0044] D926R,E388R,E462R,A691R;
[0045] D704R,A691R;
[0046] D704R,A854Y;
[0047] D704R,D926R;
[0048] E388R,T505K;
[0049] E388K,E462R,E736K;
[0050] E462R,E736D,D926E;
[0051] E388D,E462R,A854Y,D926E;
[0052] E388S,E462R,T505W;
[0053] E462R,E736D,A854Q,D926S;
[0054] E462R,T505M;
[0055] D356C,D926W;
[0056] E462C,D926C;
[0057] E350P,D926C;
[0058] D704C,D926C;
[0059] D356C,D926C;
[0060] T505Y,D704C;
[0061] E462R,T505M,E736D;
[0062] T505C,D926C;
[0063] T505F,D926C;
[0064] A854C, D926C;
[0065] E388D, E462H, T505I, E736D;
[0066] E262L, D926C;
[0067] D356C, A854C;
[0068] E350C, D926C;
[0069] T505W, A854L;
[0070] D356C, E736C;
[0071] T505F, A854L; or
[0072] D356C, A854L;
[0073] Preferably, the engineered Cas12 protein comprises a combination of mutations selected from any of the following:
[0074] E736R, E388R, E462R;
[0075] E462R, A854Y;
[0076] A691R, E388R, T505K;
[0077] E388K, E462R, E736K;
[0078] E462R, E736D, D926E;
[0079] E388D, E462R, A854Y, D926E;
[0080] E388S, E462R, T505W;
[0081] E462R, E736D, A854Q, D926S;
[0082] E462R, T505M; or
[0083] E462R, T505M, E736D.
[0084] In some embodiments, the engineered Cas12 protein described herein includes any one or more variations of the following positions relative to SEQ ID NO:137: D711X, K705X, E738X, S435X, V844X, N233X, D706X, A856X, N371X, N269X, D928X, N168X, E464X, E352X, A693X, E390X, T507X, K933X, G480X, Q951X, or K606X, wherein X is any amino acid different from the original amino acid at the indicated position.
[0085] In a preferred embodiment, the engineered Cas12 protein, relative to the number SEQ ID NO:137, comprises a mutation selected from one or more of the following: D711R, K705R, E738R, S435R, V844M, N233R, D706R, A856Y, N371R, N269R, D928R, N168R, E464R, E352R, A693R, E390R, T507R, K933R, G480R, Q951H, or K606E.
[0086] In some other preferred embodiments, the engineered Cas12 protein, relative to SEQ ID NO:137, comprises:
[0087] (i) A combination of D711R and one or more of the following: S435R, K705R, G480R, N233R, A693R, N371R, V844M, A856Y, E738R, N269R, or E352R; or
[0088] (ii) E738R in combination with one or more of the following: K705R, A856Y, A693R, S435R, E352R, V844M, N371R, G480R, N233R, N269R, or D706R; or
[0089] (iii) Select from any of the following mutation combinations:
[0090] M759V, K776N, H903N, D909E;
[0091] R662H, P716L, K776T, E961D, E966G, L989I, I1023T, M1056L;
[0092] D711G, A796S, F802S, D853E, E967D, R973L, D975V, I980F;
[0093] Y670H, H860Q;
[0094] P638L, V702A, T767A, R904C;
[0095] R812G, Q832H, V984D, G1019C;
[0096] V462F, K606N;
[0097] R662H, P716L, K776T, E961D, E966G, L989I, I1023T, M1056L;
[0098] D711G, A796S, F802S, D853E, E967D, R973L, D975V, I980F;
[0099] F646Y, E790D, R904S, K941I, V1049M;
[0100] S661C, G707S, A752V, R812G; or
[0101] H676Q, H733Q, G787A, E968V, N979I.
[0102] In some embodiments, the engineered Cas12 protein described herein includes any one or more variations of the following positions relative to SEQ ID NO:136: A693X, N233X, S435X, N269X, E352X, N371, E738X, A856X, D711X, or D706X, where X is any amino acid different from the original amino acid at the indicated position.
[0103] In some embodiments, the engineered Cas12 protein, relative to the number SEQ ID NO:136, comprises a mutation selected from one or more of the following: A693R, N233R or N233K or N233S, S435R or S435E or S435K or S435T, N269R or N269E or N269T or N269K, E352R, N371R or N371K, E738D, A856Q or A856S, D711R, or D706R.
[0104] In some embodiments, the engineered Cas12 protein described herein comprises any combination of changes to the following positional groups relative to SEQ ID NO:136:
[0105] A693X, N233X, S435X;
[0106] A693X, N269X, E352X, N371X;
[0107] N233X, N269X, N371X, S435X, E738X, A856X;
[0108] N233X, N371X, S435X, E738X;
[0109] N233X, N269X, S435X, E738X;
[0110] N233X, N269X, S435X, A856X;
[0111] A693X, N233X, E352X;
[0112] N233X, N269X, N371X, S435X, E738X;
[0113] S435X, A693X, D711X;
[0114] N233X, N371X, S435X, E738X, A856X;
[0115] N233X, D706X;
[0116] N233X, N269X, S435X, E738X, A856X; or
[0117] N233X, N269X, N371X, S435X;
[0118] Where X is any amino acid that is different from the original amino acid at the indicated position.
[0119] In a preferred embodiment, the engineered Cas12 protein, relative to SEQ ID NO:136, comprises a combination of mutations selected from any of the following:
[0120] A693R, N233R, S435R;
[0121] A693R, N269R, E352R, N371R;
[0122] N233K, N269T, N371K, S435E, E738D, A856Q;
[0123] N233K, N371R, S435E, E738D;
[0124] N233K, N269E, S435K, E738D;
[0125] N233S, N269R, S435K, A856Q;
[0126] A693R, N233R, E352R;
[0127] N233K, N269E, N371K, S435E, E738D;
[0128] S435R, A693R, D711R;
[0129] N233K, N371K, S435E, E738D, A856S;
[0130] N233R, D706R;
[0131] N233K, N269R, S435E, E738D, A856Q; or
[0132] N233K, N269K, N371K, S435T.
[0133] In some embodiments, the engineered Cas12 protein includes the PI domain shown in any one of SEQ ID NO:327-342. In some embodiments, the engineered Cas12 protein includes the Nuc domain shown in any one of SEQ ID NO:343-356. In some embodiments, the engineered Cas12 protein includes the Wed1 domain shown in any one of SEQ ID NO:358-369.
[0134] In some embodiments, the engineered Cas12 protein described herein has higher activity in cells (e.g., in mammalian cells) compared to the wild-type Cas12 protein, for example, with improved targeting efficiency (gene editing efficiency), lower off-target efficiency, and higher specificity.
[0135] In some embodiments, the engineered Cas12 protein of this disclosure can also be fused with a suitable peptide (e.g., a nuclear localization signal (NLS) sequence) to obtain a fusion protein. In some cases, the peptide exhibits enzymatic activity that modifies a target peptide associated with a target nucleic acid; for example, the peptide exhibits histone modification activity. Optionally, the engineered Cas12 protein is fused with an exonuclease.
[0136] In some cases, the engineered Cas12 protein of the invention can be fused to a peptide penetration domain to facilitate cellular uptake. Many penetration domains are known in the art and can be used with non-integrating peptides of this disclosure, including peptides, peptide mimics, and non-peptide carriers.
[0137] In a second aspect, this disclosure provides a nucleic acid molecule comprising a nucleotide sequence encoding any of the engineered Cas12 proteins of the first aspect.
[0138] In some implementations, the nucleotide sequence can be codon-optimized to suit the codon preferences of the selected host cell.
[0139] In a third aspect, this disclosure provides an engineered CRISPR-Cas system, the system comprising:
[0140] (a) a guide RNA (gRNA) or a nucleic acid sequence encoding the gRNA, wherein the gRNA comprises a direct repeat (DR) sequence and a spacer sequence, wherein the DR sequence is capable of binding to any of the engineered Cas12 proteins described in the first aspect; and
[0141] (b) any engineered Cas12 protein or nucleic acid sequence encoding it as described in the first aspect.
[0142] The engineered Cas12 protein associates with the gRNA.
[0143] The spacer sequence is a nucleotide sequence that is at least partially complementary to the target nucleic acid.
[0144] In some embodiments, the gRNA further comprises a 5' end extension sequence, optionally comprising one or more modifications selected from 2'O-methyl or thiophosphorylation modifications or combinations thereof.
[0145] In some embodiments, the gRNA comprises the sequence shown in any one of SEQ ID NO: 7-10, 118-120, 145-156, and 182 (i.e., crRNA). The gRNA comprises a direct repeat (DR) sequence and a spacer sequence, typically the spacer sequence being complementary to at least 15 nucleotides of the target nucleic acid. In some embodiments, the gRNA comprises the DR sequence shown in SEQ ID NO: 206-209.
[0146] In some embodiments, the guide RNA (gRNA) may contain one or more modifications, for example, such modifications are selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro group, thiophosphorylation modification, or combinations thereof. The modifications may be present in the direct repeat (DR) sequence and / or spacer sequence of the gRNA. In this disclosure, a gRNA containing one or more modifications is referred to as an engineered gRNA.
[0147] In some embodiments, the DR sequence in the gRNA contains one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro group, thiophosphorylation, or combinations thereof. In some embodiments, the 15th, 16th, 3rd, 4th, 1st, or 2nd base of the DR sequence is modified with 2'O-methyl. In some embodiments, the 11th or 12th base of the DR sequence is not modified with 2'O-methyl. In some embodiments, the DR sequence further contains a U7C mutation at position 7. In some embodiments, any one of the 10th, 19th, 8th, 11th, 9th, 21st, 6th, or 20th bases of the DR sequence is fluorinated. In some embodiments, the DR sequence contains: a combination of fluorinated bases at positions 4, 5, 12th, and 15; a combination of fluorinated bases at positions 4, 12th, 14th, and 15; a combination of fluorinated bases at positions 4, 5, 12th, and 14th; or a combination of fluorinated bases at positions 4, 5, 14th, and 15. In some embodiments, the DR sequence is modified through directed evolution. In some embodiments, the DR sequence comprises any one of the sequences shown in SEQ ID NO: 184-186, 187, 189-194, 202-203, 207-230, 223-230, 232-235, 273-283, 406-408, or 412-414.
[0148] In some embodiments, the spacer region in the gRNA contains one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluorine, thiophosphorylation, or combinations thereof. In some embodiments, a base at any of positions 12, 3, 8, 14, or 11 of the spacer region sequence is fluorinated. In some embodiments, bases at positions 11 and 12, 7 and 8, 5 and 6, 3 and 4, 9 and 10, 1 and 2, 15 and 16, or 13 and 14 of the spacer region sequence are fluorinated. In some embodiments, the 3' end of the spacer region sequence contains three 2'O-methyl modified bases and three thiophosphoryl modifications. In some embodiments, the spacer sequence comprises any one of the sequences shown in SEQ ID NO: 197, 205-206, 231, 240-246, 254-257, 263-270, 409-411, or 415-417.
[0149] In some embodiments, the gRNA further comprises a 3' end sequence, optionally comprising the sequence shown in SEQ ID NO:198.
[0150] In some embodiments, the engineered CRISPR-Cas system further includes a nuclear localization signal (NLS) sequence operatively linked to the engineered Cas12 protein, through which the gene-editing complex can be directed to the cell nucleus. In some embodiments, the NLS is an N-terminal NLS and / or a C-terminal NLS.
[0151] In a fourth aspect, this disclosure provides an engineered guide RNA (gRNA) comprising:
[0152] (a) A DNA targeting region containing a nucleotide sequence complementary to the target sequence in the target DNA molecule; and
[0153] (b) A protein-binding segment configured to associate with any of the engineered Cas12 proteins described in the first aspect of this disclosure;
[0154] The engineered guide RNA (gRNA) contains one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro group, thiophosphorylation, or combinations thereof.
[0155] In some implementations, the modifications in the engineered gRNA may be present in the framework region and / or spacer region sequences.
[0156] In some implementations, the engineered gRNA comprises, from 5' to 3':
[0157] (i) A 5' end extension sequence comprising one or more modifications selected from 2'O-methyl or thiophosphorylation modifications or combinations thereof, optionally, the 5' end sequence comprising any one of the nucleotide sequences described in SEQ ID NO:199-201;
[0158] (ii) a direct repeat (DR) sequence comprising one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof.
[0159] Optionally, the 15th, 16th, 3rd, 4th, 1st or 2nd base of the DR sequence is modified with 2'O-methyl, and optionally the 11th or 12th base of the DR sequence is not modified with 2'O-methyl;
[0160] Optionally, the DR sequence further includes a U7C mutation at position 7;
[0161] Optionally, any one of the bases at positions 10, 19, 8, 11, 9, 21, 6, or 20 of the DR sequence is fluorinated; or optionally, the DR sequence comprises: a combination of fluorinated bases at positions 4, 5, 12, and 15; a combination of fluorinated bases at positions 4, 12, 14, and 15; a combination of fluorinated bases at positions 4, 5, 12, and 14; or a combination of fluorinated bases at positions 4, 5, 14, and 15.
[0162] Optionally, the DR sequence is modified through directed evolution; or
[0163] Preferably, the DR sequence comprises any one of the sequences shown in SEQ ID NO: 184-186, 187, 189-194, 202-203, 207-230, 223-230, 232-235, 273-283, 406-408, or 412-414; and
[0164] (iii) A spacer sequence comprising one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof, optionally, a base at any of the 12th, 3rd, 8th, 14th, or 11th positions of the spacer sequence is fluorinated, optionally, a base at the 11th and 12th, 7th and 8th, 5th and 6th, 3rd and 4th, 9th and 10th, 1st and 2nd, 15th and 16th, or 13th and 14th positions of the spacer sequence is fluorinated in combination;
[0165] Optionally, the 3' end of the spacer sequence comprises three 2'O-methyl modified bases and three thiophosphoryl modifications; or
[0166] Optionally, the spacer sequence comprises any one of the sequences shown in SEQ ID NO: 197, 205-206, 231, 240-246, 254-257, 263-270, 409-411 or 415-417;
[0167] Optionally, the engineered gRNA further comprises a 3' end sequence, which optionally comprises the sequence shown in SEQ ID NO:198.
[0168] In a fifth aspect, this disclosure provides a recombinant expression system that expresses the engineered CRISPR-Cas system described in the third invention of this disclosure, the recombinant expression system comprising:
[0169] (i) a nucleic acid sequence encoding any of the engineered Cas12 proteins described in the first aspect of this disclosure; and
[0170] (ii) Nucleic acid sequences that encode DNA target sequences.
[0171] In some embodiments, the DNA targeting sequence comprises a guide RNA (gRNA) capable of hybridizing with the target sequence. In some embodiments, the gRNA comprises a sequence shown in any one of SEQ ID NO: 7-10, 118-120, 145-156, and 182 (i.e., crRNA). The gRNA comprises a direct repeat (DR) sequence and a spacer sequence, typically the spacer sequence being complementary to at least 15 nucleotides of the target nucleic acid.
[0172] In some embodiments, the guide RNA (gRNA) may contain one or more modifications, for example, such modifications are selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof. Therefore, in some embodiments, the gRNA is an engineered gRNA containing one or more modifications, for example, the engineered gRNA containing any one of the nucleotide sequences (i.e., DR sequences) described in SEQ ID NO: 184-186, 187, 189-194, 202-203, 207-230, 223-230, 232-235, 273-283, 406-408, or 412-414. In other embodiments, the gRNA contains any one of the spacer sequences shown in SEQ ID NO: SEQ ID NO: 197, 205-206, 231, 240-246, 254-257, 263-270, 409-411, or 415-417.
[0173] In some embodiments, the recombinant expression system comprises a recombinant expression vector in which the nucleic acid sequence encoding the engineered Cas12 protein and the nucleic acid sequence encoding the DNA target sequence are constructed in the vector (e.g., a plasmid). In some cases, the nucleic acid sequence encoding the engineered Cas12 protein and the nucleic acid sequence encoding the DNA target sequence are constructed in different vectors or in the same vector.
[0174] The recombinant expression vector may be a viral vector, for example, selected from, but not limited to, retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, and herpes simplex vectors. In some cases, the recombinant expression vector of this disclosure is a recombinant adeno-associated virus (AAV) vector. In some cases, the recombinant expression vector of this disclosure is a recombinant lentiviral vector. In some cases, the recombinant expression vector of this disclosure is a recombinant retroviral vector.
[0175] Furthermore, depending on the host / vector system used, any of a variety of suitable transcription and translation control elements can be used in the expression vector, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc.
[0176] In some embodiments, the nucleotide sequence encoding the guide RNA is operatively linked to a control element, such as a transcriptional control element, like a promoter. In some embodiments, the nucleotide sequence encoding an engineered Cas12 protein or Cas12 fusion polypeptide is operatively linked to a control element, such as a transcriptional control element, like a promoter.
[0177] In a sixth aspect, this disclosure provides a recombinant host cell comprising: any engineered Cas12 protein or nucleic acid molecule encoding it as described in the first aspect of this disclosure, an engineered CRISPR-Cas system as described in the third aspect, an engineered gRNA as described in the fourth aspect, or a recombinant expression system as described in the fifth aspect.
[0178] In some embodiments, the recombinant host cell may be a prokaryotic or eukaryotic cell, such as, but not limited to, bacterial cells, fungal cells, plant cells, insect cells, or animal cells, such as Escherichia coli cells, yeast cells, or mammalian cells, preferably mammalian cells, such as human cells or non-human primate cells. Typically, the cell does not include plant or animal cells capable of developing into a complete individual.
[0179] In a seventh aspect, this disclosure provides a pharmaceutical composition comprising the engineered CRISPR-Cas system described in the third aspect, the recombinant expression system described in the fifth aspect, or the cells described in the sixth aspect, and a pharmaceutically acceptable carrier.
[0180] In an eighth aspect, this disclosure provides a kit comprising:
[0181] (a) any engineered Cas12 protein or nucleic acid encoding it as described in the first aspect of this disclosure, the engineered CRISPR-Cas system as described in the third aspect, the engineered gRNA as described in the fourth aspect, the recombinant expression system as described in the fifth aspect, or the cell as described in the sixth aspect; and
[0182] (b) Instructions for use.
[0183] In a ninth aspect, this disclosure provides a method for editing nucleosides in a target nucleic acid sequence, the method comprising: contacting the target nucleic acid sequence with an engineered CRISPR-Cas system as described in a third aspect of this disclosure or a recombinant expression system as described in a fifth aspect, wherein a DNA targeting sequence binds to the target nucleic acid sequence, thereby editing nucleosides in the target nucleic acid sequence.
[0184] In some embodiments, the method is performed in vitro or in vivo. In some embodiments, the method can be used for therapeutic purposes or for non-therapeutic purposes.
[0185] In some embodiments, this disclosure provides the use of any of the engineered Cas12 proteins or nucleic acids encoding them described in the first aspect, the engineered CRISPR-Cas system described in the third aspect, the engineered gRNA described in the fourth aspect, the recombinant expression system described in the fifth aspect, or the cells described in the sixth aspect in the preparation of kits for editing nucleobases in target nucleic acid sequences.
[0186] In some implementations, the editing includes mutations in a single nucleobase in the target nucleic acid sequence, wherein the mutations include insertions, deletions, or substitutions.
[0187] In some implementations, the editing includes gene sequence insertion, gene sequence deletion, or gene sequence replacement.
[0188] In some implementations, the editing includes epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification includes methylation or demethylation of the target nucleic acid sequence.
[0189] In some implementations, the epigenetic modification regulates gene expression by activating or repressing target nucleic acid sequences. In some cases, regulating gene expression includes upregulating or downregulating gene expression.
[0190] In a tenth aspect, this disclosure provides a method for treating a disease or condition related to a target nucleic acid in an individual's cells, the method comprising using the method described in the ninth aspect to edit the target nucleic acid in the individual's cells, thereby treating the disease or condition.
[0191] In some embodiments, this disclosure provides the use of any of the engineered Cas12 proteins or nucleic acids encoding them described in the first aspect, the engineered CRISPR-Cas system described in the third aspect, the engineered gRNA described in the fourth aspect, the recombinant expression system described in the fifth aspect, or the cells described in the sixth aspect in the preparation of medicaments or kits for treating diseases or conditions related to target nucleic acids in the cells of an individual.
[0192] In some implementations, the disease or condition is selected from: cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
[0193] In some implementations, the individual is a mammal, preferably a human.
[0194] In the eleventh aspect, this disclosure provides a modular method for engineering Cas12 protein, the method comprising: designing variant sequences as module sequences for each of the Helical-I, Helical-II, Helical-III, NUC, PI and WED-II domains of wild-type Cas12 protein, replacing the corresponding wild-type domains with one or more module sequences, and assembling to obtain engineered Cas12 protein.
[0195] In some embodiments, the method further includes screening the resulting engineered Cas12 proteins to select engineered Cas12 proteins that have improved editing efficiency and / or reduced off-target efficiency.
[0196] In a twelfth aspect, this disclosure provides a method for preparing engineered Cas12 proteins, the method comprising a chemical synthesis method or a recombinant expression method.
[0197] In a thirteenth aspect, this disclosure provides lipid nanoparticles (LNPs) that encapsulate nucleic acid molecules encoding the engineered Cas12 protein described in the first aspect and engineered gRNA described in the fourth aspect, optionally encapsulating the nucleic acid molecules and engineered gRNA in a 1:1 mass ratio; or encapsulating the engineered CRISPR-Cas system described in the third aspect.
[0198] The engineered Cas12 protein described in this disclosure exhibits higher activity in cells (e.g., in mammalian cells) compared to the wild-type Cas12 protein, for example, with improved targeting efficiency (gene editing efficiency), lower off-target efficiency, and higher specificity. These advantages make the engineered Cas12 protein of this application and the CRISPR-Cas system containing it highly suitable for efficient and specific gene editing or gene regulation in vivo (e.g., in mammalian cells), and can be used to treat diseases or conditions related to target nucleic acids. Attached Figure Description
[0199] The embodiments and advantages of this disclosure will become more apparent when viewed in conjunction with the following figures.
[0200] Figure 1 shows the results of the Cas12-1 variant activity screening. (A) is the control group experiment conducted on agar plates containing ampicillin (Amp), and (B) is the experimental group experiment conducted on agar plates containing ampicillin and arabinose (Amp + arabinose). The detected Cas12-1 variants are shown in the table below the image according to the colony positions on the plates.
[0201] Figure 2 shows the targeting efficiency of the Cas12-1 variant in cells.
[0202] Figure 3 shows the results of activity verification of the top 10 most frequent recombinant variants (SEQ ID NO:95-SEQ ID NO:104, see Table 2) using the E. coli positive selection system. (A) Control group experiment performed on agar plates containing ampicillin (Amp), (B) Experimental group experiment performed on agar plates containing ampicillin and arabinose (Amp+arabinose). The detected Cas12-1 variants are shown in the table below the image according to the colony positions on the plates.
[0203] Figure 4 shows the targeting efficiency of the Cas12-1 variant in HEK293T cells.
[0204] Figure 5 shows the targeting efficiency of the Cas12 variant in HEK293T cells when using different gRNAs and DR (Direct Repeat) sequences.
[0205] Figure 6 shows the targeting efficiency of the Cas12 variant in HEK293T cells.
[0206] Figure 7 shows the targeting efficiency of the Cas12 variant in HEK293T cells.
[0207] Figure 8 shows the PAM-preferred base sequence of the Cas12 variant obtained by PCR library preparation and NGS sequencing (image generated by weblogo software). Detailed Implementation
[0208] definition
[0209] Unless otherwise specified, all technical and scientific terms used in this disclosure shall have the meanings commonly understood by one of ordinary skill in the art to which this disclosure pertains. Throughout the specification, several terms as defined in the following paragraphs are used. Other definitions can also be found in the text of the specification.
[0210] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” or “polynucleotide” are used interchangeably. They refer to polymers of deoxyribonucleotides or ribonucleotides, or mixtures thereof, in single-stranded or double-stranded form, and, unless otherwise stated, also encompass known analogs of natural nucleotides that can function in a similar manner to naturally occurring nucleotides. These terms cover nucleic acid-like structures with a synthetic backbone as well as amplification products. Both DNA and RNA are polynucleotides. The polymer may include natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deoxyuridine, etc.). Adenosine, 7-deazoguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose and hexose) or modified phosphate groups (e.g., thiophosphates and 5'-N-phosphoramide bonds).
[0211] As used herein, the terms “peptide” and “protein” are used interchangeably and refer to a polymer having an amino acid of any length. The polymer can be linear or branched, it can contain modified amino acids, and it can be interrupted by non-amino acid components. The term also covers modified amino acid polymers; for example, through disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeled component.
[0212] As used herein, the term "nuclease" refers to a polypeptide that can cleave phosphodiester bonds between nucleotide subunits of nucleic acids; the term "endonuclease" refers to a polypeptide that can cleave phosphodiester bonds within a polynucleotide chain.
[0213] As used herein, “guide RNA” and “gRNA” are used interchangeably and refer to RNA capable of forming a complex with the Cas12 protein and the target nucleic acid (e.g., double-stranded DNA), guiding Cas12 to recognize and cleave the target gene sequence. This paper also considers precursor guide RNA arrays that can be processed into multiple crRNAs. “crRNA” or “CRISPR RNA” contains a sequence sufficiently complementary to the target sequence of the target nucleic acid (e.g., double-stranded DNA) to guide the CRISPR complex to bind specifically to the target sequence of the target nucleic acid.
[0214] As used herein, the term "CRISPR array" refers to a nucleic acid (e.g., DNA) segment comprising CRISPR repeats and spacer regions, beginning with the first nucleotide of the first CRISPR repeat and ending with the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer region in a CRISPR array lies between two repeats. As used herein, the terms "CRISPR repeat," "CRISPR direct repeat," or "direct repeat (DR)" refer to multiple short direct repeat sequences that exhibit very little or no sequence variation in the CRISPR array. Appropriately, direct repeats may form stem-loop structures.
[0215] In the early days of CRISPR, the function of gRNA was achieved by two sequences, crRNA and tracrRNA, working together. Later, through modification, crRNA (guide RNA) and tracrRNA (trans-activating RNA) were fused together to form an integrated single gRNA, called sgRNA. crRNA typically consists of a direct repeat (DR) sequence and a spacer region sequence.
[0216] As used herein, “variant” is defined as a polynucleotide or polypeptide that differs from a reference polynucleotide or polypeptide but retains the necessary characteristics. A typical variant of a polynucleotide differs from the nucleic acid sequence of another reference polynucleotide. Changes in the nucleic acid sequence of a variant may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. Nucleotide changes can result in amino acid substitutions, additions, deletions, fusions, and truncations in the polypeptide encoded by the reference sequence, as described below. A typical variant of a polypeptide differs from another reference polypeptide in its amino acid sequence. Typically, the differences are limited, making the sequences of the reference polypeptide and the variant very similar overall and identical in many regions. The amino acid sequences of the variant and the reference polypeptide can differ by any combination of one or more substitutions, additions, or deletions. The substituted or inserted amino acid residues may or may not be amino acid residues encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring (such as allelic variants) or may be variants of unknown natural origin. Non-natural variants of polynucleotides and polypeptides can be prepared by mutagenesis, by direct synthesis, and by other recombinant methods known to those skilled in the art.
[0217] As used herein, the term "wildtype" has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, gene, or trait that distinguishes it from mutants or variants when it exists in nature. It can be isolated from resources in nature and is not deliberately modified.
[0218] As used herein, the terms “non-naturally occurring” and “engineered” are used interchangeably. When these terms are used to describe a nucleic acid molecule or polypeptide / protein, it means that the nucleic acid molecule or polypeptide / protein is at least substantially free of at least one other component that is naturally associated with or naturally occurring.
[0219] As used herein, the term "identity" is used to indicate sequence matching between two polypeptides or two nucleic acids. Two compared sequences are considered identical at that position when a position is occupied by the same base or amino acid monomer subunit (e.g., a position in each of two DNA molecules occupied by adenine, or a position in each of two polypeptides occupied by lysine). The "percentage of identity" between the two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions to be compared x 100. For example, if six out of ten positions in two sequences match, the two sequences have 60% identity. For example, the DNA sequences CGGACT and CAGGTT have 50% identity (three out of six positions match). Typically, such comparisons are performed when two sequences are aligned to produce the greatest possible identity. Such alignments can be performed, for example, by the method described in Needleman et al. (1970) J. Mol. Biol. 48: 443-453, which can be conveniently performed using computer programs such as the Align program (DNAstar, Inc.). Alternatively, a PAM 120 weighted residue table can be used, with the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4: 11-17 (1988)) integrated into the ALIGN program (version 2.0). A gap length penalty of 12 and a gap penalty of 4 are used to determine the percentage of identity between two amino acid sequences. Alternatively, the Needleman and Wunsch algorithm (J MoI Biol. 48: 444-453 (1970)) integrated into the GCG software package (available at www.gcg.com) can be used, employing a Blossum 62 matrix or a PAM250 matrix, with gap weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6, to determine the percentage of identity between two amino acid sequences.
[0220] The term "cell" as used herein should be understood not only to a specific single cell, but also to its offspring or potential offspring. Because certain modifications may occur in offspring due to mutations or environmental influences, such offspring may indeed differ from the parent cell, but are still included within the scope of this terminology.
[0221] The term "in vivo" refers to the organism from which cells are obtained. "Ex vivo" or "in vitro" refers to the organism from which cells are obtained outside of the organism.
[0222] As used herein, “treatment” is a method for obtaining a beneficial or desired outcome (including clinical outcomes). For the purposes of this disclosure, a beneficial or desired clinical outcome includes, but is not limited to, one or more of the following: relief of one or more symptoms caused by a disease; reduction of the severity of the disease; stabilization of the disease (e.g., prevention or delay of disease progression); prevention or delay of disease spread (e.g., metastasis); prevention or delay of disease recurrence; reduction of the recurrence rate of the disease; delay or slowing of disease progression; improvement of disease status; provision of (partial or complete) remission of the disease; reduction of the dosage of one or more other medications required to treat the disease; delay of disease progression; improvement of quality of life; and / or prolongation of survival. “Treatment” also includes reduction of symptoms, condition, or pathological consequences of the disease. The methods of this disclosure consider any one or more of these aspects of treatment.
[0223] The terms “individual,” “subject,” or “patient” are used interchangeably herein. For therapeutic purposes, it refers to any animal classified as a mammal, including humans, livestock, and farm animals, as well as zoo, farm, or pet animals such as dogs, horses, cats, and cattle. In some embodiments, the individual is a human individual.
[0224] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Those skilled in the art should understand that the embodiments are for illustrative purposes only and do not limit the scope of protection of this disclosure in any way. Based on the embodiments in this disclosure, those skilled in the art can determine that equivalent implementations of the embodiments are also within the scope of protection of this disclosure.
[0225] Those skilled in the art should also understand that, unless otherwise stated, the cells, plasmids, reagents, etc. used in the examples are commercially available.
[0226] Example
[0227] Example 1: Optimization of Cas12-1 protein domains using computational protein design techniques
[0228] To improve the targeting efficiency of the Cas12-1 protein (SEQ ID NO:2) (Table 1-1), we redesigned the sequences of its six domains—Helical-I (SEQ ID NO:11), Helical-II (SEQ ID NO:12), Helical-III (SEQ ID NO:13), NUC (SEQ ID NO:14), PI (SEQ ID NO:15), and WED-II (SEQ ID NO:16)—using computational protein design techniques (Table 1-3). Ten new sequences were generated for each domain, totaling 60 (SEQ ID NO:17-SEQ ID NO:76) (Table 1-3). These new sequences, after codon optimization, replaced the corresponding sequences in the original wild-type domains using Gibson assembly technology. To verify the activity of the variants, we designed a forward screening system for E. coli, which includes a Cas protein expression plasmid controlled by the Tet-on promoter (SEQ ID NO:1) and a plasmid co-expressing gRNA and inducible ccdB protein (SEQ ID NO:6) (Table 1-1). This system utilizes Cas12-mediated cleavage of the plasmid expressing the bacterial virulence-encoding gene ccdB to relieve the lethal effect of ccdB expression on bacteria, thereby indicating whether the Cas protein has cleavage activity.
[0229] The specific procedure is as follows: Approximately 50 ng of the Cas12 variant plasmid (ampicillin (Amp) resistant) and 50 ng of the co-expression plasmid encoding ccdB toxin protein and gRNA (chloramphenicol resistant) were co-transformed into *E. coli* Top10 strains. After chemical transformation, the culture was incubated in 200 μL LB medium at 37°C for 1 hour, followed by induction with dehydrotetracycline at a concentration of 100 ng / μL for 1 hour to activate Cas12 protein expression. Then, 2 μL of bacterial culture was inoculated onto two agar plates: one containing ampicillin (control group, Figure 1A), and the other containing arabinose and ampicillin (experimental group, Figure 1B). Variants showing colony growth on the agar plate containing arabinose were considered to be Cas12-1 variants with cleavage activity. The Cas12-1 variant activity screening results are shown in Figure 1.
[0230] To further verify the targeting efficiency of the Cas12-1 variants (SEQ ID NO:83-SEQ ID NO:94) (Table 2-1), which exhibit high activity in bacteria, in HEK293T cells, we designed three genomic in situ targeting gRNAs (gRNA4, gRNA5, and gRNA6, sequences SEQ ID NO:118-SEQ ID NO:120, respectively) (Table 2-2). Using Fugene HD transfection reagent, 88 ng of the Cas12-1 variant plasmid and a mixture of three gRNA expression plasmids (24 ng / plasm) were co-transfected into HEK293T cells (96-well plates, 2×10⁶ cells / well). 4 (HEK293T cells were purchased from the Cell Bank of the Chinese Academy of Sciences Type Culture Collection Committee). After 72 hours of culture, the cells were lysed to extract genomic DNA, and then the intracellular targeting efficiency of the Cas12-1 variant was verified by PCR library construction and NGS sequencing (Figure 2).
[0231] The results in Figure 2 indicate that the Cas12-1 variants (SEQ ID NO:83-SEQ ID NO:94) generated by computation and screened experimentally can achieve high intracellular cleavage activity.
[0232] Example 2: Computational protein design for sequence recombination of Cas12-1 protein domains
[0233] Based on Example 1, we randomly combined the six optimized structural domains to construct a theoretical diversity of 1×102 6 The plasmid library was used. This library was introduced into *E. coli* Top10 electroporated competent cells (electroplated competent cells containing plasmids co-expressing gRNA and inducible ccdB protein) for active variant selection. After electroporation, the cells were incubated in 10 ml LB medium at 37°C for 1 hour, followed by induction with 100 ng / μL dehydrotetracycline for 1 hour. Positive colonies on arabinose-containing plates were then collected, plasmids were extracted, and three rounds of repeated electroporation selection were performed before Sanger sequencing. The top 10 recombinant variants with the highest sequencing frequency (SEQ ID NO: 95-SEQ ID NO: 104) (Table 2-1) were then used for activity verification using an *E. coli* positive selection system. The experimental results are shown in Figure 3.
[0234] To further validate and optimize the targeting efficiency of the aforementioned Cas12-1 variants, we performed intracellular targeting efficiency verification on the variants (SEQ ID NO:95-SEQ ID NO:100) (Table 2-1). We then used two variants, Cas12-1-top1 (SEQ ID NO:95) and Cas12-1-top3 (SEQ ID NO:97), to generate new Cas12-1 variants (SEQ ID NO:105-SEQ ID NO:114) by restoring the calculated protein-designed domain sequences one by one to the wild-type (WT) sequences (Table 2-1). Furthermore, we introduced four arginine mutations (D233R, D267R, N369R, S433R) (SEQ ID NO:115-SEQ ID NO:116) into the two Cas12-1 variants (Table 2-1). To evaluate the editing activity of these variants, we performed targeting efficiency verification in HEK293T cells, following the experimental methods described in Example 1. The experimental results are shown in Figure 4.
[0235] The results in Figure 4 indicate that the Cas12-1 recombinant variants, after bacterial screening and engineering, can achieve high intracellular cleavage activity.
[0236] Example 3: Recombination of Cas12 homologous protein domains
[0237] We divided the domains of four naturally occurring Cas12 homologous proteins, Cas12-1, Cas12-2, Cas12-3, and Cas12-4 (SEQ ID NO:2-SEQ ID NO:5, Table 1-1), into six regions (WED-I & Helical-I: region 1, PI: region 2, Helical-I (downstream) & Helical-II: region 3, WED-II: region 4, RuvC-I & Helical-III & BH: region 5, RuvC-II & Nuc & RuvC-III: region 6) (SEQ ID NO:157-SEQ ID NO:178, Table 3-3), and generated a library with a theoretical diversity of 4096 through random recombination by PCR and Gibson assembly. After bacterial screening, we obtained six Cas12 variants (SEQ ID NO:121-SEQ ID NO:126, Table 3-1), the main feature of which is the substitution of the PI domain (SEQ ID NO:15, SEQ ID NO:77-SEQ ID NO:78, Table 1-3). To verify the intracellular targeting efficiency of the above Cas12 variants, we designed three gRNAs, namely gRNA1, gRNA2, and gRNA3, and used four DR (Direct Repeat) sequences corresponding to Cas12 (SEQ ID NO:145-SEQ ID NO:156, Table 3-2). The cell transfection and NGS detection procedures are as described in Example 1, and the experimental results are shown in Figure 5. The results in Figure 5 show that, through the exchange of the PI domain, the variants Cas12-1, Cas12-2, and Cas12-3 can maintain high intracellular targeting efficiency, among which the variants Cas12-1-PIV2 and Cas12-3-PIV2 have better targeting efficiency than their wild-type counterparts.
[0238] Meanwhile, we exchanged the Nuc domains (SEQ ID NO:14, SEQ ID NO:79-SEQ ID NO:81, Tables 1-3) of four naturally occurring Cas12 proteins, and then co-transfected these variants (SEQ ID NO:127-SEQ ID NO:132, Table 3-1) with gRNA1, gRNA2, and gRNA3 (SEQ ID NO:145-SEQ ID NO:156, Table 3-2) into HEK293T cells. The cellular targeting efficiency of these recombinant Cas12 nucleases was verified by NGS analysis. The cell transfection and NGS detection procedures are described in Example 1, and the experimental results are shown in Figure 6. The results in Figure 6 indicate that exchanging the Nuc domains of the other three proteins Cas12-2, Cas12-3, and Cas12-4 on the Cas12-1 protein backbone (SEQ ID NO:79-SEQ ID NO:81) can maintain high intracellular targeting efficiency. Among them, Cas12-1-NucV2 and Cas12-1-NucV3 significantly improved intracellular targeting efficiency compared to wild-type proteins.
[0239] Example 4: Engineering modifications further enhance the activity of the Cas12 variant.
[0240] We constructed a series of novel Cas12 variants (SEQ ID NO:133-SEQ ID NO:144, Table 3-1) by combining computationally designed protein sequences and recombinant domains on the Cas12-1 protein. The targeting efficiency of these variants in HEK293T cells was evaluated using three gRNA targets (gRNA4, gRNA5, and gRNA6, SEQ ID NO:118-SEQ ID NO:120, Table 2-2), with experimental methods as described in Example 1. The results are shown in Figure 7.
[0241] The results in Figure 7 indicate that the further modified Cas12 variants (SEQ ID NO:133-SEQ ID NO:144) have significantly better intracellular targeting efficiency than wild-type Cas12-1.
[0242] Example 5: PAM analysis of engineered Cas12 variants
[0243] First, we designed a plasmid library targeting the Cas12 and gRNA ribonucleoprotein complex (SEQ ID NO:181, Table 4-1). The target sequence in this plasmid library consists of 6 random bases upstream of the 5' end (5'-NNNNNN-3', where N represents any nucleotide) (the plasmid library serves as an in vitro cleavage substrate for the Cas12 variant). Then, using Fugene HD, we co-transfected the Cas variant expression plasmid (80 ng) (Cas variant sequences are SEQ ID NO:133-SEQ ID NO:141, Table 3-1) and the gRNA expression plasmid (40 ng) (gRNA sequence is SEQ ID NO:182, Table 4-2) into HEK293T cells (96-well plates, 2 × 10⁻⁶ cells per cell line). 4 (Number of cells). After 48 hours of culture, cells were lysed to obtain the Cas12 and gRNA ribonucleoprotein complex. The plasmid library was then used to perform in vitro cutting experiments with the ribonucleoprotein complex, and the PAM-biased base sequences of the Cas12 variants were obtained by PCR library construction and NGS sequencing. Images were then generated using weblogo software. The experimental results are shown in Figure 8. Figure 8 shows that the engineered Cas12 variant exhibits a relatively clear 5'-NTTN-3' or 5'-NATN-3' PAM bias.
[0244] Example 6: Methoxylation of gRNA direct repeat (DR) sequence and modification of gRNA 5' / 3' end and functional verification
[0245] We chemically modified the gRNA sequence of Cas12-1 to enhance its intracellular stability and thus improve its intracellular targeting efficiency. First, we selected a gRNA, QKS-g325 (SEQ ID NO:120, Table 2-2), and performed methoxy (2'-OMe) scanning modification on the direct repeat (DR) sequence (SEQ ID NO:184-186 and SEQ ID NO:188-196, Table 5; SEQ ID NO:187 without methoxy modification served as a control). We used Lipofectamine MessengerMAX transfection reagent and divided the mRNA:gRNA transfection dose into three gradients: 100 ng:100 ng, 25 ng:25 ng, and 6.25 ng:6.25 ng, respectively, and co-transfected Huh-7 cells (96-well plates, 8 × 10⁶ cells per cell). 3(Huh-7 cells were purchased from the Cell Bank of the Chinese Academy of Sciences Type Culture Collection Committee). After 72 hours of culture, the cells were lysed to extract genomic DNA, and then the activity of Cas12-1 gRNAs with different chemical modifications in the cells was verified by PCR library construction and NGS sequencing (Table 5). The experimental results showed that: (1) modification of the 15th, 16th, 3rd, 4th, 1st and 2nd methoxy bases in the DR sequence of gRNA significantly improved the intracellular targeting efficiency (SEQ ID NO: 184-186); (2) when the 11th or 12th or 21st base in the DR sequence was modified by methoxy (Seq ID No: 195-196), the targeting efficiency was greatly reduced, indicating that the 11th or 12th or 21st base in the DR sequence cannot be modified by methoxy.
[0246] We then combined the highly active methoxy modifications from the above experiments and added chemically modified auxiliary sequences (SEQ ID NO:199-SEQ ID NO:201) to the 5' end of the gRNA, and modified the last three bases of the 3' end with three thio and three methoxy groups, and added a U7C mutation at position 7 of the DR sequence (SEQ ID NO:203) (Table 6) to better maintain the stability of the gRNA in cells or mouse liver. Cell transfection and gRNA activity verification methods were as described above, and the experimental results are shown in Table 6, Experiments 1 and 2. The experimental results showed that adding chemically modified auxiliary sequences (SEQ ID NO:199-SEQ ID NO:201) to the 5' end of the gRNA and modifying the last three bases of the 3' end of the gRNA with three thio and three methoxy groups (SEQ ID NO:206) significantly improved intracellular targeting activity.
[0247] To confirm the activity of the selected modified and optimized gRNA combination in mouse liver, we selected Cas12-ART9 mRNA (encoding the amino acid sequence SEQ ID NO:141) and the modified gRNA imouse-g10 (SEQ ID NO:297) for downstream testing. Nucleic acids were packaged in lipid nanoparticles (LNPs) using a formulation of ALC-0315 (converted according to the method described in patent application WO / 2023 / 185697A2, which combines Cas12-ART9 mRNA and modified gRNA (Finn et al., 2018) and encapsulated within the LNPs) at a mass ratio of 1:1 (mRNA and gRNA, with LNP doses of 1 mg / kg body weight (mpk) and 0.5 mpk, respectively). Mice were injected via tail vein. Liver tissue DNA was extracted 7 days later, and the targeting activity of different modified gRNAs (Seq ID No: 284-Seq ID No: 288) in mouse liver was verified by PCR library construction and NGS sequencing. Experimental results showed that adding a chemically modified auxiliary sequence at the 5' end of the gRNA (Seq ID No: 287), modifying the last three bases of the 3' end with three thio and three methoxy groups (Seq ID No: 285), and the DR sequence with a U7C mutation at position 7 (Seq ID No: 288) were effective in targeting different types of gRNAs. No:288) all contribute to improving the targeting activity of gRNA in mouse liver (Table 12).
[0248] To further test the effect of adding a chemically modified helper sequence (Seq ID No:201) to the 5' end of gRNA on enhancing the activity of different Cas12 cells, we selected Cas12-1 (SEQ ID NO:2), Cas12-2 (SEQ ID NO:3), Cas12-3 (SEQ ID NO:4), and Cas12-4 (SEQ ID NO:5), with corresponding gRNA spacer sequences of SEQ ID NO:206, SEQ ID NO:409, SEQ ID NO:410, and SEQ ID NO:411, respectively, for cellular level assays. Lipofectamine MessengerMAX transfection reagent was used. We divided the mRNA:gRNA transfection dose into three gradients: 100 ng:100 ng, 10 ng:10 ng, and 1 ng:1 ng, and co-transfected Huh-7 cells (96-well plates, 8 × 10⁶ cells per cell line) with the corresponding mRNA:gRNA doses of 100 ng:100 ng, 10 ng:10 ng, and 1 ng:1 ng. 3(cells). After 72 hours of culture, the genomic DNA of the cells was lysed, followed by PCR library construction and NGS sequencing. The experimental results showed that the addition of chemically modified auxiliary sequences (SEQ ID NO:201) to the 5' end of different Cas12 gRNAs significantly improved their targeting activity in huh-7 cells (Table 6, Experiment 3).
[0249] Example 7: Fluorination modification of gRNA direct repeat sequences (DR)
[0250] Since 2′-F modification may increase the interaction between gRNA and Cas nuclease or reduce the degradation of nucleases in cells or in vivo, thereby improving the targeting activity of the CRISPR system in cells or in vivo, we selected imouse-g10 (SEQ ID NO:297) to perform fluorination modification scans on the direct repeat (DR) sequences of gRNA (SEQ ID NO:207-221 and 223-230) (Table 7, where SEQ ID NO:222 without fluorination modification is used as a control). The experiment used Lipofectamine MessengerMAX transfection reagent. We used Cas12-ART9 mRNA and the fluorinated gRNA imouse-g10, dividing the mRNA:gRNA transfection dose into three gradients: 100ng:100ng, 10ng:10ng, and 1ng:1ng, respectively, and co-transfected Hepa1-6 cells (96-well plates, 8×10⁶ cells per cell). 3 (Number of cells). After 72 hours of culture, the genomic DNA of the cells was lysed, and the targeting activity of gRNAs with different fluorination modifications in the cells was verified by PCR library construction and NGS sequencing (Table 7). The experimental results in Table 7 show that fluorination modification at any position of DR sequence 15, 5, 14, 23, 12, 4, 3, 2, 1, 22, 17, 18, 7, 13 or 16 (SEQ ID NO: 207-221) significantly increased the intracellular targeting activity, while fluorination modification at any position of DR sequence 10, 19, 8, 11, 9, 21, 6 or 20 (SEQ ID NO: 223-230) basically maintained the intracellular targeting activity.
[0251] To further increase the number of fluorinated bases in the DR sequence (SEQ ID NO:232-237) and thus enhance the intracellular stability of gRNA, we selected two gRNAs, imouse-g10 (SEQ ID NO:297) and imouse-g18 (SEQ ID NO:298), and performed multi-fluorinated combination modifications on the better-performing single-fluorinated modifications. The experimental methods are as described above, and the results are shown in Table 8. The experimental results in Table 8 show that the combined fluorinated base modifications at positions 4, 5, 12, and 15 of the gRNA DR sequence (SEQ ID NO:232), 4, 12, 14, and 15 (SEQ ID NO:233), 4, 12, and 14 (SEQ ID NO:234), and 4, 5, 14, and 15 (SEQ ID NO:235) can all improve the total editing activity in cells.
[0252] Example 8: Fluorination modification of gRNA spacer region
[0253] To further enhance gRNA activity, we selected gRNA, imouse-g18 (SEQ ID NO: 298), and performed a fluorination modification scan on the gRNA spacer region (SEQ ID NO: 240-244) (Table 9, where SEQ ID NO: 245 is unfluorinated and serves as a control). The experimental method was the same as in Example 7. The experimental results are shown in Table 9. The results indicate that fluorination modification at any position (12, 3, 8, 14, or 11) of the gRNA spacer region can improve intracellular editing activity.
[0254] To further increase the number of fluorinated molecules in the spacer region and thus enhance the stability of intracellular gRNA, we selected two gRNAs, imouse-g10 (SEQ ID NO:297) and imouse-g18 (SEQ ID NO:298), and combined various single fluorinated modifications. The experimental method is as described in Example 7, and the experimental results are shown in Table 10. The results show that combined fluorinated modifications at positions 11 and 12, 7 and 8, 5 and 6, 3 and 4, 9 and 10, 1 and 2, 15 and 16, or 13 and 14 in the gRNA spacer region (SEQ ID NO:263-270) can improve the editing activity of gRNA in cells.
[0255] To confirm that fluorination of the spacer region helps improve targeting activity in mouse liver, we designed six modifications (1-6 fluorinations) for the spacer region (Seq ID No: 289-293) and delivered mRNA and gRNA with different numbers of fluorinations via LNP. The experimental method is the same as in Example 6. The results showed that the targeting activity of the six different fluorinated gRNAs in mouse liver was superior to that of the unfluorinated gRNA (Table 13).
[0256] To confirm that the combination of fluorination modification of the spacer region and fluorination modification of the DR sequence helps to improve the targeting activity in mouse liver, we selected preferred fluorination modifications of the spacer region and the DR sequence for combined modification (Seq ID No: 294-296). The mouse experimental protocol is described in Example 6. The experimental results showed that the combined modification of the spacer region and the DR sequence significantly improved the targeting activity in mouse liver (Table 14).
[0257] Example 9: Backbone Modification of gRNA Direct Repeat (DR) Sequences
[0258] To further optimize the gRNA, we modified the gRNA DR sequence backbone, with the spacer sequence being QKS-354-15 (SEQ ID NO:299). A total of 137 sequences were designed by performing saturation substitutions, insertions, and deletions (by position) on 23 bases of the DR sequence. After sequence synthesis, these sequences were inserted (left homologous arm - GGATCTTTGACAGCTAGCTCAGTCC, right homologous arm - CTTGGGGCCTCTAAACGGGTCT) into the gRNA expression vector backbone QKS-430 (sequence not shown; those skilled in the art can construct it according to standard molecular biology methods) using Gibson assembly technology, thus constructing a plasmid library stably expressing gRNA-DR. Then, the top 10 electrocompetent cells of the co-expressing gRNA-DR library and ccdB expression plasmid (sequence not shown, which can be constructed by those skilled in the art according to standard molecular biology methods) were prepared. After electroporating the top 10 electrocompetent cells with the Cas12-ART9 protein expression plasmid QKE-260-B0031 (sequence not shown, which can be constructed by those skilled in the art according to standard molecular biology methods), they were recovered in 5 ml LB medium at 37°C and 220 rpm for 1 hour. Then, Cas12-ART9 protein expression was induced by adding 100 ng / ml ATC for 1 hour. E. coli were plated on LB plates containing 1% arabinose and 1×amp (ccdb is a toxic protein; arabinose-induced expression of this plasmid will lead to bacterial death. It can be detoxified by electroporating the Cas12-ART9 expression plasmid and interacting with gRNA to cleave the ccdB expression plasmid. Theoretically, gRNA with targeting activity will be retained. The stronger the DR variant activity, the stronger the ability to detoxify ccdB, and the higher the frequency of retained DR variants) (Seq ID). No: 273-283), the bacteria were collected the next day and library construction and NGS sequencing were performed. The preferred DR sequence was selected for sequence chemical synthesis and verified at the cell level. The experimental method is described in Example 6 and the experimental results are shown in Table 11.
[0259] The results in Table 11 show that gRNAs containing DR variants (Seq ID No: 273-283) modified through directed evolution are more effective at targeting mammalian cells than wild-type DR (Seq ID No: 187).
[0260] Example 10: Directed Evolution of Cas12 Protein (Cas12-ART9)
[0261] To enhance the targeting activity of Cas12-ART9, we modified the Cas12-ART9 protein using directed evolution. First, we constructed single-amino acid mutant libraries of Cas12-ART9, selecting nine regions (i.e., 169-234aa, 235-300aa, 301-366aa, 367-432aa, 433-498aa, 499-564aa, 565-630aa, 837-902aa, and 997-1062aa). For each region, we performed saturation mutation scanning and single-amino acid insertion or deletion. After chemical synthesis of oligonucleotide fragments, they were recombined into the Cas12-ART9 expression vector QKE-260-B0031 (sequence not shown; those skilled in the art can construct it according to standard molecular biology methods) via Gibson assembly. Finally, the nine library plasmids were merged for downstream E. coli screening experiments.
[0262] We designed a forward screening system for E. coli based on ccdB virulence plasmid cleavage. This system includes a Cas12-ART9 mutant plasmid library controlled by the Tet-on promoter, three co-expressed gRNAs (different gRNAs representing different selection pressures), and an arabinose-inducible ccdB protein expression plasmid (sequence not shown, which can be constructed by those skilled in the art according to standard molecular biology methods). This system utilizes Cas12-ART9-mediated plasmid cleavage of the bacterial virulence-encoding gene ccdB to relieve the lethal effect of ccdB expression on bacteria. Theoretically, the stronger the activity of the Cas12-ART9 variant, the stronger the cleavage activity of the ccdB plasmid, and the greater the possibility of bacterial survival. Finally, the variants are genotyped by NGS sequencing, and the preferred Cas12-ART9 variants are selected based on enrichment frequency. The experimental procedure is as follows: (1) First, prepare the top 10 electroporation competent cells expressing 3 gRNAs respectively; (2) Electroporate the above Cas12-ART9 variant plasmid library into the top 10 electroporation competent cells of 3 gRNAs respectively; (3) Recover in 5 ml LB medium at 37°C and 220 rpm for 1 hour; (4) Add 100 ng / ml ATC to induce Cas12-ART9 variant protein expression for 1 hour; (5) Spread E. coli on LB plates containing 1% arabinose and 1×amp; (6) Collect the bacteria the next day and perform library construction and NGS sequencing (Table 15).
[0263] Similarly, we cloned the above 9 fragments into a vector expressing TadA 8e-Cas12-ART9 (E844A mutation, inactivating the double-strand cleavage activity of Cas12-ART9) (sequence not shown, which can be constructed by those skilled in the art according to standard molecular biology methods). The 9 library plasmids were merged for downstream E. coli screening experiments. We designed an ABE-based screening system for antibiotic resistance recovery. This system includes a vector expressing TadA8e-Cas12-ART9 (E844A mutation, inactivating the double-strand cleavage activity of Cas12-ART9) controlled by an arabinose promoter (sequence not shown, but can be constructed by those skilled in the art according to standard molecular biology methods), a gRNA expression vector g4 or g4V2 (Seq ID No: 302-303), and reporter system vectors cam reporter1 and cam reporter2 (sequences not shown, but can be constructed by those skilled in the art according to standard molecular biology methods, with TAA or TAG stop codons introduced into the cam gene; only when the ABE system activates (base A>G mutation), mutating TAA or TAG to TGG, can the bacteria acquire cam resistance). Theoretically, the stronger the activity of the TadA8e-Cas12-ART9 variant, the greater the probability of restoring the TAA or TAG stop codon on the cam gene to the TGG codon, and thus the greater the likelihood of the bacteria retaining in a cam-resistant environment. The experimental procedure is as follows: (1) First, two top10 electroporation competent states expressing g4+cam reporter1 and g4V2+cam reporter2 were prepared respectively; (2) The above Cas12-ART9 variant plasmid library was electroporated into the two top10 electroporation competent states respectively; (3) In 10 ml (4) In LB medium, recover at 37°C and 220 rpm for 1 hour. Add 1% arabinose + 50 ug / ml Amp + 150 ug / ml spe + 150 ug / ml cam + 100 ug / ml kan, shake overnight at 37°C and 220 rpm. (5) Make two copies of each sample. Spread one copy on an LB plate containing 50 ug / ml Amp + 150 ug / ml spe + 150 ug / ml cam + 100 ug / ml kan, and spread the other copy on an LB plate containing 50 ug / ml Amp + 150 ug / ml spe + 100 ug / ml kan as a cam-free resistance screening control. (6) On the third day, collect the bacteria and perform library construction and NGS sequencing (Table 15).
[0264] Table 15 shows the experimental results of variants with significantly increased enrichment folds (fold change > 3) relative to Cas12-ART9 selected by two directed evolution screening systems. High enrichment folds mean that these Cas12 variants have superior editing activity in E. coli compared to Cas12-ART9.
[0265] Example 11: Evaluation of the firing efficiency of the Cas12-ART9 variant based on directed evolution and rational design
[0266] We rationally designed a targeting strategy based on the conservation of the amino acid homology sequence of the Cas12-ART9 protein, enhancing the interaction between amino acids and nucleic acids. Combined with the mutation sites selected through directed evolution in Example 10, we validated the cell-level targeting efficiency. We selected three genomic in situ targeting sites: QKS-g325 (SEQ ID NO:120), QKS-g437 (SEQ ID NO:316), and gRNA2 (SEQ ID NO:211). Using Fugene HD transfection reagent, 88 ng of the Cas12-ART9 variant plasmid and three gRNA plasmids (24 ng / each) were co-transfected into HEK293T cells (96-well plates, 2×10⁶ cells / well). 4 (cells). After 72 hours of culture, the genomic DNA of the cells was lysed, and then the intracellular targeting efficiency of the Cas12-ART9 variant was verified by PCR library construction and NGS sequencing (Table 16).
[0267] The experimental results in Table 16 show that the amino acid mutation sites mutation1-mutation71, selected through rational design or directed evolution, have better editing efficiency than wild-type Cas12-ART9 at different gRNA sites in the cell. Mutation72-mutation160 has at least one or more gRNAs on different gRNAs, and its editing activity is better than Cas12-ART9.
[0268] To further improve the targeting efficiency of Cas12-ART9 (SEQ ID NO:141) at the cellular level, we selected high-performing single-point mutations for direct combination (combinations 1-22), and used a computationally designed method to select high-performing single-site mutations for amino acid regeneration and combination (combinations 23-46). Fugene HD transfection reagent was used to co-transfect 100 ng of the Cas12-ART9 variant plasmid and two gRNA plasmids (30 ng / each) into HEK293T cells (96-well plates, 2×10⁶ cells per well). 4(cells). After 72 hours of culture, the genomic DNA of the cells was lysed, and then the intracellular targeting efficiency of the Cas12-ART9 variant was verified by PCR library construction and NGS sequencing (Table 17).
[0269] The experimental data in Table 17 show that the above amino acid mutation combinations (combination 1-22 and combination 23-46) have at least one or more gRNAs on different gRNAs with superior editing activity compared to Cas12-ART9.
[0270] To further confirm the activity of the above Cas12-ART9 combination variants at the mRNA level, we selected the better-performing combination variants (combinations 1-3, 23-28, and 35) for mRNA level testing. The experiment used Lipofectamine MessengerMAX transfection reagent to co-transfect 100 ng of the Cas12-ART9 combination variants with three gRNAs: QKS-g325 (SEQ ID NO: 120), QKS-g437 (SEQ ID NO: 316), and gRNA2 (SEQ ID NO: 211) into Huh-7 cells (96-well plates, 8 × 10⁶ cells per well). The result was approximately 33.3 ng / gRNA. 3 (cells). After 72 hours of culture, the genomic DNA of the cells was lysed, and then the intracellular targeting efficiency of the Cas12-ART9 combined variant was verified by PCR library construction and NGS sequencing (Table 18).
[0271] The experimental data in Table 18 show that the above amino acid mutant combinations (combination 1-3, combination 23-28, combination 35) have superior intracellular editing activity on different gRNAs compared to Cas12-ART9.
[0272] To confirm the activity of the selected Cas12-ART9 combination variants in mouse liver, we packaged Cas12-ART9 combination variants combination23 and combination27 using LNP (formulation ALC-0315), selected gRNA (SEQ ID NO:291), and injected mice via tail vein after LNP packaging (mRNA and gRNA mass ratio 1:1, doses of 0.5 mpk and 0.3 mpk, respectively). Seven days later, liver tissue DNA was extracted, and the targeting activity of Cas12-ART9 combination variants combination23 and combination27 in mouse liver was verified by PCR library construction and NGS sequencing. The experimental results showed that Cas12-ART9 combination variants combination23 and combination27 both enhanced the targeting activity in mouse liver compared to wild type (Table 19).
[0273] Example 12: Replacement of the Cas12-ART9 protein homology domain (PI / Nuc domain / Wed1)
[0274] Through bioinformatics analysis and sequence alignment of Cas12-ART9, we replaced three domains of naturally occurring Cas12-ART9 homologous proteins—comprising 15 PI domains (SEQ ID NO: 328-342), 14 Nuc domains (SEQ ID NO: 343-356), and 12 Wed1 domains (SEQ ID NO: 358-369)—specifically replacing the PI domain (SEQ ID NO: 327), Nuc domain (SEQ ID NO: 80), and Wed1 domain (SEQ ID NO: 357) of Cas12-ART9. The experimental protocol was the same as in Example 11, and the results are shown in Table 20.
[0275] The experimental data in Table 20 show that the three domains of the naturally occurring Cas12-ART9 homologs, PI (e.g., SEQ ID NO:327-342), Nuc (e.g., SEQ ID NO:343-356) and Wed1 (e.g., SEQ ID NO:358-369), can replace the corresponding domains of Cas12-ART9 and maintain the double-strand cutting activity of mammalian cell DNA.
[0276] Example 13: Evaluation of the targeting efficiency of Cas12-ART9 and exonuclease fusion expression
[0277] T5E or Trex2 exonuclease was fused to the N-terminus or C-terminus of the Cas12-ART9 protein (SEQ ID NO:317-319) to construct three mRNA IVT plasmids: Cas12-ART9-T5E (SEQ ID NO:320) , Cas12-ART9-Trex2 (SEQ ID NO:321) , and Trex2-Cas12-ART9 (SEQ ID NO:322) . mRNA was produced in vitro via IVT, and intracellular targeting activity was assessed using the mRNA. Lipofectamine MessengerMAX transfection reagent was used to co-transfect Huh-7 cells (96-well plates, 8×10⁶ cells per cell line) with the mRNA and two gRNAs, QKS-g323 (SEQ ID NO:118) and QKS-g324 (SEQ ID NO:119). 3 Three dose gradients (150ng+50ng+50ng, 50ng+16.7ng+16.7ng, 16.7ng+5.6ng+5.6ng) were co-transfected with mRNA and two gRNAs. After 72 hours of culture, the genomic DNA of the cells was lysed, and the targeting efficiency of the fusion protein in cells was verified by PCR library construction and NGS sequencing (Table 21).
[0278] The experimental data in Table 21 show that fusing exonucleases such as T5E or Trex2 to the N-terminus or C-terminus of the Cas12-ART9 protein can significantly improve the targeting activity of Cas12-ART9 in mammalian cells.
[0279] Example 14: Cas12-ART9 mRNA codon optimization
[0280] We optimized the sequence of Cas12-ART9 mRNA (SEQ ID NO: 323-324) using a computational codon optimization method and validated the mRNA at the cellular level. The experimental protocol is the same as in Example 11. The data show that the codon-optimized mRNA sequence we designed can effectively improve intracellular targeting efficiency. The experimental results are shown in Table 22.
[0281] Example 15: Target Efficiency Evaluation of Cas12-ART4 and Cas12-ART5 Variants
[0282] Preferred single-amino acid mutation sites of Cas12-ART9 were selected and mapped to homologous positions of Cas12-ART5 (mutation 164-mutation 186), and the results were verified using plasmids on HEK293T. The experimental protocol was the same as in Example 11, and the results are shown in Table 23. The experimental results in Table 23 show that the preferred single-amino acid mutation sites of Cas12-ART9 (mutation 164-mutation 181) mapped to Cas12-ART5 (SEQ ID NO:137) can significantly improve the intracellular targeting activity of Cas12-ART5.
[0283] We selected high-performing single-point mutations for combination on the Cas12-ART5 protein (combinations 47-68), and based on computational design, we selected preferred single-point mutations for amino acid regeneration and combination on the Cas12-ART4 protein (combinations 69-81). The experimental protocol was the same as in Example 11, and the results are shown in Table 24. The results in Table 24 indicate that direct combination of the above-mentioned preferred Cas12-ART5 single amino acids (combinations 47-68) significantly improved the intracellular targeting activity of Cas12-ART5, and the selection of preferred single-point mutations for amino acid regeneration and combination (combinations 69-81) improved the intracellular targeting activity of Cas12-ART4 (SEQ ID NO:136).
[0284] To further confirm the activity of the above-mentioned combined variants at the mRNA level, we selected preferred combined variants (combination 47-48, combination 58-59, combination 69-72, combination 74, combination 76) and tested them using mRNA in the HEK293T-gRNA library. The gRNA library construction process is as follows: (1) Design approximately 5000 gRNA sequences. The oligo form adopts a design method called Indelphi, with the left homologous arm (U6 promoter region) - gRNA DR sequence - gRNA spacer sequence - flanking region - PAM+gRNA spacer - flanking region - right homologous arm. (2) The chemically synthesized oligos are recombined into the lentiviral expression vector using the Gibson assembly method. (3) The lentivirus is used to infect cells at 5000 eIs. 6HEK293T cells (MOI=1) were cultured for two weeks in complete medium containing 2 μg / ml puromycin to construct the HEK293T-gRNA library. This design works by integrating an oligo sequence into the genome, expressing a DR-containing gRNA (involved in DNA cleavage), forming a targeting complex with Cas proteins, and cleaving a PAM-containing gRNA sequence on the same strand (the target). Indels are formed through intracellular DNA repair, which can be detected using NGS. The experimental protocol was as follows: 24 hours before transfection, the HEK293T-gRNA library was seeded in 24-well plates with 1e5 cells per well. Transfection was performed using Lipofectamine MessengerMAX transfection reagent, transfecting 1 μg of mRNA per well. After 72 hours of culture, the genomic DNA was lysed, followed by PCR library construction and NGS sequencing. Data were quantified using the mean and median of all gRNAs. The experimental results are shown in Table 25.
[0285] The experimental results in Table 25 show that the Cas12-ART5 and Cas12-ART4 amino acid combination mutant variants (combination47-48, combination58-59, combination69-72, combination74, combination76) selected through the above procedure were all validated in the HEK293T-gRNA library cell line, and their targeted editing efficiency was significantly improved compared with the original version.
[0286] Example 16: Evaluation of the targeting efficiency of Cas12-ART5 variants generated by ePCR on gRNA libraries
[0287] We randomly mutated the Cas12-ART5 protein (SEQ ID NO: 137) using error-prone PCR. Library construction was performed according to the kit parameters 4.5–9 mutations / kb (GeneMorph II Random Mutagenesis Kit, angilent technologies). The selection pressure was QKS-354-M1516 (Seq ID No: 300). The experimental procedure was as follows: (1) First, top10 electroporation competent cells expressing QKS-354-M1516 were prepared; (2) The above Cas12-ART5 variant plasmid library was electroporated into the Top10 electroporation competent cells; (3) In 5 ml (3) Recover at 37°C and 220 rpm for 1 hour in LB medium. (4) Add 100 ng / ml ATC to induce Cas12-ART5 variant protein expression for 1 hour. (5) Spread E. coli on LB plates containing 1% arabinose and 1×amp. (6) Pick single clones for Sanger sequencing, and then construct them into mRNA-IVT plasmids according to genotype. After in vitro transcription of the mRNA mutant, it was tested in the HEK293T-gRNA library (Table 26). The experimental procedure is as described in Example 15.
[0288] The experimental data in Table 26 show that the editing activity of the Cas12-ART5 variants (combination 82-85; combination 88-95) obtained by error-prone PCR in the HEK293T-gRNA library is superior to that of Cas12-ART5.
[0289] Table 1-1. Sequences of plasmids and wild-type Cas12
[0290] Table 1-2. gRNA (crRNA) and PAM sequences
[0291] Table 1-3. Sequence of each domain of Cas12
[0292] Table 2-1. Sequences of plasmids and Cas12 variants
[0293] Table 2-2 gRNA (crRNA) and PAM sequences
[0294] Table 3-1. Sequences of Cas12 variants
[0295] Table 3-2. Sequences of gRNA (crRNA) and PAM
[0296] Note:
[0297] gRNA (crRNA) is formed by directly linking a DR sequence (at the 5' end) and a spacer sequence (at the 3' end). For example, the DR sequence shown in SEQ ID NO:206 is directly linked with the spacer sequence shown in SEQ ID NO:210 to form gRNA1 (crRNA) shown in SEQ ID NO:145, wherein the DR sequence is at the 5' end of gRNA1 (crRNA) and the spacer sequence is at the 3' end of gRNA1 (crRNA).
[0298] Table 3-3. Sequence of each domain of Cas12
[0299] Table 4-1. Plasmid sequences
[0300] Table 4-2 - gRNA (crRNA) and PAM sequences
[0301] Table 16 Evaluation of single-point mutation targeting efficiency based on directed evolution and rational design (for Cas12-ART9 normalization) Note: _insert_ means inserting an amino acid at the position indicated by the preceding number, and _drop means deleting the amino acid at the position indicated by the preceding number; "NA" means no analysis was performed and the data was normalized relative to the control.
[0302] Table 17 Evaluation of the targeting efficiency of combined mutations (for wild-type Cas12-ART9 normalization) Note: Data were normalized relative to the control group.
[0303] Table 18 Evaluation of the targeting efficiency of combined mutant mRNAs on huh-7
[0304] Table 19 Evaluation of the targeting efficiency of combined mutant mRNAs in mice
[0305] Table 20 Target Efficiency Evaluation After Replacement of Cas12i Homologous Domains (PI and Nuc Domains)
[0306] Table 21 Evaluation of Targeting Efficiency of Cas12 and Exonuclease Fusion Expression
[0307] Table 22 Evaluation of Cas12 codon optimization targeting efficiency Note: Data were normalized relative to the control group.
[0308] Table 23 Evaluation of Cas12-ART5 Single-Point Mutation Targeting Efficiency Note: Data were normalized relative to the control group.
[0309] Table 25 Evaluation of the targeting efficiency of Cas12-ART5 and Cas12-ART4 mutant combination mRNAs on gRNA libraries
[0310] Table 26 Evaluation of the targeting efficiency of Cas12-ART5 mutant combinatorial mRNA generated by ePCR on gRNA libraries - 1
[0311] Table 27 Mouse genome targeting gRNA sequence information
[0312] Table 28 gRNA sequence information for directed evolution systems
[0313] Table 29 Human Genome Targeted gRNA Sequence Information
[0314] Table 30 Exonuclease amino acid and vector sequence information
[0315] Table 31 mRNA sequence and vector sequence information for in vitro transcription of mRNA
[0316] Table 32. Amino acid sequence information of the PI domain of natural homologous proteins of Cas12.
[0317] Table 33. Amino acid sequence information of the Nuc domain of natural homologs of Cas12.
[0318] Table 34. Amino acid sequence information of the Wed-1 domain of Cas12 natural homologs.
[0319] The preferred embodiments of this disclosure have been described in detail above. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, those skilled in the art can make various changes to the technical solutions of this disclosure and still obtain the desired technical effects. All such changes fall within the protection scope of this disclosure.
Claims
1. An engineered Cas12 protein, said engineered Cas12 protein comprising, relative to wild-type Cas12 protein: (i) A Helical-I domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:17-26, or (ii) A Helical-II domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:27-36, or (iii) A Helical-III domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:37-46, or (iv) A NUC domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO:47-56, 79-81, and 343-356, or (v) A PI domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO:57-66, 77-78, and 327-342, or (vi) A WED-II domain comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of the amino acids in SEQ ID NO:67-76, or (vii) Any combination of two or more of the terms in (i)-(vi) above, The wild-type Cas12 protein mentioned therein is any one of the wild-type Cas12 proteins selected from wild-type Cas12-1, Cas12-2, Cas12-3 and Cas12-4.
2. An engineered Cas12 protein, which contains, relative to SEQ ID, Any one or more changes to the following positions of number NO:141: E462X, D926X, T505X, E494X, D293X, T505X, D709X, E350X, A691X, E736X, E179X, D165X, D704X, E399X, Y292X, K931X, E262X, K703X, E1035X, V842X, E675X, E875X, E601X, E504X, C567X, K277X, D356X, D904X, K563X, E285X, E388X, G478X, E271X, H520X, K809X, L967X, E 815X, V842X, V602X, T354X, K138X, I657X, Y468X, N879X, Q186X, S318X, D 356X, I265X, D591X, K404X, V842X, K270X, H858X, E321X, W1046X, K187X, G 232X, F891X, L97X, L13X, L74X, M618X, I1048X, S582X, D120X, V1047X, S24 3X, H858X, N168X, E681X, Q620X, K401X, T943X, W927X, E253X, A854X, G402 X, K774X, W1046X, E793X, E252X, R302X, E447X, V1039X, R330X, D906X, S2 40X, A624X, T235X, R902X, V1014X, R902X, D851X, F583X, Y537X, 443_inse rt_X, 1046_insert_X, 221_insert_X, D907X, D684X, V1047X, L274X, E157X, W1046X, E479X, R302X, N456X, 1046_drop, I440X, K181X, P275X, 1048_ drop, S547X, R902X, R233X, A130R, R902X, R433X, V189X, R902X, V1047X, A239X, H901X, V532X, 1045_insert_X, R267X, K1052X, E434X, S548X, S470 X, 265_insert_X, R433X, S664R, E244X, R369X, Q209X, R302X, R263X, S625X, R902X, E422X, R267X, L192X, 1047_drop, E328X, V1022X, I199X, D864X,F891X, K281X, T238X, W648X, 609_insert_X, E518X, A239X, V173X, E658X, I414X, K886X, or R433X, where X is any amino acid different from the original amino acid at the indicated position, _insert_ indicates the insertion of any amino acid at the position indicated by the preceding number, and _drop_ indicates the deletion of the amino acid at the position indicated by the preceding number. Preferably, it includes any one or more changes selected from the following positions: E462X, D926X. T505X, E494X, D293X, T505X, D709X, E350X, A691X, E736X, E179X, D165X, D704X, E399X, Y292X, K931X, E262X, K703X, E1035X, V842X, D356X, D904X, E388X, K809X, V842X, N879X, S318X, D356X, V842X, H858X, I1048X, H858X, or A854X; or... It includes any one or more changes at the following positions relative to SEQ ID NO:137: D711X, K705X, E738X, S435X, V844X, N233X, D706X, A856X, N371X, N269X, D928X, N168X, E464X, E352X, A693X, E390X, T507X, K933X, G480X, Q951X, or K606X, where X is any amino acid different from the original amino acid at the indicated position; or It includes any one or more changes at the following positions relative to the SEQ ID NO:136 number: A693X, N233X, S435X, N269X, E352X, N371, E738X, A856X, D711X, or D706X, where X is any amino acid that is different from the original amino acid at the indicated position.
3. The engineered Cas12 protein according to claim 1 or 2, wherein the engineered Cas12 protein comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO: 83-116, 121-144, or 370-404, or is composed of any one of said amino acid sequences. Preferably, the engineered Cas12 protein comprises the amino acid sequence of any one of SEQ ID NO: 141 or 370-394, 133, 134, 135, 136 or 395-399, 137 or 400-404, or is composed of the amino acid sequence of any one of SEQ ID NO: 141 or 370-394, 133, 134, 135, 136 or 395-399, 137 or 400-404, or Optionally, relative to the number in SEQ ID NO:141, the engineered Cas12 protein comprises a mutation selected from one or more of the following: E462R, D926R, T505K, E494R, D293R, T505R, D709R, E350R, A691R, E736R, E179R, D165R, D704R, E399R, Y292R, K931R, E262R, K703R, E1035R, V842M, E675R, E875R, E601R, E504R, C567R, K277R, D356N, D904R, K563R, E285R, E388R, G478R, E271R, H520 Y, K809R, L967K, E815R, V842A, V602R, T354S, K138R, I657L, Y468R, N879T , Q186R, S318F, D356H, I265R, D591R, K404R, V842L, K270R, H858M, E321R, W 1046R, K187R, G232D, F891Q, L97F, L13K, L74W, M618L, I1048V, S582R, D12 0R, V1047K, S243R, H858R, N168F, E681R, Q620L, K401A, T943R, W927R, E253 R, A854Y, G402W, K774R, W1046S, E793R, E252R, R302G, E447R, V1039H, R33 0P, D906R, S240R, A624R, T235R, R902G, V1014I, R902E, D851R, F583L, Y53 7H, 443_insert_A, 1046_insert_D, 221_insert_V, D907R, D684F, V1047D, L274R, E157R, W1046E, E479R, R302A, N456G, 1046_drop, I440M, K181R, P2 75K, 1048_drop, S547D, R902H, R233W, A130R, R902D, R433V, V189R, R902I , V1047E, A239R, H901M, V532F, 1045_insert_N, R267L, K1052N, E434R, S54 8V, S470R, 265_insert_G, R433M, S664R, E244K, R369Y, Q209H, R302V, R26 3A, S625P, R902W, E422R, R267G, L192R, 1047_drop, E328R, V1022A, I199L,D864R, F891E, K281R, T238R, W648D, 609_insert_R, E518R, A239C, V173C, E658R, I414L, K886N, or R433I, where _insert_ indicates the insertion of an amino acid at the position indicated by the preceding number, and _drop_ indicates the deletion of an amino acid at the position indicated by the preceding number; preferably, it contains mutations selected from any one or more of the following: E462R, D926R, T505 K, E494R, D293R, T505R, D709R, E350R, A691R, E736R, E179R, D165R, D704R, E399R, Y292R, K931R, E262R, K703R, E1035R, V842M, D356N, D904R, E388R, K809R, V842A, N879T, S318F, D356H, V842L, H858M, I1048V, H858R, or A854Y; or Optionally, relative to the number in SEQ ID NO:141, the engineered Cas12 protein comprises a combination of mutations selected from one or more of the following: E736R, E388R, E462R; E462R, A854Y; A691R, E388R, T505K; E462R, A691R; E462R, V842M; D926R, A691R, E736R; E462R, A691R, E736R, D926R; D926R, A691R; D704R, I1048V; E388R, A854Y; D704R, D356H; D704R, T505K; E388R, D926R; A854Y, T505K, A691R, E736R; E388R, A691R; E388R, I1048V; D704R, E388R; D926R, E388R, E462R, A691R; D704R, A691R; D704R, A854Y; D704R, D926R; E388R, T505K; E388K, E462R, E736K; E462R, E736D, D926E; E388D, E462R, A854Y, D926E; E388S, E462R, T505W; E462R, E736D, A854Q, D926S; E462R, T505M; D356C, D926W; E462C, D926C; E350P, D926C; D704C, D926C; D356C, D926C; T505Y, D704C; E462R, T505M, E736D; T505C, D926C; T505F, D926C; A854C, D926C; E388D, E462H, T505I, E736D; E262L, D926C; D356C, A854C; E350C, D926C; T505W, A854L; D356C, E736C; T505F, A854L; or D356C, A854L; Preferably, it comprises a combination of mutations selected from any of the following: E736R, E388R, E462R; E462R, A854Y; A691R, E388R, T505K; E388K, E462R, E736K; E462R, E736D, D926E; E388D, E462R, A854Y, D926E; E388S, E462R, T505W; E462R, E736D, A854Q, D926S; E462R, T505M; or E462R, T505M, E736D; or Optionally, relative to the number in SEQ ID NO:137, the engineered Cas12 protein comprises a mutation selected from one or more of the following: D711R, K705R, E738R, S435R, V844M, N233R, D706R, A856Y, N371R, N269R, D928R, N168R, E464R, E352R, A693R, E390R, T507R, K933R, G480R, Q951H, or K606E; or Optionally, relative to SEQ ID NO:137, the engineered Cas12 protein comprises: D711R and combinations thereof selected from one or more of the following: S435R, K705R, G480R, N233R, A693R, N371R, V844M, A856Y, E738R, N269R, or E352R; or E738R and combinations thereof selected from one or more of the following: K705R, A856Y, A693R, S435R, E352R, V844M, N371R, G480R, N233R, N269R, or D706R; or Choose from any of the following mutation combinations: M759V, K776N, H903N, D909E; R662H, P716L, K776T, E961D, E966G, L989I, I1023T, M1056L; D711G, A796S, F802S, D853E, E967D, R973L, D975V, I980F; Y670H, H860Q; P638L, V702A, T767A, R904C; R812G, Q832H, V984D, G1019C; V462F, K606N; R662H, P716L, K776T, E961D, E966G, L989I, I1023T, M1056L; D711G, A796S, F802S, D853E, E967D, R973L, D975V, I980F; F646Y, E790D, R904S, K941I, V1049M; S661C, G707S, A752V, R812G; or H676Q, H733Q, G787A, E968V, N979I; or Optionally, relative to the number in SEQ ID NO:136, the engineered Cas12 protein comprises a mutation selected from one or more of the following: A693R, N233R or N233K or N233S, S435R or S435E or S435K or S435T, N269R or N269E or N269T or N269K, E352R, N371R or N371K, E738D, A856Q or A856S, D711R, or D706R; or Optionally, relative to the number SEQ ID NO:136, the engineered Cas12 protein comprises a combination of mutations selected from any of the following: A693R, N233R, S435R; A693R, N269R, E352R, N371R; N233K, N269T, N371K, S435E, E738D, A856Q; N233K, N371R, S435E, E738D; N233K, N269E, S435K, E738D; N233S, N269R, S435K, A856Q; A693R, N233R, E352R; N233K, N269E, N371K, S435E, E738D; S435R, A693R, D711R; N233K, N371K, S435E, E738D, A856S; N233R, D706R; N233K, N269R, S435E, E738D, A856Q; or N233K, N269K, N371K, S435T; or Optionally, the engineered Cas12 protein comprises a PI domain as shown in any one of SEQ ID NO:327-342; or Optionally, the engineered Cas12 protein comprises a Nuc domain as shown in any one of SEQ ID NO:343-356; or Optionally, the engineered Cas12 protein comprises the Wed1 domain shown in any one of SEQ ID NO:358-369.
4. The engineered Cas12 protein according to claim 1, wherein the engineered Cas12 protein comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID NO:157-180.
5. The engineered Cas12 protein according to any one of claims 1-4, wherein the engineered Cas12 protein is fused with a polypeptide, and optionally, the engineered Cas12 protein is fused with an exonuclease.
6. A nucleic acid molecule comprising a nucleotide sequence encoding an engineered Cas12 protein according to any one of claims 1-5, wherein the nucleotide sequence is optionally codon-optimized.
7. An engineered CRISPR-Cas system, the system comprising: (a) a guide RNA (gRNA) or a nucleic acid sequence encoding said gRNA, wherein said gRNA comprises a direct repeat (DR) sequence and a spacer region sequence, wherein said DR sequence is capable of binding to the engineered Cas12 protein of any one of claims 1-5; and (b) the engineered Cas12 protein or the nucleic acid sequence encoding it as described in any one of claims 1-5. The engineered Cas12 protein associates with the gRNA. The spacer sequence is a nucleotide sequence that is at least partially complementary to the target nucleic acid. Optionally, the gRNA further comprises a 5' end extension sequence, which optionally contains one or more modifications selected from 2'O-methyl or thiophosphorylation modifications or combinations thereof.
8. The engineered CRISPR-Cas system of claim 7, wherein the gRNA comprises one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro group, thiophosphorylation, or combinations thereof.
9. The engineered CRISPR-Cas system of claim 7, wherein the DR sequence comprises one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof. Optionally, the 15th, 16th, 3rd, 4th, 1st or 2nd base of the DR sequence is modified with 2'O-methyl, and optionally the 11th or 12th base of the DR sequence is not modified with 2'O-methyl; Optionally, the DR sequence further includes a U7C mutation at position 7; Optionally, any one of the bases at positions 10, 19, 8, 11, 9, 21, 6, or 20 of the DR sequence is fluorinated; or optionally, the DR sequence comprises: a combination of fluorinated bases at positions 4, 5, 12, and 15; a combination of fluorinated bases at positions 4, 12, 14, and 15; a combination of fluorinated bases at positions 4, 5, 12, and 14; or a combination of fluorinated bases at positions 4, 5, 14, and 15. Optionally, the DR sequence is modified through directed evolution; or Preferably, the DR sequence comprises any one of the sequences shown in SEQ ID NO: 184-186, 187, 189-194, 202-203, 207-230, 223-230, 232-235, 273-283, 406-408, or 412-414; and / or The spacer region comprises one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof. Optionally, the base at any of the 12th, 3rd, 8th, 14th, and 11th positions of the spacer region sequence is fluorinated. Optionally, the base at the 11th and 12th, 7th and 8th, 5th and 6th, 3rd and 4th, 9th and 10th, 1st and 2nd, 15th and 16th, or 13th and 14th positions of the spacer region sequence is fluorinated. Optionally, the 3' end of the spacer sequence comprises three 2'O-methyl modified bases and three thiophosphoryl modifications; or Optionally, the spacer sequence comprises any one of the sequences shown in SEQ ID NO: 197, 205-206, 231, 240-246, 254-257, 263-270, 409-411 or 415-417; Optionally, the gRNA further comprises a 3' end sequence, which optionally comprises the sequence shown in SEQ ID NO:
198.
10. The engineered CRISPR-Cas system according to any one of claims 7-9, wherein the engineered CRISPR-Cas system further comprises a nuclear localization signal (NLS) sequence operatively linked to the engineered Cas12 protein, wherein the NLS is an N-terminal NLS and / or a C-terminal NLS.
11. An engineered guide RNA (gRNA) comprising: (a) A DNA targeting region containing a nucleotide sequence complementary to the target sequence in the target DNA molecule; and (b) A protein-binding segment configured to associate with the engineered Cas12 protein according to any one of claims 1-5; The engineered guide RNA (gRNA) contains one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro group, thiophosphorylation, or combinations thereof.
12. The engineered gRNA of claim 11, wherein the engineered gRNA comprises, from 5' to 3': (i) A 5' end extension sequence comprising one or more modifications selected from 2'O-methyl or thiophosphorylation modifications or combinations thereof, optionally, the 5' end sequence comprising any one of the nucleotide sequences described in SEQ ID NO:199-201; (ii) a direct repeat (DR) sequence comprising one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof. Optionally, the 15th, 16th, 3rd, 4th, 1st or 2nd base of the DR sequence is modified with 2'O-methyl, and optionally the 11th or 12th base of the DR sequence is not modified with 2'O-methyl; Optionally, the DR sequence further includes a U7C mutation at position 7; Optionally, any one of the bases at positions 10, 19, 8, 11, 9, 21, 6, or 20 of the DR sequence is fluorinated; or optionally, the DR sequence comprises: a combination of fluorinated bases at positions 4, 5, 12, and 15; a combination of fluorinated bases at positions 4, 12, 14, and 15; a combination of fluorinated bases at positions 4, 5, 12, and 14; or a combination of fluorinated bases at positions 4, 5, 14, and 15. Optionally, the DR sequence is modified through directed evolution; or Preferably, the DR sequence comprises any one of the sequences shown in SEQ ID NO: 184-186, 187, 189-194, 202-203, 207-230, 223-230, 232-235, 273-283, 406-408, or 412-414; and (iii) A spacer sequence comprising one or more modifications selected from 2'O-methyl (i.e., 2' methylation), 2' fluoro, thiophosphorylation, or combinations thereof, optionally, a base at any of the 12th, 3rd, 8th, 14th, or 11th positions of the spacer sequence is fluorinated, optionally, a base at the 11th and 12th, 7th and 8th, 5th and 6th, 3rd and 4th, 9th and 10th, 1st and 2nd, 15th and 16th, or 13th and 14th positions of the spacer sequence is fluorinated in combination; Optionally, the 3' end of the spacer sequence comprises three 2'O-methyl modified bases and three thiophosphoryl modifications; or Optionally, the spacer sequence comprises any one of the sequences shown in SEQ ID NO: 197, 205-206, 231, 240-246, 254-257, 263-270, 409-411 or 415-417; Optionally, the engineered gRNA further comprises a 3' end sequence, which optionally comprises the sequence shown in SEQ ID NO:
198.
13. A recombinant expression system expressing the engineered CRISPR-Cas system according to any one of claims 7-10, the recombinant expression system comprising: (i) a nucleic acid sequence encoding the engineered Cas12 protein as described in any one of claims 1-5; and (ii) Nucleic acid sequences that encode DNA target sequences.
14. The recombinant expression system of claim 13, wherein the DNA targeting sequence comprises a guide RNA (gRNA), optionally, the gRNA comprising any one of the sequences shown in SEQ ID NO: 7-10, 118-120, 182, 184-186, 187, 189-194, 202-203, 207-230, 223-230, 232-235, 273-283, 406-408 or 412-414, or comprising any one of the sequences shown in SEQ ID NO: SEQ ID NO: 197, 205-206, 231, 240-246, 254-257, 263-270, 409-411 or 415-417.
15. The recombinant expression system according to claim 13, wherein (i) and (ii) are contained in the same vector or in different vectors.
16. The recombinant expression system of claim 15, wherein the vector is a viral vector, for example, selected from retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, and herpes simplex vectors.
17. A recombinant host cell comprising: an engineered Cas12 protein according to any one of claims 1-5, or a nucleic acid molecule according to claim 6, or an engineered CRISPR-Cas system according to any one of claims 7-10, or an engineered gRNA according to claim 11 or 12, or a recombinant expression system according to any one of claims 13-16.
18. A pharmaceutical composition comprising: an engineered CRISPR-Cas system according to any one of claims 7-10, a recombinant expression system according to any one of claims 13-16, or the cell according to claim 17, and a pharmaceutically acceptable carrier.
19. A kit comprising: (a) the engineered Cas12 protein of any one of claims 1-5, or the nucleic acid molecule of claim 6, or the engineered CRISPR-Cas system of any one of claims 7-10, or the engineered gRNA of claim 11 or 12, or the recombinant expression system of any one of claims 13-16, or the cell of claim 17, and (b) Instructions for use.
20. A method for editing nucleobases in a target nucleic acid sequence, the method comprising: The target nucleic acid sequence is contacted with an engineered CRISPR-Cas system according to any one of claims 7-10 or a recombinant expression system according to any one of claims 13-16, wherein the DNA target sequence binds to the target nucleic acid sequence, thereby editing the nucleobases in the target nucleic acid sequence.
21. The method of claim 20, wherein the editing comprises a mutation of a single nucleobase in a target nucleic acid sequence, wherein the mutation comprises an insertion, deletion, or substitution.
22. The method of claim 20, wherein the editing includes gene sequence insertion, gene sequence deletion, or gene sequence replacement.
23. The method of claim 20, wherein the editing comprises epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification comprises methylation or demethylation of the target nucleic acid sequence.
24. A method of treating a disease or disorder associated with a target nucleic acid in a cell of an individual, the method comprising: Contacting the target nucleic acid sequence in the cells of the individual with the engineered CRISPR-Cas system of any one of claims 7-10 or the recombinant expression system of any one of claims 13-16, wherein the DNA target sequence binds to the target nucleic acid sequence, thereby editing the nucleobases in the target nucleic acid sequence to treat a disease or condition related to the target nucleic acid.
25. The method of claim 24, wherein the individual is a mammal, preferably a human.
26. The method of claim 24, wherein the disease or condition is selected from: cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
27. A modular method of engineering a Casl2 protein, the method comprising: Variant sequences were designed as module sequences for each of the Helical-I, Helical-II, Helical-III, NUC, PI, and WED-II domains of the wild-type Cas12 protein. One or more module sequences were used to replace the corresponding wild-type Cas12 protein domains to assemble the engineered Cas12 protein.
28. A lipid nanoparticle encapsulating a nucleic acid molecule encoding the engineered Cas12 protein of any one of claims 1-5 and the engineered gRNA of any one of claims 11 or 12, or encapsulating the engineered CRISPR-Cas system of any one of claims 7-10.