Cas protein and its mutants, as well as corresponding gene editing systems and their use
Cas protein variants with targeted amino acid mutations address the limitations of existing CRISPR/Cas systems by enhancing cleavage activity and binding efficiency, achieving significant improvements in target sequence recognition and editing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2026-04-14
AI Technical Summary
Existing CRISPR/Cas systems have varying advantages and disadvantages in terms of Cas protein size, guide RNA, and PAM recognition, necessitating the development of novel Cas proteins and mutants to meet diverse usage needs.
Development of Cas protein variants with specific amino acid mutations at key sites, enhancing cleavage activity and binding efficiency, and providing improved spacer-specific endonuclease cleavage activity.
The Cas protein variants exhibit enhanced cleavage activity, improved binding to target sites, and altered editing bias, achieving at least 5% to 10-fold improvements in target sequence cleavage compared to wild-type proteins.
Smart Images

Figure 2026511580000050 
Figure 2026511580000051 
Figure 2026511580000052
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene editing, and specifically to Cas proteins and their mutants, as well as corresponding gene editing systems and uses.
Background Art
[0002] The Clustered regularly interspaced short palindromic repeats (CRISPR) system is formed by bacteria and archaea to defend themselves against the DNA of invading phages. Among them, the most common is the CRISPR / Cas9 system. The Cas9 protein can process pre-crRNA into mature crRNA that binds to tracrRNA with the assistance of trans-coding small RNA (tracrRNA). Subsequently, it was discovered that by artificially constructing a single-stranded chimeric guide RNA (guide RNA, gRNA) that mimics the crRNA-tracrRNA complex, the recognition and cleavage of target sites by the Cas9 protein can be effectively mediated. Among them, the three bases adjacent to the 3' end of the target site must be in the form of 5'-NGG-3', thereby forming the PAM (protospacer adjacent motif) structure necessary for the Cas / crRNA complex to recognize the target site.
[0003] However, the various existing CRISPR / Cas have different advantages and disadvantages. For example, there are differences in the sizes of different Cas proteins, guide RNAs, and PAMs.
[0004] Therefore, even now, there is a need to develop novel Cas proteins and their mutants, as well as CRISPR-Cas systems, to meet diverse usage needs.
Summary of the Invention
[0005] The main objective of the present invention is to provide novel Cas proteins and their mutants, as well as gene editing systems and their uses, to meet the above-mentioned usage needs. Based on this, the present invention also provides novel CRISPR-Cas compositions, as well as gene editing methods and nucleic acid detection methods based on the system.
[0006] A first aspect of the present invention provides a Cas protein selected from the following group.
[0007] (a) Polypeptides having the amino acid sequence shown in Sequence ID No. 1; (b) Polypeptides having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% or more homology (or identity) with the amino acid sequence shown in Sequence ID No. 1, and having the biological function of Sequence ID No. 1; (c) An inducible polypeptide having one or more (preferably 1 to 20, more preferably 2 to 18, more preferably 3 to 16) amino acid residues substituted, deleted, or added in the amino acid sequence shown in Sequence ID No. 1, and retaining the biological function of Sequence ID No. 1.
[0008] In some embodiments, the Cas protein includes wild-type and mutant Cas proteins (i.e., orthologues, homologs, variants, or functional fragments of the Cas protein, where the orthologues, homologs, variants, or functional fragments substantially retain the biological function of the sequence from which they are derived).
[0009] In some embodiments, the Cas protein is a variant of the protein having the amino acid sequence shown in SEQ ID NO: 1.
[0010] In some embodiments, the Cas protein is a non-naturally occurring protein.
[0011] In some embodiments, the Cas protein is a Cas protein variant, the variant is a non-natural protein, and the variant corresponds to E118, G119, D120, V121, Q122, Y123, K241, S242, S243, S246, L272, S275, C277, T285, S286, C287, D288, E289, A2 in SEQ ID NO: 1, which corresponds to the wild-type Cas protein. The molecule contains a mutation in one or more important amino acid sites related to cleavage activity, selected from the group consisting of 90, I292, K299, R300, W301, L394, D395, R396, D397, R398, K399, E400, D401, E402, D403, S404, C405, R406, F407, E408, K502, K557, L657, C672, P758, L789, K827, and V860.
[0012] In some embodiments, the Cas protein variant exhibits enhanced cleavage activity of a target molecule complementary to the guide RNA sequence at the target sequence, or activity equivalent to that of the wild-type Cas protein. Alternatively, the CRISPR composition containing the Cas protein variant exhibits enhanced binding to the binding site or altered editing bias.
[0013] In some embodiments, the Cas protein mutant exhibits significantly improved spacer-specific endonuclease cleavage activity against target sequences of target DNA complementary to the spacer sequence compared to the wild-type Cas protein (for example, an improvement of at least 5%, preferably at least 10%, more preferably at least 20%, for example at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%, for example at least 1-fold improvement, at least 2-fold improvement, at least 3-fold improvement, at least 4-fold improvement, at least 5-fold improvement, at least 6-fold improvement, at least 8-fold improvement, at least 9-fold improvement, or at least 10-fold improvement).
[0014] In some embodiments, the Cas protein is compared to a polypeptide having the amino acid sequence shown in SEQ ID NO: 1. (1) The mutation in E118 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, and more preferably from E118D, E118K, E118I, E118Q, E118P, E118S, E118T, and E118Y.
[0015] (2) The mutation in G119 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from G119P, G119A, G119V, G119L, G119I, G119D, G119M, G119R, and G119S, more preferably from G119D, G119M, G119L, G119P, G119R, G119S, and G119A, more preferably from G119D, G119M, G119L, G119P, G119R, and G119S.
[0016] (3) The mutation in D120 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, and more preferably from D120E, D120K, D120L, D120N, D120G, D120S, and D120T.
[0017] (4) The mutation in V121 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, preferably from V121I, V121L, V121M, V121F, V121A, V121G, V121E, V121P, V121R, V121S, and V121T, preferably from V121L, V121A, V121E, V121G, V121P, V121R, V121S, and V121T, and more preferably from V121A, V121E, V121G, V121P, V121R, V121S, and V121T.
[0018] (5) Mutations in Q122 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, and more preferably the mutations are selected from Q122N, Q122S, Q122T, Q122K, Q122F, Q122K, Q122H, Q122G, Q122R, Q122S, Q122T, and Q122Y. Preferably, the mutation is selected from Q122K, Q122F, Q122K, Q122H, Q122G, Q122R, Q122S, Q122T, Q122Y, and Q122N, and more preferably, the mutation is selected from Q122K, Q122F, Q122K, Q122H, Q122G, Q122R, Q122S, Q122T, and Q122Y.
[0019] (6) The mutation in Y123 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from Y123F, Y123L, Y123I, Y123S, Y123T, and Y123W, more preferably from Y123F, Y123L, Y123I, Y123S, Y123T, and F, and more preferably from Y123F, Y123L, Y123I, Y123S, and Y123T.
[0020] (7) The mutation in K241 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from K241A, K241R, K241Q, K241N, K241H, K241A, and K241R, and more preferably from K241A.
[0021] (8) The mutation in S242 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from S242T, S242N, S242Q, S242A, S242F, and S242Y, more preferably from S242F, S242Y, and S242T, and more preferably from S242F and S242Y.
[0022] (9) The mutation in S243 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably the mutation is selected from S243A, S243C, S243N, S243Q, and S243T, more preferably from S243A, S243C, and S243T, and more preferably from S243A and S243C.
[0023] (10) The mutation in S246 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from S246S, S246T, S246N, S246Q, S246A, and S246R, more preferably from S246R and S246T.
[0024] (11) The mutation in L272 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from L272I, L272V, L272M, L272A, L272F, L272H, and L272G, more preferably from L272H and L272I, and even more preferably from L272H.
[0025] (12) The mutation in S275 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from S275T, S275A, S275R, S275N, and S275Q, more preferably from S275T and S275R, and even more preferably from S275R.
[0026] (13) The mutation in C277 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from C277S, C277M, C277L, and C277P, preferably from C277S and C277L, and even more preferably from C277L.
[0027] (14) The mutation at T285 is selected from any amino acid, preferably from the amino acid types shown in Table A, Table B, and Table C, more preferably, the mutation is selected from T285A, T285S, T285R, T285N, T285Q, more preferably, selected from T285S, T285R, and even more preferably, selected from T285R.
[0028] (15) The mutation at S286 is selected from any amino acid, preferably from the amino acid types shown in Table A, Table B, and Table C, more preferably, the mutation is selected from S286A, S286T, S286F, S286N, S286Q, more preferably, selected from S286F, S286T, and even more preferably, selected from S286F.
[0029] (16) The mutation at C287 is selected from any amino acid, preferably from the amino acid types shown in Table A, Table B, and Table C, more preferably, the mutation is selected from C287S, C287M, C287I, C287P, more preferably, selected from C287S, C287I, and even more preferably, selected from C287I.
[0030] (17) The mutation at D288 is selected from any amino acid, preferably from the amino acid types shown in Table A, Table B, and Table C, more preferably, the mutation is selected from D288W, D288E, and more preferably, selected from D288W.
[0031] (18) The mutation at E289 is selected from any amino acid, preferably from the amino acid types shown in Table A, Table B, and Table C, more preferably, the mutation is selected from E289D, E289A, E289L, E289R, E289S, preferably, selected from E289A, E289L, E289R, E289S.
[0032] (19) The mutation in A290 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, and more preferably from A290L, A290I, A290N, A290H, A290V, A290G, A290S, and A290T, and more preferably from A290N, A290H, and A290V.
[0033] (20) The mutation in I292 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from I292L, I292V, I292M, I292A, I292F, and I292G, more preferably from I292L and I292V.
[0034] (21) The mutation in K299 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from K299R, K299Q, K299N, and K299H, and more preferably from K299R.
[0035] (22) The mutation in R300 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from R300K, R300Q, R300N, R300D, and R300H, more preferably from R300K and R300D, and even more preferably from R300D.
[0036] (23) The mutation in W301 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from W301Y and W301F, more preferably from W301Y and W301F, and even more preferably from W301F.
[0037] (24) The mutation at L394 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from L394I, L394V, L394M, L394A, L394F, L394A, L394H, L394R, L394P, and L394G, more preferably from L394I, L394A, L394H, L394R, and L394P, and more preferably from L394A, L394H, L394R, and L394P.
[0038] (25) The mutation in D395 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from D395E, D395C, D395V, D395P, and D395S, more preferably from D395C, D395V, D395P, and D395S.
[0039] (26) The mutation in R396 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from R396K, R396Q, R396N, R396G, R396T, R396Y, and R396H, more preferably from R396K, R396G, R396T, and R396Y, and even more preferably from R396G, R396T, and R396Y.
[0040] (27) The mutation in D397 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, and more preferably the mutation is selected from D397E, D397F, D397H, D397N, D397R, D397S, and D397T, and more preferably from D397F, D397H, D397N, D397R, D397S, and D397T.
[0041] (28) Mutations in R398 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from R398K, R398Q, R398N, R398A, R398D, R398G, R398N, R398I, R398S, and R398H, more preferably from R398A, R398D, R398G, R398N, R398I, R398S, and R398K, and even more preferably from R398A, R398D, R398G, R398N, R398I, and R398S.
[0042] (29) Mutations in K399 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, and more preferably the mutations are K399R, K399Q, K399N, K399D, K399P, K399L, K399I, K399V, K399E, K399R, K399S, K399T, K399P, K399H Selected from, more preferably selected from K399R, K399D, K399P, K399L, K399I, K399V, K399E, K399R, K399S, K399T, K399P, and even more preferably selected from K399D, K399P, K399L, K399I, K399V, K399E, K399R, K399S, K399T, K399P.
[0043] (30) The mutation in E400 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from E400D, E400A, E400C, E400F, E400P, E400R, E400S, E400T, and E400V, and even more preferably from E400A, E400C, E400F, E400P, E400R, E400S, E400T, and E400V.
[0044] (31) The mutation in D401 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from D401E, D401C, D401F, D401H, D401G, D401M, D401K, D401P, D401R, D401S, D401T, and D401V, and even more preferably from D401C, D401F, D401H, D401G, D401M, D401K, D401P, D401R, D401S, D401T, and D401V.
[0045] (32) Mutations in E402 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from E402D, E402A, E402F, E402G, E402L, E402I, E402R, E402S, E402T, and E402Y, and even more preferably from E402A, E402F, E402G, E402L, E402I, E402R, E402S, E402T, and E402Y.
[0046] (33) The mutation in D403 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from D403E, D403G, D403P, D403R, D403S, and D403T, and more preferably from D403G, D403P, D403R, D403S, and D403T.
[0047] (34) Mutations in S404 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from S404A, S404T, S404K, S404Q, S404R, S404P, S404Y, and S404N, more preferably from S404T, S404K, S404Q, S404R, S404P, and S404Y, and even more preferably from S404K, S404Q, S404R, S404P, and S404Y.
[0048] (35) Mutations in C405 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from C405S, C405M, C405A, C405K, C405L, C405R, and C405P, more preferably from C405S, C405A, C405K, C405L, C405R, and C405P, and even more preferably from C405A, C405K, C405L, C405R, and C405P.
[0049] (36) The mutation in R406 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from R406K, R406Q, R406N, R406R, R406L, R406H, R406G, R406T, and R406P, more preferably from R406K, R406L, R406H, R406G, R406T, and R406P, and even more preferably from R406L, R406H, R406G, R406T, and R406P.
[0050] (37) Mutations in F407 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from F407L, F407V, F407I, F407A, F407Y, F407W, F407M, F407S, F407L, F407P, and F407R, more preferably from F407L, F407M, F407S, F407L, F407P, and F407R, and even more preferably from F407M, F407S, F407L, F407P, and F407R.
[0051] (38) The mutation at E408 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from E408D, E408L, E408I, E408R, and E408V, and even more preferably from E408L, E408I, E408R, and E408V.
[0052] (39) Mutations in K502 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from K502R, K502Q, K502N, and K502H, and even more preferably from K502R.
[0053] (40) The mutation at K557 is selected from any amino acid, more preferably from K502R, K502Q, and K502N, and more preferably from K502N.
[0054] (41) The mutation at L657 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from L657I, L657V, L657M, L657A, L657F, and L657G, more preferably from L657I and L657F, and even more preferably from L657F.
[0055] (42) The mutation in C672 is selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from C672S, C672M, C672R, and C672P, more preferably from C672S and C672R, and even more preferably from C672R.
[0056] (43) Mutations in P758 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from P758L, P758A, P758C, and P758M, and more preferably from P758L.
[0057] (44) Mutations at L789 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from L789I, L789V, L789M, L789A, L789F, L789S, and L789G, more preferably from L789I and L789S, and even more preferably from L789S.
[0058] (45) Mutations in K827 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from K827R, K827Q, K827N, K827E, and K827H, more preferably from K827R and K827E, and even more preferably from K827E.
[0059] (46) Mutations in V860 are selected from any amino acid, preferably from the amino acid types shown in Tables A, B, and C, more preferably from V860I, V860L, V860M, V860F, V860A, and V860G, and more preferably from V860L.
[0060] In some embodiments, the mutations are E118D, E118K, E118I, E118Q, E118P, E118S, E118T, E118Y, G119D, G119M, G119L, G119P, G119R, G119S, D120E, D120K, D120L, D120N, D120G, D120S, D120T, V121A, V121E, V121G, V121P, V121R, V121S, V121T, Q122K, Q122F, Q122K, Q122H, Q122G, Q122R, Q122S, Q122T, Q122Y, Y123F, Y 123L, Y123I, Y123S, Y123T, K241A, S242F, S242Y, S243A, S243C, S246R, L27 2H, S275R, C277L, T285R, S286F, C287I, D288W, E289A, E289L, E289R, E289S, A290N, A290H, A290V, I292L, I292V, K299R, R300D, W301F, L394A, L394H, L3 94R, L394P, D395C, D395V, D395P, D395S, R396G, R396T, R396Y, D397F, D397H , D397N, D397R, D397S, D397T, R398A, R398D, R398G, R398N, R398I, R398S, K 399D, K399P, K399L, K399I, K399V, K399E, K399R, K399S, K399T, K399P, E40 0A, E400C, E400F, E400P, E400R, E400S, E400T, E400V, D401C, D401F, D401H , D401G, D401M, D401K, D401P, D401R, D401S, D401T, D401V, E402A, E402F, E4 02G, E402L, E402I, E402R, E402S, E402T, E402Y, D403G, D403P, D403R, D403 S, D403T, S404K, S404Q, S404R, S404P, S404Y, C405A, C405K, C405L, C405R, C 405P, R406L, R406H, R406G, R406T, R406P, F407M, F407S, F407L, F407P, F40 7R, E408L, E408I, E408R, E408V, K502R, K557N, L657F, C672R, P758L, L789S,Selected from the group consisting of K827E, V860L, or combinations thereof.
[0061] In some embodiments, the mutations include mutations at the L272H, S275R, and C277L sites.
[0062] In some embodiments, the Cas protein further includes mutations in two or more key amino acid sites related to cleavage activity, selected from the group consisting of E289, A290, and I292 in Sequence ID No. 1.
[0063] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of K399, E400, D401, E402, and D403 in SEQ ID NO: 1.
[0064] In some embodiments, the Cas protein further includes mutations in three or more key amino acid sites related to cleavage activity, selected from the group consisting of D403, S404, C405, R406, F407, and E408 in SEQ ID NO: 1.
[0065] In some embodiments, the Cas protein further includes mutations in four or more key amino acid sites related to cleavage activity, selected from the group consisting of L394, D395, R396, D397, and R398 in SEQ ID NO: 1.
[0066] In some embodiments, the Cas protein further includes mutations in five or more key amino acid sites related to cleavage activity, selected from the group consisting of D397, R398, K399, E400, D401, and E402 in SEQ ID NO: 1.
[0067] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of S242, S243, and S246 in SEQ ID NO: 1.
[0068] In some embodiments, the Cas protein further includes mutations in four or more key amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, and Y123 in SEQ ID NO: 1.
[0069] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of D397, R398, E400R, D401, E402, and L789 in SEQ ID NO: 1.
[0070] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of T285, S286, C287, D288, D403, C405, R406, F407, E408, K502, and K557 in SEQ ID NO: 1.
[0071] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of E289, I292, K399, E400, D401, E402, and D403 in Sequence ID No. 1.
[0072] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of E289, A290, I292, P758, and V860 in SEQ ID NO: 1.
[0073] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of K299, R300, W301, C672, and P758 in SEQ ID NO: 1.
[0074] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of D397, R398, E400, D401, E402, L657, and L789 in SEQ ID NO: 1.
[0075] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, C672, and P758 in SEQ ID NO: 1.
[0076] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of K399, E400, D401, E402, D403, P758, and K827 in SEQ ID NO: 1.
[0077] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, L394, D395, R396, D397, R398, C672, and P758 in SEQ ID NO: 1.
[0078] In some embodiments, the Cas protein further includes mutations in 12 or more key amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, D403, S404, C405, R406, F407, E408, C672, and P758 in SEQ ID NO: 1.
[0079] In some embodiments, the Cas protein further includes mutations in key amino acid sites related to cleavage activity, selected from the group consisting of K241, S242, S243, S246, D397, R398, E400, D401, E402, and L789 in SEQ ID NO: 1.
[0080] In some embodiments, the mutation is selected from the group consisting of the following:
[0081] (1) L272H + S275R + C277L; (2)L272H+S275R+C277L+E289A+I292L; (3)L272H+S275R+C277L+E289L+A290N+I292L; (4)L272H+S275R+C277L+E289L+A290H+I292V; (5)L272H+S275R+C277L+K399P+E400F+D401S+E402Y+D403P; (6)L272H+S275R+C277L+K399I+E400C+D401T+E402A+D403P; (7)L272H+S275R+C277L+K399V+E400P+D401V+E402L+D403S; (8)L272H+S275R+C277L+K399E+E400A+D401C+E402R+D403P; (9)L272H+S275R+C277L+K399R+E400V+D401G+E402T+D403R; (10)L272H+S275R+C277L+K399T+E400P+D401R+E402F+D403S; (11)L272H+S275R+C277L+K399P+E400A+D401R+E402F+D403T; (12)L272H+S275R+C277L+K399S+E400P+D401H+E402T+D403T; (13)L272H+S275R+C277L+S404Q+C405A+E408I; (14)L272H+S275R+C277L+S404R+C405R+R406T+E408L; (15)L272H+S275R+C277L+S404Y+C405K+R406G+F407P+E408V; (16)L272H+S275R+C277L+D403R+S404K+C405R+R406P+F407L+E408R; (17)L272H+S275R+C277L+D403S+C405P+R406P+F407R+E408V; (18)L272H+S275R+C277L+L394H+D395C+R396Y+D397F+R398S; (19)L272H+S275R+C277L+L394A+D395V+D397S+R398A; (20)L272H+S275R+C277L+D397S+R398I+K399D+E400S+D401K+E402I; (21)L272H+S275R+C277L+D397N+R398G+E400R+D401T+E402S; (22)L272H+S275R+C277L+D397H+R398A+K399S+E400A+D401M+E402S; (23)L272H+S275R+C277L+D397R+R398D+K399R+E400S+D401P+E402S; (24)L272H+S275R+C277L+D397N+R398G+E400R+D401T+E402S; (25)S242Y+S243A+S246R+L272H+S275R+C277L; (26)G119P+D120L+V121S+Q122K+Y123F+L272H+S275R+C277L; (27)E118D+G119S+D120K+V121S+Q122G+Y123F+L272H+S275R+C277L; (28)E118Y+G119S+D120G+V121P+Q122G+Y123L+L272H+S275R+C277L; (29)E118S+G119D+D120S+V121A+Q122F+Y123L+L272H+S275R+C277L; (30)E118K+G119S+D120T+V121G+Q122G+Y123I+L272H+S275R+C277L; (31)E118Q+G119S+D120S+V121T+Q122G+Y123F+L272H+S275R+C277L; (32)E118T+G119D+D120T+V121S+Q122K+Y123I+L272H+S275R+C277L; (33)E118I+G119P+D120E+V121P+Q122H+Y123F+L272H+S275R+C277L; (34)E118P+G119R+D120T+V121G+Q122R+L272H+S275R+C277L; (35)E118P+G119P+D120N+V121R+Q122R+Y123I+L272H+S275R+C277L; (36)E118S+G119M+D120S+Q122S+Y123T+L272H+S275R+C277L; (37)G119S+V121S+Q122K+Y123S+L272H+S275R+C277L; (38)E118S+G119D+V121E+Q122Y+Y123T+L272H+S275R+C277L; (39)G119L+D120S+V121S+Q122T+L272H+S275R+C277L; (40)L272H+S275R+C277L+D397N+R398G+E400R+D401T+E402S+L789S; (41)L272H+S275R+C277L+T285R+S286F+C287I+D288W+D403S+C405P+R406P+F407R+E408V+K502R+K557N; (42)L272H+S275R+C277L+E289R+I292L+K399S+E400P+D401H+E402T+D403T; (43)L272H+S275R+C277L+E289S+A290V+I292L+P758L+V860L; (44)L272H+S275R+C277L+K299R+R300D+W301F+C672R+P758L; (45)L272H+S275R+C277L+D397N+R398G+E400R+D401T+E402S+L657F+L789S; (46)E118P+G119R+D120T+V121G+Q122R+L272H+S275R+C277L+C672R+P758L; (47)L272H+S275R+C277L+K399L+E400T+D401F+E402G+D403G+P758L+K827E; (48)E118P+G119R+D120T+V121G+Q122R+L272H+S275R+C277L+L394R+D395S+R396T+D397T+R398N+C672R+P758L; (49)E118P+G119R+D120T+V121G+Q122R+L272H+S275R+C277L+L394P+D395P+R396G+D397F+C672R+P758L; (50)E118P+G119R+D120T+V121G+Q122R+L272H+S275R+C277L+D403S+C405R+R406L+F407S+E408R+C672R+P758L; (51)E118P+G119R+D120T+V121G+Q122R+L272H+S275R+C277L+D403S+S404P+C405L+R406H+F407M+E408V+C672R+P758L; (52)K241A+S242F+S243C+S246R+L272H+S275R+C277L+D397N+R398G+E400R+D401T+E402S+L789S (as numbered in Sequence ID: 1).
[0082] In some embodiments, the Cas protein variant has the same or substantially the same amino acid sequence as the wild-type Cas protein, except for the mutations (positions 118, 119, 120, 121, 122, 123, 241, 242, 243, 246, 272, 275, 277, 285, 286, 287, 288, 289, 290, 292, 299, 300, 301, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 502, 557, 657, 672, 758, 789, 860 and / or 827).
[0083] In some embodiments, being substantially identical means that at most 50 amino acids (preferably 1 to 20, more preferably 1 to 10, more preferably 1 to 5) are different, where "different" includes amino acid substitution, deletion, or addition, and the gene editing activity of the Cas protein variant is improved.
[0084] In some embodiments, the homology between the Cas protein mutant and the wild-type Cas protein is at least 80%, preferably at least 85% or 90%, more preferably at least 95%, and most preferably at least 98% or 99%.
[0085] In some embodiments, the Cas protein variant is formed by mutating the wild-type Cas protein.
[0086] A second aspect of the present invention provides a protein variant in which the variant is a non-natural protein and the variant comprises a mutation in one or more important amino acid sites related to cleavage activity, selected from the group consisting of the following in Sequence ID No. 1, which corresponds to the wild-type protein.
[0087] The 659th aspartic acid (D) site; and / or The 711th aspartic acid (D) site; and / or Glutamate (E) site 895; and / or The 1069th aspartic acid (D) site.
[0088] In some embodiments, the protein variant is a variant of an effector protein in the CRISPR / Cas system.
[0089] In another preferred example, the protein variant exhibits reduced (e.g., 50%, 60%, 70%, 80%, 90%, 95%, or more) cleavage activity of the target molecule complementary to the guide RNA sequence compared to the wild-type protein, or is substantially lost.
[0090] In some embodiments, the 659th aspartic acid (D) is mutated into any amino acid, preferably into the amino acid types shown in Tables A, B, and C, preferably into one or more amino acids selected from the group consisting of Ala(A), Val(V), Leu(L), and Ile(I), preferably into Ala(A) and Val(V), and more preferably into Ala(A).
[0091] In some embodiments, the 711th aspartic acid (D) is mutated into any amino acid, preferably into the amino acid types shown in Tables A, B, and C, preferably into one or more amino acids selected from the group consisting of Ala(A), Val(V), Leu(L), and Ile(I), preferably into Ala(A) and Val(V), and more preferably into Ala(A).
[0092] In some embodiments, the 895th glutamic acid (E) is mutated into any amino acid, preferably into the amino acid types shown in Tables A, B, and C, preferably into one or more amino acids selected from the group consisting of Ala(A), Val(V), Leu(L), and Ile(I), preferably into Ala(A) and Val(V), and more preferably into Ala(A).
[0093] In some embodiments, the 1069th aspartic acid (D) is mutated into any amino acid, preferably into the amino acid types shown in Tables A, B, and C, preferably into one or more amino acids selected from the group consisting of Ala(A), Val(V), Leu(L), and Ile(I), preferably into Ala(A) and Val(V), and more preferably into Ala(A).
[0094] In some embodiments, the 659th aspartic acid (D) is mutated to alanine (A).
[0095] In some embodiments, the 711th aspartic acid (D) is mutated to alanine (A).
[0096] In some embodiments, the 895th glutamic acid (E) is mutated to alanine (A).
[0097] In some embodiments, the 1069th aspartic acid molecule (D) is mutated to alanine (A).
[0098] In some embodiments, the mutation is selected from the group consisting of D659A, D711A, E895A, D1069A, or combinations thereof.
[0099] In some embodiments, the amino acid sequence of the protein variant is shown in any of SEQ ID NOs: 44-47.
[0100] In some embodiments, the protein variant is a polypeptide having the amino acid sequence shown in any of SEQ ID NOs: 44-47, an active fragment thereof, or a conserved mutant polypeptide thereof.
[0101] In some embodiments, the protein variant has the same or substantially the same amino acid sequence as the wild-type protein or its variant described in the first aspect of the present invention, except for the mutations (e.g., positions 659, 711, 895, and / or 1069).
[0102] In some embodiments, being substantially identical means that at most 50 amino acids (preferably 1 to 20, more preferably 2 to 18, more preferably 3 to 16) are different, where "different" includes amino acid substitution, deletion, or addition, and the cleavage activity of the protein variant is reduced.
[0103] In some embodiments, the homology between the mutant and the wild-type protein is at least 80%, preferably at least 85% or 90%, more preferably at least 95%, and most preferably at least 98% or 99%.
[0104] In some embodiments, the protein variant is selected from the following group.
[0105] (a) Polypeptides having any of the amino acid sequences shown in Sequence IDs 44-47; (b) A polypeptide derived from (a) having one or more (e.g., two, three, four, or five) amino acid residues substituted, deleted, or added in any of the amino acid sequences shown in SEQ ID NOs: 44-47, and having reduced cleavage activity.
[0106] In some embodiments, the homology between the derived polypeptide and the sequence shown in any of SEQ ID NOs: 44-47 is at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90%, for example, 95%, 97%, or 99%.
[0107] In some embodiments, the protein variant is formed by mutating the wild-type protein.
[0108] A third aspect of the present invention provides a fusion protein comprising a Cas protein described in the first aspect of the present invention or a protein variant described in the second aspect of the present invention and one or more functional domains.
[0109] In some embodiments, the functional domain is selected from localization signals, reporter proteins, Cas protein targeting moieties, DNA binding domains, epitope tags, transcriptional activation domains, transcriptional repression domains, nucleases, deaminase domains, methylases, demethylases, transcription termination factors, HDACs, polypeptides with cleavage activity, ligases, integrases, transposases, recombinases, polymerases, and base excision repair inhibitors (e.g., uracil-DNA glycosylase inhibitors (UGIs)).
[0110] In some embodiments, the functional domain includes one or more enzymatic activities among methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., derived from O-GlcNAc transferase), and deglycosylation activity against a target sequence.
[0111] In some embodiments, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.
[0112] In some embodiments, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.
[0113] In some embodiments, the adenosine deaminase catalytic domain includes an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in SEQ ID NO: 28 (a 005V1 deaminase selected from CN114634923A, which in this application has the amino acid sequence shown in SEQ ID NO: 2), and retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 28.
[0114] In some embodiments, the amino acid sequence of the adenosine deaminase catalytic domain has amino acid additions, insertions, deletions, and substitutions relative to the amino acid sequence shown in SEQ ID NO: 28.
[0115] In some embodiments, the adenosine deaminase catalytic domain includes Q148G+Q149M+P150R, a mutant of the amino acid sequence shown in SEQ ID NO: 28, and is named deaminase 005V1-10-3 (SEQ ID NO: 29).
[0116] In some embodiments, the adenosine deaminase catalytic domain includes an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in SEQ ID NO: 54 (a 004V1 deaminase selected from CN114634923A, which in this application has the amino acid sequence shown in SEQ ID NO: 1), and retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 54.
[0117] In some embodiments, the amino acid sequence of the adenosine deaminase catalytic domain has amino acid additions, insertions, deletions, and substitutions relative to the amino acid sequence shown in SEQ ID NO: 54.
[0118] In some embodiments, the functional domain is the full-length or functional fragment of TadA8e (SEQ ID NO: 55).
[0119] In some embodiments, the localization signal includes a nuclear localization signal (NLS) and / or a nuclear export signal (NES).
[0120] In some embodiments, the sequence of the nuclear localization signal is PKKKRKV (SEQ ID NO: 41) or as shown in SEQ ID NOs: 35-40, 42.
[0121] In some embodiments, the sequence of the nuclear localization signal is located at, near, or close to the terminal (e.g., N-terminus or C-terminus) of the Cas protein described in the first embodiment.
[0122] In some embodiments, the nuclear export signal includes protein tyrosine kinase 2 (e.g., human protein tyrosine kinase 2).
[0123] In some embodiments, the reporter protein includes glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, and an autofluorescent protein.
[0124] In some embodiments, the autofluorescent protein includes green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), and blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire).
[0125] In some embodiments, the DNA-binding domain includes a methylation-binding protein, LexADBD, and Gal4DBD.
[0126] In some embodiments, the epitope tag includes histidine tags, V5 tags, FLAG tags, influenza virus hemagglutinin tags, Myc tags, VSV-G tags, thioredoxin tags, and streptavidin tags.
[0127] In some embodiments, the transcriptional activation domain includes VP64 and / or VPR.
[0128] In some embodiments, the transcriptional repression domain includes KRAB and / or SID.
[0129] In some embodiments, the nuclease comprises FokI.
[0130] In some embodiments, the polypeptide having cleavage activity includes polypeptides having single-strand RNA cleavage activity, polypeptides having double-strand RNA cleavage activity, polypeptides having single-strand DNA cleavage activity, or polypeptides having double-strand DNA cleavage activity.
[0131] In some embodiments, the ligase includes DNA ligase and / or RNA ligase.
[0132] In some embodiments, the functional domain is ligated to the N-terminus and / or C-terminus of the Cas protein.
[0133] In some embodiments, the functional domain is inserted between the N-terminus and C-terminus of the Cas protein.
[0134] In some embodiments, the one or more functional domains are optionally linked to the N-terminus and / or C-terminus of the Cas protein via a linker.
[0135] In some embodiments, the functional domain is inserted between the N-terminus and C-terminus of the Cas protein via a linker.
[0136] In some embodiments, the fusion protein has the following structure from the N-terminus to the C-terminus.
[0137] Z1-Z2 (I); or Z2-Z1(II); or Z3-Z1-Z4(III); Here, Z1 is cytosine deaminase or adenosine deaminase, Z2 is the Cas protein described in the first aspect of the present invention or the Cas protein variant described in the second aspect of the present invention. Z3 is the N-terminal fragment of the Cas protein described in the first aspect of the present invention or the Cas protein variant described in the second aspect of the present invention. Z4 is the C-terminal fragment of the Cas protein described in the first aspect of the present invention or the Cas protein variant described in the second aspect of the present invention. And each "-" is an independent linker or connector.
[0138] In some embodiments, the fusion protein has the amino acid sequence shown in SEQ ID NO: 43.
[0139] A fourth aspect of the present invention provides an isolated polynucleotide that encodes a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, or a fusion protein according to the third aspect of the present invention.
[0140] In some embodiments, the polynucleotide is selected from the following group.
[0141] (a) A polynucleotide having the sequence shown in either SEQ ID NO: 2 or 34; (b) A polynucleotide comprising a nucleotide sequence having 70% or more homology (preferably 80% or more, more preferably 90% or more, more preferably 95% or more, most preferably 99% or more) to the sequence shown in either SEQ ID NO: 2 or 34, and encoding a polypeptide shown in either SEQ ID NO: 1 or 43; (c) A polynucleotide complementary to any of the polynucleotides described in (a) to (b).
[0142] In some embodiments, the isolated nucleotides include an optimized humanized sequence.
[0143] In some embodiments, the polynucleotide further includes, in the flanking region of the ORF of the variant, an auxiliary element selected from the group consisting of a signal peptide, a secreted peptide, a tag sequence (e.g., 6His), or a combination thereof.
[0144] In some embodiments, the polynucleotide is selected from the group consisting of a genome sequence, a cDNA sequence, an RNA sequence, or a combination thereof.
[0145] In some embodiments, the polynucleotide further comprises a promoter operably ligated to the ORF sequence of the variant.
[0146] In some embodiments, the promoter is selected from the group consisting of a constitutive promoter, a tissue-specific promoter, an inducible promoter, or a strong promoter.
[0147] In some embodiments, the polynucleotide is a codon-optimized polynucleotide in accordance with the codon bias of the host cell.
[0148] In some embodiments, the host cells include prokaryotic cells or eukaryotic cells.
[0149] In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).
[0150] In some embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.
[0151] In some embodiments, the yeast cells are yeasts of one or more origins selected from the group consisting of Pichia yeast, Cluyveromyces, or a combination thereof, and preferably the yeast cells include Cluyveromyces, more preferably Cluyveromyces marcyanas, and / or Cluyveromyces lactis.
[0152] In some embodiments, the host cells are selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or combinations thereof.
[0153] A fifth aspect of the present invention includes or is composed of an array selected from the following: (i) Sequence ID: the sequence shown in 5; (ii) Sequences in which one or more bases are substituted, deleted, or added to the sequence shown in Sequence ID No. 5 (for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases are substituted, deleted, or added); (iii) Sequences having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, and at least 95% sequence identity with the sequence shown in sequence number 5; (iv) A sequence that hybridizes with any one of the sequences described in (i) to (iii) under stringent conditions; or (v) A sequence complementary to any one of the sequences described in (i) to (iii); Furthermore, any sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it originates. The isolated nucleic acid molecule is, for example, RNA. The isolated nucleic acid molecule provides an isolated nucleic acid molecule that includes, for example, a direct repeat sequence in a CRISPR / Cas system.
[0154] In some embodiments, the nucleic acid molecule comprises one or more stem-loops or optimized secondary structures.
[0155] For example, any sequence described in any one of (ii) to (v) retains the secondary structure of the sequence from which it originates.
[0156] In some embodiments, the nucleic acid molecule is a nucleic acid molecule that is codon-optimized according to the codon bias of the host cell.
[0157] In some embodiments, the nucleic acid molecule includes or is composed of a sequence selected from the following:
[0158] (a) Nucleotide sequence shown in Sequence ID: 5; (b) A sequence that hybridizes with the sequence described in (a) under stringent conditions; or (c) Sequence number: A sequence complementary to the nucleotide sequence shown in 5.
[0159] A sixth aspect of the present invention provides a guide RNA (gRNA) comprising a direct repeat (DR) sequence capable of binding to the Cas protein described in the first aspect of the present invention and a spacer sequence capable of targeting a target sequence.
[0160] A seventh aspect of the present invention is: (i) A protein component selected from the group consisting of the Cas protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, or a combination thereof; and (ii) A nucleic acid component selected from the group consisting of the guide RNA described in the sixth aspect of the present invention, the nucleic acid encoding the guide RNA described in the sixth aspect of the present invention, the guide RNA precursor RNA described in the sixth aspect of the present invention, the nucleic acid encoding the precursor RNA of the guide RNA described in the sixth aspect of the present invention, or a combination thereof. Includes, Here, the protein component and the nucleic acid component are bound to each other to form a complex. Provide a composite.
[0161] In some embodiments, the direct repeat (DR) sequence in the guide RNA (gRNA) is ligated to the 3' or 5' end of the nucleic acid molecule.
[0162] In some embodiments, the spacer sequence in the guide RNA (gRNA) includes a sequence complementary to the target sequence.
[0163] An eighth aspect of the present invention provides a vector comprising a polynucleotide as described in the fourth aspect of the present invention or a nucleic acid molecule as described in the fifth aspect of the present invention.
[0164] In some embodiments, the vector is (1) A first regulatory element operably linked to a nucleotide sequence encoding a Cas protein according to a first aspect of the present invention, a nucleotide sequence encoding a protein variant according to a second aspect of the present invention, or a nucleotide sequence encoding a fusion protein according to a third aspect of the present invention; and (2) A second regulatory element operably ligated to the nucleotide sequence encoding the guide RNA. Includes, The aforementioned guide RNA is (a) A spacer sequence that can hybridize with the target sequence, and (b) Direct Repeat (DR) sequences linked to the spacer sequence, which guide the Cas protein variant described in the first aspect of the present invention or the protein variant described in the second aspect of the present invention to bind to the guide RNA and target the target sequence, forming a complex according to the seventh aspect of the present invention. Includes.
[0165] In some embodiments, the first tuning element and the second tuning element are located on the same or different vectors.
[0166] In some embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inductive promoter, a constitutive promoter, a ubiquitous promoter, or a cell type or tissue-specific promoter.
[0167] In some embodiments, the vector comprises one or more promoters, the promoters operably linked to the nucleic acid sequence, enhancers, transcription termination signals, polyadenylation sequences, replication start sites, selection markers, nucleic acid restriction sites, and / or homologous recombination sites.
[0168] In some embodiments, the vector includes plasmids and viral vectors.
[0169] In some embodiments, the viral vector is selected from the group consisting of adeno-associated viruses (AAV), adenoviruses, lentiviruses, retroviruses, herpesviruses, SV40, poxviruses, or combinations thereof.
[0170] In some embodiments, the vectors include cloning vectors, transformation vectors, expression vectors, shuttle vectors, embedded vectors, and multifunctional vectors.
[0171] A ninth aspect of the present invention is: (i) A first component selected from the group consisting of a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, a fusion protein according to the third aspect of the present invention, a nucleotide sequence encoding a Cas protein according to the first aspect of the present invention or a protein variant according to the second aspect of the present invention or a fusion protein according to the second aspect of the present invention, and any combination thereof; and (ii) A second component which is a nucleotide sequence comprising one or more guide RNAs described in the sixth aspect of the present invention, or a nucleotide sequence encoding the one or more guide RNAs described in the sixth aspect of the present invention. Includes, The guide RNA can form a complex with the protein or fusion protein described in (i). A CRISPR-Cas composition is provided.
[0172] In some embodiments, the guide RNA includes a direct repeat sequence and a spacer sequence in the 5' to 3' direction, the spacer sequence being hybridizable with the target sequence.
[0173] In some embodiments, the composition further comprises a pharmaceutically acceptable carrier.
[0174] In some embodiments, the composition includes a pharmaceutical composition.
[0175] In some embodiments, the dosage form of the composition is selected from the group consisting of lyophilized formulations, liquid formulations, or combinations thereof.
[0176] In some embodiments, the dosage form of the composition is a liquid formulation.
[0177] In some embodiments, the dosage form of the composition is an injectable preparation.
[0178] In some embodiments, the composition is a cell preparation.
[0179] A tenth aspect of the present invention includes one or more vectors, wherein the one or more vectors are (i) a nucleotide sequence encoding a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, or a fusion protein according to the third aspect of the present invention, which is optionally operably linked to the first regulatory element; and (ii) A second nucleic acid encoding a nucleotide sequence containing the guide RNA described in the sixth aspect of the present invention, which is optionally operably linked to the second regulatory element. Includes, Here, The first nucleic acid and the second nucleic acid are present in the same or different vectors. The guide RNA can form a complex with the protein or fusion protein described in (i). We provide CRISPR-Cas systems.
[0180] In some embodiments, the vector includes plasmids and viral vectors.
[0181] In some embodiments, the guide RNA includes a spacer sequence that can hybridize with a target sequence; and a direct repeat (DR) sequence that is linked to the spacer sequence and forms a CRISPR-Cas composition or complex that guides the protein to bind to the guide RNA and target the target sequence.
[0182] In some embodiments, the guide RNA includes unmodified guide RNA and modified guide RNA.
[0183] In some embodiments, the modified guide RNA includes chemical modifications of bases.
[0184] In some embodiments, the chemical modification includes methylation modification, methoxy modification, fluorination modification, or thio modification.
[0185] In some embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inductive promoter.
[0186] In some embodiments, at least one component in the composition is either unnaturally occurring or modified.
[0187] In some embodiments, the spacer array is attached to the 3' end of the direct repeat (DR) array.
[0188] In some embodiments, the spacer sequence includes a sequence complementary to the target sequence.
[0189] In some embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of a protospacer adjacent motif (PAM), and the PAM has a sequence in which the 5'-PAM is shown as TTN, where N is A, T, C, or G.
[0190] In some embodiments, the target sequence is DNA derived from prokaryotic or eukaryotic cells, or a DNA sequence formed by reverse transcription of RNA, or the target sequence is DNA that exists in nature, or a DNA sequence formed by reverse transcription of RNA.
[0191] In some embodiments, the target sequence includes a cDNA sequence.
[0192] In some embodiments, the target sequence includes a single-stranded DNA or a double-stranded DNA sequence.
[0193] In some embodiments, the target sequence is located inside a cell.
[0194] In some embodiments, the target sequence is located within the cell nucleus or cytoplasm (e.g., in an organelle).
[0195] In some embodiments, the cells are eukaryotic cells.
[0196] In some embodiments, the cells are prokaryotic cells.
[0197] In some embodiments, the target sequence is located extracellularly.
[0198] In some embodiments, the Cas protein described in the first aspect of the present invention is ligated to one or more NLS sequences, or the fusion protein comprises one or more NLS sequences.
[0199] In some embodiments, the NLS sequence is ligated to the N-terminus or C-terminus of the Cas protein described in the first aspect of the present invention.
[0200] In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the Cas protein described in the first aspect of the present invention.
[0201] An eleventh aspect of the present invention provides a kit comprising one or more components selected from the Cas protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the ninth aspect of the present invention, or the system described in the tenth aspect of the present invention.
[0202] In some embodiments, the kit further includes a label or instructions.
[0203] In some embodiments, the kit is used for one or more of the following: gene or genome editing, disease treatment, targeting of a target gene, or cutting of a target gene or a non-target gene.
[0204] A twelfth aspect of the present invention provides a delivery vector and a delivery composition comprising one or more selected from the Cas protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the ninth aspect of the present invention, or the system described in the tenth aspect of the present invention.
[0205] In some embodiments, the delivery vector is a particle.
[0206] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microbubbles, gene guns, or viral vectors (e.g., replication-deficient retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0207] A thirteenth aspect of the present invention provides a host cell comprising a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, a fusion protein according to the third aspect of the present invention, a polynucleotide according to the fourth aspect of the present invention, a complex according to the seventh aspect of the present invention, a vector according to the eighth aspect of the present invention, a composition according to the ninth aspect of the present invention, a system according to the tenth aspect of the present invention, or a delivery composition according to the twelfth aspect of the present invention.
[0208] In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).
[0209] In some embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.
[0210] In some embodiments, the yeast cells are yeasts of one or more origins selected from the group consisting of Pichia yeast, Cluyveromyces, or a combination thereof, and preferably the yeast cells include Cluyveromyces, more preferably Cluyveromyces marcyanas, and / or Cluyveromyces lactis.
[0211] In some embodiments, the host cells are selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or combinations thereof.
[0212] A fourteenth aspect of the present invention provides an enzyme preparation comprising a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, a fusion protein according to the third aspect of the present invention, a complex according to the seventh aspect of the present invention, a CRISPR-Cas composition according to the ninth aspect of the present invention, a system according to the tenth aspect of the present invention, or a delivery composition according to the twelfth aspect of the present invention.
[0213] In some embodiments, the enzyme preparation includes an injectable preparation and / or a lyophilized preparation.
[0214] A fifteenth aspect of the present invention provides a pharmaceutical kit comprising a first container and a compound according to the seventh aspect of the present invention, a composition according to the ninth aspect of the present invention, or a system according to the tenth aspect of the present invention, which are disposed within the first container, or a compound according to the seventh aspect of the present invention, a composition according to the ninth aspect of the present invention, or a system according to the tenth aspect of the present invention.
[0215] In some embodiments, the pharmaceutical product in the first container is a single-component formulation comprising the complex described in the seventh aspect of the present invention, the composition described in the ninth aspect of the present invention, or the system described in the tenth aspect of the present invention.
[0216] In some embodiments, the dosage form of the pharmaceutical product is selected from the group consisting of lyophilized formulations, liquid formulations, or combinations thereof.
[0217] In some embodiments, the dosage form of the pharmaceutical product is an oral dosage form or an injectable dosage form.
[0218] In some embodiments, the pharmaceutical kit further includes instructions.
[0219] A sixteenth aspect of the present invention is: (a1) A first container, and a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, a fusion protein according to the third aspect of the present invention, or their coding genes or expression vectors, which are placed inside the first container; or a pharmaceutical product containing a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, a fusion protein according to the third aspect of the present invention, or their coding genes or expression vectors; (b1) A second container of any choice, and a guide RNA or expression vector thereof according to the sixth aspect of the present invention, or a pharmaceutical product containing the guide RNA or expression vector thereof according to the sixth aspect of the present invention, which is placed in the second container. We provide a pharmaceutical kit that includes the following:
[0220] In some embodiments, the first container and the second container are different containers.
[0221] In some embodiments, the pharmaceutical product in the first container is a single-component formulation containing the Cas protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, or their coding genes or expression vectors.
[0222] In some embodiments, the pharmaceutical product in the second container is a single-component formulation containing the guide RNA or expression vector described in the sixth aspect of the present invention.
[0223] In some embodiments, the dosage form of the pharmaceutical product is selected from the group consisting of lyophilized formulations, liquid formulations, or combinations thereof.
[0224] In some embodiments, the dosage form of the pharmaceutical product is an oral dosage form or an injectable dosage form.
[0225] In some embodiments, the pharmaceutical kit further includes instructions.
[0226] A 17th aspect of the present invention provides a method for targeting and editing or cleaving a target gene, comprising contacting the target gene with the Cas protein described in the 1st aspect of the present invention, the protein variant described in the 2nd aspect of the present invention, the fusion protein described in the 3rd aspect of the present invention, the complex described in the 7th aspect of the present invention, the composition described in the 9th aspect of the present invention, the system described in the 10th aspect of the present invention, the delivery composition described in the 12th aspect of the present invention, the enzyme preparation described in the 14th aspect of the present invention, or the pharmaceutical kit described in the 15th or 16th aspect of the present invention, wherein the target gene contains a target sequence.
[0227] In some embodiments, the target gene is located within a cell.
[0228] In some embodiments, the cells are prokaryotic cells.
[0229] In some embodiments, the cells are eukaryotic cells, such as mammalian cells (e.g., human cells) or plant cells.
[0230] In some embodiments, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro.
[0231] In some embodiments, editing or cleaving the target gene includes cleaving a target sequence, for example, cleaving a double strand of DNA or a single strand of RNA, or inserting an exogenous nucleic acid into the cleavage site.
[0232] In some embodiments, the target gene includes DNA.
[0233] In some embodiments, the DNA includes single-stranded DNA or double-stranded DNA.
[0234] An 18th aspect of the present invention provides a method for inducing a change in cellular state, comprising contacting a target gene in a cell with a Cas protein described in the 1st aspect of the present invention, a protein variant described in the 2nd aspect of the present invention, a fusion protein described in the 3rd aspect of the present invention, a complex described in the 7th aspect of the present invention, a composition described in the 9th aspect of the present invention, a system described in the 10th aspect of the present invention, a delivery composition described in the 12th aspect of the present invention, an enzyme preparation described in the 14th aspect of the present invention, or a pharmaceutical kit described in the 15th or 16th aspect of the present invention.
[0235] A 19th aspect of the present invention provides a method for modifying the expression of a gene product, comprising contacting a nucleic acid molecule encoding the gene product with a Cas protein described in the 1st aspect of the present invention, a protein variant described in the 2nd aspect of the present invention, a fusion protein described in the 3rd aspect of the present invention, a complex described in the 7th aspect of the present invention, a composition described in the 9th aspect of the present invention, a system described in the 10th aspect of the present invention, a delivery composition described in the 12th aspect of the present invention, an enzyme preparation described in the 14th aspect of the present invention, or a pharmaceutical kit described in the 15th or 16th aspect of the present invention, or delivering the nucleic acid molecule containing the nucleic acid molecule, wherein the nucleic acid molecule contains the target sequence.
[0236] In some embodiments, the nucleic acid molecule is present in an in vitro nucleic acid molecule (e.g., a plasmid).
[0237] In some embodiments, the expression of the gene product is modified (e.g., enhanced or reduced).
[0238] In some embodiments, the gene product is a protein.
[0239] In some embodiments, the protein, fusion protein, polynucleotide, isolated nucleic acid molecule, complex, vector, or composition is contained within the delivery vector.
[0240] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, and viral vectors (e.g., replication-deficient retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0241] In some embodiments, cells, cell lines, or organisms are modified by altering one or more target sequences in a nucleic acid molecule encoding a target gene or target gene product.
[0242] A 20th aspect of the present invention provides cells or their offspring obtained by a method according to any one of the 17th to 19th aspects of the present invention, which include modifications not present in the wild type, or modifications not present in the corresponding cells not modified by the said method, or in which abnormal traits are resolved or improved in the modified cells.
[0243] A 21st aspect of the present invention provides a cell product of a cell or its progeny described in a 20th aspect of the present invention.
[0244] A 22nd aspect of the present invention provides in-vitro, ex-vivo, or in-vivo cells or cell lines or their offspring comprising a Cas protein as described in the 1st aspect of the present invention, a protein variant as described in the 2nd aspect of the present invention, a fusion protein as described in the 3rd aspect of the present invention, a polynucleotide as described in the 4th aspect of the present invention, a complex as described in the 7th aspect of the present invention, a vector as described in the 8th aspect of the present invention, a CRISPR-Cas composition as described in the 9th aspect of the present invention, a system as described in the 10th aspect of the present invention, or a delivery composition as described in the 12th aspect of the present invention.
[0245] In some embodiments, the cells are prokaryotic cells.
[0246] In some embodiments, the cells are eukaryotic cells, such as mammalian cells (e.g., human cells) or plant cells.
[0247] In some embodiments, the cells are stem cells or stem cell lines.
[0248] A 23rd aspect of the present invention provides a use of the Cas protein described in the 1st aspect of the present invention, the protein variant described in the 2nd aspect of the present invention, the fusion protein described in the 3rd aspect of the present invention, the polynucleotide described in the 4th aspect of the present invention, the complex described in the 7th aspect of the present invention, the vector described in the 8th aspect of the present invention, the CRISPR-Cas composition described in the 9th aspect of the present invention, the system described in the 10th aspect of the present invention, the kit described in the 11th aspect of the present invention, the delivery composition described in the 12th aspect of the present invention, the enzyme preparation described in the 14th aspect of the present invention, or the pharmaceutical kit described in the 15th or 16th aspect of the present invention, which is used in the manufacture of a pharmaceutical or preparation for nucleic acid editing (e.g., gene or genome editing).
[0249] In some embodiments, the gene or genome editing includes gene modification, gene knockout, alteration of gene product expression, mutation repair, and / or polynucleotide insertion.
[0250] A 24th aspect of the present invention provides a use of a Cas protein described in the first aspect of the present invention, a protein variant described in the second aspect of the present invention, a fusion protein described in the third aspect of the present invention, a polynucleotide described in the fourth aspect of the present invention, a complex described in the seventh aspect of the present invention, a vector described in the eighth aspect of the present invention, a CRISPR-Cas composition described in the ninth aspect of the present invention, a system described in the tenth aspect of the present invention, a kit described in the eleventh aspect of the present invention, a delivery composition described in the twelfth aspect of the present invention, an enzyme preparation described in the fourteenth aspect of the present invention, or a pharmaceutical kit described in the fifteenth or sixteenth aspect of the present invention, which is used in the manufacture of a pharmaceutical or preparation for one or more items selected from the following group.
[0251] (i) ex vivo gene or genome editing; (ii) Detection of single-stranded DNA in ex vivo; (iii) Modification of living or non-human organisms by editing target sequences at target gene loci; (iv) Treatment of disorders caused by deletion of target sequences at target gene loci; (v) Treatment of the disability or disease of the subject in need.
[0252] In some embodiments, the disorder or disease includes cancer, infectious diseases, neurological diseases, ophthalmic diseases, and hearing diseases.
[0253] In some embodiments, the disease or disorder is cystic fibrosis, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy, α1-antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington's disease, Fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell anemia, β-thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, hypercholesterolemia, transthyretin amyloidosis, retinopathy, macular degeneration. This includes Wilms' tumor, Ewing's sarcoma, neuroendocrine tumors, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, multiple myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and bladder cancer.
[0254] In some embodiments, the disorder or disease is caused by a pathogenic point mutation.
[0255] A 25th aspect of the present invention provides a method for detecting the presence of a target nucleic acid molecule in a sample, comprising contacting a sample with a Cas protein according to the first aspect of the present invention, a protein variant according to the second aspect of the present invention, a fusion protein according to the third aspect of the present invention, a complex according to the seventh aspect of the present invention, a CRISPR-Cas composition according to the ninth aspect of the present invention, a system according to the tenth aspect of the present invention, a kit according to the eleventh aspect of the present invention, a delivery composition according to the twelfth aspect of the present invention, or an enzyme preparation according to the fourteenth aspect of the present invention and a non-target sequence, and detecting a detectable signal produced by cleaving the non-target sequence, wherein the non-target sequence does not hybridize with guide RNA.
[0256] In some embodiments, when the non-target sequence is cleaved by a protein in the complex, CRISPR-Cas composition, system, or delivery composition, it indicates that the target nucleic acid molecule is present in the sample; on the other hand, when the non-target sequence is not cleaved by a protein in the complex, CRISPR-Cas composition, system, or delivery composition, it indicates that the target nucleic acid molecule is not present in the sample.
[0257] In some embodiments, the target nucleic acid molecule is target DNA.
[0258] In some embodiments, the target DNA includes DNA formed by RNA reverse transcription.
[0259] In some embodiments, the target DNA includes cDNA.
[0260] In some embodiments, the target DNA is selected from the group consisting of single-stranded DNA, double-stranded DNA, or a combination thereof.
[0261] It should be understood that each of the above-described constituent elements of the present invention and each of the constituent elements specifically described below (for example, in the embodiments) can be combined with each other within the scope of the present invention to constitute a new technical solution or a preferred technical solution. Due to the limitation of the paper space, detailed descriptions of individual items are omitted here.
Brief Description of Drawings
[0262] [Figure 1A] Shows the CasY6 recombinant expression plasmid map. [Figure 1B] Shows the LbCpf1 recombinant expression map. [Figure 2] Shows the secondary structure prediction of the direct repeat (DR) sequence corresponding to CasY6. [Figure 3] Shows the Target plasmid map. [Figure 4] Shows the comparison of the editing efficiencies of CasY6 and LbCpf1 in Escherichia coli. [Figure 5] Shows the PHK09T plasmid map. [Figure 6] Shows the comparison of the editing efficiencies of CasY6 and LbCpf1 in HEK293T cells. [Figure 7] Shows four mutants of CasY6 with lost catalytic activity (cleavage activity) compared to the wild-type CasY6 protein. [Figure 8A] Figure 8 shows the detection of base editing activity at the target site of the base editor containing dCasY6. Figure 8A shows the single-base editing (A>G) efficiency at the A2, A4, A13, A15-A17 sites of the base editor composed of dCasY6. [Figure 8B] Figure 8B shows the single-base editing (A>G) efficiency at the A7, A13, A20 sites of the base editor composed of dCasY6. [Figure 9] Shows the dual luciferase reporter plasmid map. [Figure 10] Shows the flowchart for measuring the cleavage activity of CasY6 mutants by the dual luciferase reporter system. [Modes for carrying out the invention]
[0263] As a result of diligent research, the inventors have unexpectedly discovered a novel Cas protein. The Cas protein of the present invention has excellent gene editing activity and can effectively edit or cleave target genes, thereby effectively treating disorders or diseases in subjects requiring such treatment. Compared to Cas enzymes disclosed in the prior art, the Cas protein of the present invention has superior editing efficiency and provides more options for base editing tools. Base editors constructed with the Cas enzymes disclosed in the present invention can effectively perform base editing and have potential applications. Furthermore, the present invention has discovered a novel Cas protein variant for the first time. The Cas protein variant of the present invention has gene editing activity equivalent to or better than that of the wild-type Cas protein, and can effectively edit or cleave target genes, thereby effectively treating disorders or diseases in subjects requiring such treatment. Based on this finding, the inventors have completed the present invention.
[0264] term The following examples are for illustrative purposes only and do not limit the present invention. Unless otherwise specified, the experiments and methods described in the examples are generally carried out in accordance with methods well known in the art and conventional methods described in various references.
[0265] Furthermore, in the examples where specific conditions are not explicitly stated, the procedures are carried out under normal conditions or conditions recommended by the manufacturer. Reagents and equipment whose manufacturers are not specified are all common products available commercially. Those skilled in the art will understand that while the examples illustrate the present invention, they are not intended to limit the scope of protection of the present invention. All disclosures and other reference materials referred to herein are incorporated herein by reference in their entirety.
[0266] To facilitate understanding of this disclosure, we first define some terms. Unless otherwise specified, the following terms should have the meanings set forth below, as used in this application. Other definitions are provided throughout this application.
[0267] The term “approximately” may also refer to a value or composition within an acceptable margin of error for a particular value or composition as determined by those skilled in the art, and depending in part to how that value or composition is measured or determined. For example, as used herein, the expression “approximately 100” includes 99, 101 and all values in between (e.g., 99.1, 99.2, 99.3, 99.4, etc.).
[0268] As used herein, the terms “contain” or “include” may be open, semi-closed, or closed. In other words, these terms also include “substantially composed of” or “composed of.”
[0269] Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window (which may be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of the reference nucleotide sequence or protein) and determining the number of positions in which the same residues appear. This is usually expressed as a percentage. The method for measuring sequence identity of nucleotide sequences is well known to those skilled in the art.
[0270] Wild-type Cas protein As used herein, “wild-type Cas protein” refers to a naturally occurring, unmodified Cas protein whose nucleotide sequence can be obtained by genetic engineering techniques such as genome sequencing or polymerase chain reaction (PCR), and whose amino acid sequence can be inferred from the nucleotide sequence. In one preferred example of the present invention, the wild-type Cas protein is CasY6, whose sequence is shown in Sequence ID No. 1.
[0271] Cas protein In the present invention, the Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably. The Cas protein is used in the broadest sense and includes wild-type Cas proteins, their derivatives or mutants, analogs, and functional fragments thereof, such as oligonucleotide-binding fragments.
[0272] The term "wild-type" has the meaning generally understood by those skilled in the art and represents the typical form of an organism, strain, gene, protein, or a characteristic that is distinguishable from mutant or variant forms when present in nature, can be isolated from a natural source, and has not been artificially and intentionally modified.
[0273] The terms "mutant", "derivative", and "analog" refer to polypeptides that substantially retain the function or activity of the Cas protein of the present invention.
[0274] Generally, derivatization of a protein does not adversely affect the desired activity of the protein (e.g., the activity of binding to guide RNA, endonuclease activity, the activity of binding and cleaving a specific site of a target sequence under the instruction of guide RNA), that is, the derivative of the protein has the same activity as the protein. The modified form of the "derivative" includes that one or more amino acids of the protein may be deleted, inserted, modified, and / or substituted. The terms "non-naturally occurring" or "engineered" can be used interchangeably and indicate artificial intervention.
[0275] In one aspect, the present invention provides a Cas protein comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence of SEQ ID NO: 1 and substantially retaining the biological function of the sequence from which it is derived.
[0276] In one embodiment, the amino acid sequence of the Cas protein has one or more amino acid substitutions, deletions, or additions compared to the amino acid sequence of SEQ ID NO: 1, and substantially retains the biological function of the sequence from which it is derived.
[0277] In one embodiment, the Cas protein comprises the amino acid sequence shown in Sequence ID No. 1, Alternatively, the sequence may include one or more amino acids substituted, deleted, or added (for example, one, two, three, four, five, six, seven, eight, nine, or ten amino acids substituted, deleted, or added) compared to the sequence shown in Sequence ID No. 1, or it may include a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with the amino acid sequence shown in Sequence ID No. 1.
[0278] In one embodiment, the protein has the amino acid sequence shown in Sequence ID No. 1.
[0279] As will be apparent to those skilled in the art, the structure of a protein can be altered without adversely affecting its activity and functionality. For example, one or more conserved amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule.
[0280] Those skilled in the art are familiar with examples and embodiments of conservative amino acid substitutions. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted; that is, a nonpolar amino acid residue can be substituted with another nonpolar amino acid residue, a polar uncharged amino acid residue can be substituted with another polar uncharged amino acid residue, a basic amino acid residue can be substituted with another basic amino acid residue, and an acidic amino acid residue can be substituted with another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, such as the substitution of an amino acid residue with another amino acid belonging to the same group, are within the scope of the present invention, as long as the substitution does not inactivate the biological activity of the protein. Accordingly, the protein of the present invention may contain one or more conservative substitutions in its amino acid sequence, and these conservative substitutions are preferably made according to Table A. Furthermore, the present invention also includes proteins containing one or more other nonconservative substitutions, as long as they do not significantly affect the desired function and biological activity of the protein of the present invention.
[0281] Conservative amino acid substitutions can be made at one or more predicted non-essential amino acid residues. "Non-essential" amino acid residues are those that can be modified (deleted, substituted, or added) without altering biological activity, while "essential" amino acid residues are those required for biological activity. A "conservative amino acid substitution" is one in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can be made in non-conservative regions of Cas enzymes. Generally, such substitutions are not made at conservative amino acid residues or those located within conservative motifs, because such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have a small number of conservative or non-conservative modifications in conserved regions. [Table 1] Those skilled in the art know that one or more amino acid residues from the N-terminus and / or C-terminus of a protein can be modified (substituted, deleted, truncated, or inserted) while retaining its functional activity. Therefore, a protein having one or more amino acid residues modified from the N-terminus and / or C-terminus of the Cas protein while retaining its desired functional activity is also within the scope of the present invention. These modifications may include modifications introduced by modern molecular biological methods, such as PCR. Such methods include PCR amplification, which modifies or extends a protein-coding sequence by incorporating the amino acid-coding sequence into oligonucleotides used for PCR amplification.
[0282] Proteins can be modified in a variety of ways, including amino acid substitution, deletion, truncation, and insertion, and methods for such operations are generally well known to those skilled in the art.
[0283] For example, amino acid sequence variants of the Cas protein can be prepared by inducing mutations in DNA. This can also be achieved by other forms of mutagenesis and / or directed evolution. For instance, known mutagenesis, recombination, and / or shuffling methods can be used in combination with relevant screening methods to induce one or more amino acid substitutions, or one to multiple amino acid deletions and / or one to multiple amino acid insertions.
[0284] Those skilled in the art will understand that these minute amino acid changes in the Cas protein of the present invention can occur (e.g., naturally occurring mutations) or be produced (e.g., using r-DNA technology) without causing loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domain of the protein, the properties of the polypeptide will change, but the polypeptide activity can be preserved. If the existing mutations are not in close proximity to the catalytic domain, active site, or other functional domain, their impact is expected to be small.
[0285] Those skilled in the art can identify the essential amino acids of the Cas protein by methods known in the art, such as site-directed mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domain, active site, or other functional domain of the protein can also be determined by physical analysis of its structure, for example, by techniques such as nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, in combination with mutations of amino acids in the predicted key sites.
[0286] orthologue (ortholog) As used herein, the term “orthologue” has the meaning commonly understood by those skilled in the art. Further explanation is provided that the “orthologues” of proteins described herein are proteins belonging to different species, and that the protein and its orthologue perform the same or similar functions.
[0287] The nucleic acid cleavage of the present invention includes cleavage of DNA or RNA in a target nucleic acid by the Cas protein (Cis cleavage), and cleavage of the collateral nucleic acid substrate (single-stranded nucleic acid substrate) of DNA or RNA by the collateral cleavage activity of the Cas protein (i.e., nonspecific or non-targeted cleavage, trans cleavage, or collateral cleavage activity). In some embodiments, the cleavage is of double-stranded DNA. In some embodiments, the cleavage is of single-stranded DNA or single-stranded RNA.
[0288] Trans cleavage refers to the phenomenon where, under certain conditions, activated Cas12 family proteins maintain their activity after binding to a target sequence, continuing to nonspecifically cleave non-target oligonucleotides. This collateral cleavage activity makes it possible to detect the presence of specific target oligonucleotides using the Cas system. For example, the Cas12i system can be genetically modified to nonspecifically cleave ssDNA or transcripts. Collateral cleavage activity is used in a highly sensitive and specific nucleic acid detection platform called SHERLOCK, and is available for use in many clinical diagnostics (Gootenberg, JS et al., Nucleic acid detection with CRISPR-Cas13a / C2c2. Science 356, 438-442 (2017)).
[0289] Cas protein variants and their encoding nucleic acids As used herein, the terms “Cas protein variant,” “variant of the present invention,” “gene-edited mutant protein of the present invention,” and “mutant protein” are all interchangeable and refer to a non-naturally occurring mutant Cas protein, wherein the mutant protein has a mutation in one or more important amino acid sites related to cleavage activity, selected from the following group in Sequence ID No. 1, which corresponds to the wild-type Cas protein.
[0290] E118, G119, D120, V121, Q122, Y123, K241, S242, S243, S246, L272, S275, C277, T285, S286, C287, D288, E289, A290, I292, K299, R300, W301, L394, D395, R396, D397, R398, K399, E400, D401, E402, D403, S404, C405, R406, F407, E408, K502, K557, L657, C672, P758, L789, K827 and V860.
[0291] The term "key amino acid" refers to a sequence based on the wild-type Cas protein and having at least 80%, for example 84%, 85%, 90%, 92%, 95%, 98%, or 99% homology to the wild-type Cas protein, where the corresponding site is one of the specific amino acids described herein. For example, based on the wild-type gene-edited protein, the key amino acids are as follows:
[0292] E118, G119, D120, V121, Q122, Y123, K241, S242, S243, S246, L272, S275, C277, T285, S286, C287, D288, E289, A290, I292, K299, R300, W301, L394, D395, R396, D397, R398, K399, E400, D401, E402, D403, S404, C405, R406, F407, E408, K502, K557, L657, C672, P758, L789, K827 and V860.
[0293] Furthermore, mutant proteins obtained by mutating the above-mentioned important amino acids either have higher cleavage activity than the wild-type Cas protein (SEQ ID NO: 1), or the cleavage activity of the mutant protein is equivalent to that of the wild-type Cas protein. Alternatively, CRISPR compositions containing Cas protein mutants exhibit enhanced binding to the binding site or altered editing bias.
[0294] Preferably, in the present invention, the important amino acids of the present invention are mutated as follows.
[0295] E118D、E118K、E118I、E118Q、E118P、E118S、E118T、E118Y、G119D、G119M、G119L、G119P、G119R、G119S、D120E、D120K、D120L、D120N、D120G、D120S、D120T、V121A、V121E、V121G、V121P、V121R、V121S、V121T、Q122K、Q122F、Q122K、Q122H、Q122G、Q122R、Q122S、Q122T、Q122Y、Y123F、Y123L、Y123I、Y123S、Y123T、K241A、S242F、S242Y、S243A、S243C、S246R、L272H、S275R、C277L、T285R、S286F、C287I、D288W、E289A、E289L、E289R、E289S、A290N、A290H、A290V、I292L、I292V、K299R、R300D、W301F、L394A、L394H、L394R、L394P、D395C、D395V、D395P、D395S、R396G、R396T、R396Y、D397F、D397H、D397N、D397R、D397S、D397T、R398A、R398D、R398G、R398N、R398I、R398S、K399D、K399P、K399L、K399I、K399V、K399E、K399R、K399S、K399T、K399P、E400A、E400C、E400F、E400P、E400R、E400S、E400T、E400V、D401C、D401F、D401H、D401G、D401M、D401K、D401P、D401R、D401S、D401T、D401V、E402A、E402F、E402G、E402L、E402I、E402R、E402S、E402T、E402Y、D403G、D403P、D403R、D403S、D403T、S404K、S404Q、S404R、S404P、S404Y、C405A、C405K、C405L、C405R、C405P、R406L、R406H、R406G、R406T、R406P、F407M、F407S、F407L、F407P、F407R、E408L、E408I、E408R、E408V、K502R、K557N、L657F、C672R、P758L、L789S、K827E、V860L。
[0296] It should be understood that the amino acid numbers in the mutant proteins of the present invention are based on the wild-type Cas protein, and when the sequence homology between a particular mutant protein and the wild-type Cas protein reaches 80% or more, the amino acid numbers of the mutant protein may be shifted relative to the amino acid numbers of the wild-type Cas protein, for example, by 1 to 100 positions to the N-terminus or C-terminus. It is common knowledge to those skilled in the art that such shifts are within a reasonable range using the sequence alignment techniques common in the art. Therefore, mutant proteins having 80% or more homology (e.g., 90%, 95%, 98%) and having the same or similar gene editing activity (e.g., cleavage activity) as or greater than that of the wild-type Cas protein (SEQ ID NO: 1) should not be excluded from the scope of mutant proteins of the present invention due to the shift in amino acid numbers.
[0297] The mutant proteins of the present invention are synthetic or recombinant proteins, that is, they may be products of chemical synthesis or produced from a prokaryotic or eukaryotic host (e.g., bacteria, yeast, or plants) using recombinant technology. Depending on the host used in the recombinant production protocol, the mutant proteins of the present invention may or may not be glycosylated. The mutant proteins of the present invention may or may not contain an initiating methionine residue.
[0298] The present invention further comprises fragments, derivatives, and analogues of the mutant protein. As used herein, the terms “fragment,” “derivative,” and “analogy” refer to proteins that substantially retain the same biological function or activity as the mutant protein.
[0299] The mutant protein fragments, derivatives, or analogues of the present invention may include: (i) mutant proteins in which one or more conserved or non-conserved amino acid residues (preferably conserved amino acid residues) are substituted, and the substituted amino acid residues may or may not be encoded by genetic codons; (ii) mutant proteins having substituents on one or more amino acid residues; (iii) mutant proteins formed by fusing a mature mutant protein with another compound (e.g., a compound that extends the half-life of the mutant protein, e.g., polyethylene glycol); (iv) mutant proteins formed by fusing an additional amino acid sequence to the mutant protein sequence (e.g., a leader sequence, a secretion sequence, a sequence for purifying the mutant protein, a proprotein sequence, or a fusion protein with an IgG fragment). Under the teachings of this specification, these fragments, derivatives, and analogues are known to those skilled in the art.
[0300] In some embodiments, the following groups of other amino acids are considered to be conserved substitutions with each other: [Table 2] In some embodiments, the following groups of other amino acids are considered to be conserved substitutions with each other (see, for example, Creighton, *Proteins* (1984)): [Table 3] In some embodiments, the following groups of other amino acids are considered to be conserved substitutions with each other:
[0301] [Table 4] The active mutant protein of the present invention has gene editing activity (e.g., cleavage activity) equivalent to or greater than that of the wild-type Cas protein (SEQ ID NO: 1).
[0302] Furthermore, the mutant proteins of the present invention can also be modified. Modifications (usually without changing the primary structure) include chemically derived forms of the mutant protein in vivo or in vitro, such as acetylation or carboxylation. Modifications also include glycosylation, such as mutant proteins resulting from glycosylation modification during the synthesis and processing of the mutant protein or in a further processing step. Such modifications can be carried out by exposing the mutant protein to a glycosylation enzyme (e.g., mammalian glycosylase or deglycosylase). Modifications also include sequences having phosphorylated amino acid residues (e.g., phosphotyrosine, phosphoserine, phosphothreonine). Furthermore, mutant proteins modified to improve their resistance to protein hydrolysis or to optimize their solubility are also included.
[0303] The term "polynucleotide encoding a mutant protein" may refer to a polynucleotide comprising a sequence encoding the mutant protein of the present invention, and may further comprise a polynucleotide comprising additional coding sequences and / or non-coding sequences.
[0304] The present invention also relates to variants of the polynucleotides encoding polypeptide or mutant protein fragments, analogs, and derivatives having the same amino acid sequence as the present invention. These nucleotide variants include substitutional variants, deletion variants, and insertion variants. As is well known to those skilled in the art, allelic variants are one of the alternative forms of polynucleotides in which one or more nucleotides may be substituted, deleted, or inserted, but the function of the mutant protein they encode is not substantially altered.
[0305] The present invention also relates to a polynucleotide that hybridizes with the above sequence and has a homology between the two sequences of at least 50%, preferably at least 70%, and more preferably at least 80%. The present invention particularly relates to a polynucleotide that can hybridize with the polynucleotide described in the present invention under stringent conditions (or strict conditions). In the present invention, “stringent conditions” means (1) hybridization and elution at relatively low ionic strength and relatively high temperature, such as 0.2 × SSC, 0.1% SDS, 60°C; or (2) hybridization in the presence of a denaturing agent, such as 50% (v / v) formamide, 0.1% fetal bovine serum / 0.1% Ficoll, 42°C; or (3) conditions under which hybridization occurs only when the homology between the two sequences is at least 90%, more preferably 95%.
[0306] The mutant proteins and polynucleotides of the present invention are preferably provided in isolated form, and more preferably homogeneously purified.
[0307] The full-length sequences of the polynucleotides of the present invention can typically be obtained by PCR amplification, recombination, or artificial synthesis. PCR amplification involves designing primers based on the relevant nucleotide sequences disclosed in the present invention, particularly the open reading frame sequences, and amplifying them using a commercially available cDNA library or a cDNA library prepared according to a common method known to those skilled in the art as a template to obtain the relevant sequences. For long sequences, it is often necessary to perform PCR amplification two or more times, after which the fragments amplified each time are joined in the correct order.
[0308] Once the relevant sequence is obtained, a large quantity of it can be obtained by recombination. Typically, the sequence is cloned into a vector, introduced into cells, and then the relevant sequence is isolated from the proliferated host cells using conventional methods.
[0309] Furthermore, related sequences can be synthesized using artificial synthesis methods, which are particularly suitable when the fragment length is short. Typically, a very long fragment can be obtained by first synthesizing multiple small fragments and then concatenating them.
[0310] Currently, the DNA sequence encoding the protein (or fragment thereof or derivative thereof) of the present invention can be obtained entirely by chemical synthesis. This DNA sequence can then be introduced into various existing DNA molecules (or, for example, vectors) and cells known in the art. Furthermore, mutations can be introduced into the protein sequence of the present invention by chemical synthesis.
[0311] To obtain the polynucleotides of the present invention, a method of amplifying DNA / RNA using PCR technology is preferred. In particular, when it is difficult to obtain full-length cDNA from a library, the RACE method (RACE-cDNA terminal rapid amplification method) can be preferably used. The primers used for PCR can be appropriately selected according to the sequence information of the present invention disclosed herein and can be synthesized by conventional methods. The amplified DNA / RNA fragments can be isolated and purified by conventional methods such as gel electrophoresis.
[0312] Fusion protein In one embodiment, the present invention provides a fusion protein comprising a Cas protein described in any one of the above-mentioned items and one or more functional domains.
[0313] In one embodiment, the functional domain includes one or more of the following: a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcription activation domain, a transcription repression domain, a nuclease, a deaminase domain, a methylase, a demethylase, a transcription termination factor, an HDAC, a polypeptide having cleavage activity, and a ligase.
[0314] In one example, "methylase" can refer to, for example, HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3, ZMET2, CMT1, CMT2, and the like.
[0315] Demethylases are enzymes that remove methyl (CH3-) groups from nucleic acids, proteins (e.g., histones), and other molecules. Demethylases are important in epigenetic modification mechanisms. Demethylase proteins alter the transcriptional regulation of the genome by controlling the level of methylation that occurs in DNA and histones, and further regulate the chromatin state at specific gene loci in living organisms. Examples include TET1 (ten-eleven translocation 1), 10-11 translocation (TET) dioxygenase 1 (TET1CD), DME, DML1, DML2, and ROS1.
[0316] In some embodiments, the transcription termination factor may be, for example, eukaryotic transcription termination factor 1 (ERF1) or eukaryotic transcription termination factor 3 (ERF3).
[0317] In one embodiment, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.
[0318] In one embodiment, the localization signal includes a nuclear localization signal and / or a nuclear export signal.
[0319] Preferably, the nuclear export signal includes human protein tyrosine kinase 2.
[0320] Preferably, the reporter protein comprises one or more of glutathione-S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, or autofluorescent protein.
[0321] Preferably, the autofluorescent protein includes one or more of the following: green fluorescent protein, HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, or blue fluorescent protein.
[0322] Preferably, the DNA-binding domain includes one or more of the methylation-binding proteins, LexADBD, or Gal4DBD.
[0323] Preferably, the epitope tag includes one or more of the following: a histidine tag, a V5 tag, a FLAG tag, an influenza virus hemagglutinin tag, a Myc tag, a VSV-G tag, or a thioredoxin tag.
[0324] Preferably, the transcriptional activation domain includes VP64 and / or VPR.
[0325] Preferably, the transcriptional repression domain includes KRAB and / or SID.
[0326] Preferably, the nuclease contains FokI.
[0327] Preferably, the deaminase domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.
[0328] In one embodiment, the deaminase domain includes an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with the amino acid sequence described in any of SEQ ID NOs: 28-29, 54-55.
[0329] In one embodiment, the functional domain is the full length or functional fragment of TadA8e (SEQ ID NO: 55).
[0330] Preferably, the polypeptide having cleavage activity includes a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity.
[0331] Preferably, the ligase includes DNA ligase and / or RNA ligase.
[0332] Polynucleotides In one embodiment, the present invention provides a polynucleotide which is a polynucleotide sequence encoding the Cas protein or a polynucleotide sequence encoding the fusion protein.
[0333] In one embodiment, the polynucleotide is a DNA molecule that has been codon-optimized according to the codon bias of the host cell.
[0334] The optimization described herein may require mutating the nucleotide sequence encoding a protein (e.g., the Cas protein of this disclosure) to mimic the codon bias expected in a host organism or cell encoding the same protein. Thus, the codons may change, but the encoded protein remains unchanged. For example, if the target cell is expected to be a human cell, a nucleotide sequence encoding a protein optimized with human codons can be used. In another non-limiting embodiment, if the host cell is expected to be an animal cell (e.g., a mouse cell or an insect cell), a nucleotide sequence encoding a protein optimized with the codons of that animal can be generated. In another non-limiting embodiment, if the host cell is expected to be a plant cell, a nucleotide sequence encoding a protein optimized with plant codons can be generated.
[0335] A selection list of codons is readily available; for example, the "Codon Usage Database" can be obtained from www.kazusa.or.jp / codon. In some cases, the nucleic acids of this disclosure comprise a nucleotide sequence encoding CasY6 and its variants or fusion proteins, wherein the nucleotide sequence is codon-optimized for expression in eukaryotic cells. In some cases, the nucleic acids of this disclosure comprise a nucleotide sequence encoding CasY6 and its variants or fusion proteins, wherein the nucleotide sequence is codon-optimized for expression in animal cells. In some cases, the nucleic acids of this disclosure comprise a nucleotide sequence encoding CasY6 and its variants or fusion proteins, wherein the nucleotide sequence is codon-optimized for expression in fungal cells. In some cases, the nucleic acids of this disclosure comprise a nucleotide sequence encoding CasY6 and its variants or fusion proteins, wherein the nucleotide sequence is codon-optimized for expression in plant cells.
[0336] In one embodiment, the host cells include prokaryotic cells or eukaryotic cells.
[0337] CRISPR system The terms “Clustered Regular Interspaced Short Palindromic Repeat (CRISPR)-CRISPR-Related (Cas) (CRISPR-Cas) System” or “CRISPR System” can be used interchangeably and have meanings generally understood by those skilled in the art, and typically include transcripts or other elements associated with the expression of CRISPR-related ("Cas") genes, or transcripts or other elements capable of directing the activity of said Cas gene.
[0338] CRISPR-Cas composition In one embodiment, the present invention also provides a CRISPR-Cas composition comprising the following:
[0339] (1) Protein component: The Cas protein, the fusion protein, or a nucleic acid molecule encoding the Cas protein or the fusion protein; (2) RNA component: guide RNA, one or more nucleic acids encoding the guide RNA, precursor RNA of the guide RNA, or nucleic acid encoding the precursor RNA of the guide RNA.
[0340] The protein component and the nucleic acid component bind to each other to form a complex.
[0341] In one embodiment, the composition is an activated CRISPR complex further comprising a target sequence of a target nucleic acid bound to the guide RNA.
[0342] In one embodiment, the CRISPR-Cas composition comprises one or more vectors, and the one or more vectors are (1) A first regulatory element operably linked to the nucleotide sequence encoding the Cas protein or the nucleotide sequence encoding the fusion protein; and (2) A second regulatory element operably linked to the nucleotide sequence encoding the guide RNA. Includes, The aforementioned guide RNA is (a) A spacer sequence that can hybridize with the target sequence of the target nucleic acid, and (b) Direct Repeat (DR) sequences linked to the spacer sequence, which guide the Cas protein to bind to the guide RNA and form a CRISPR-Cas complex that targets the target sequence. Includes.
[0343] Here, the first adjustment element and the second adjustment element are located in the same or different vectors in the CRISPR-Cas vector system.
[0344] In one embodiment, the first regulatory element or the second regulatory element includes a promoter, the promoter includes one or more of an inductive promoter, a constitutive promoter, or a tissue-specific promoter.
[0345] In one embodiment, the promoter includes one or more of T7, SP6, T3, CMV, EF1a, SV40, PGK1, humanβ-actin, CAG, U6, H1, T7, T7lac, araBAD, trp, lac, or Ptac.
[0346] In one embodiment, the first adjustment element and the second adjustment element are located on the same or different vectors.
[0347] In one embodiment, the vector includes a retroviral vector, a lentiviral vector, an adenovirus vector, an adeno-associated virus vector, a herpes simplex virus vector, or a phagemid vector.
[0348] In one embodiment, the vector includes a plasmid vector.
[0349] In one embodiment, the target nucleic acid includes DNA derived from a eukaryote or DNA derived from a prokaryote.
[0350] In one embodiment, the eukaryote includes animals or plants.
[0351] In one embodiment, the target nucleic acid includes non-human mammalian DNA, human DNA, insect DNA, avian DNA, reptile DNA, amphibian DNA, rodent DNA, fish DNA, worm DNA, nematode DNA, or yeast DNA.
[0352] In one embodiment, the non-human mammalian DNA includes non-human primate DNA.
[0353] CRISPR / Cas complex The term "CRISPR / Cas complex" refers to a complex formed by the binding of a guide RNA, gRNA (guide RNA), or mature crRNA (or guide RNA) to a Cas protein, and includes a direct repeat sequence that hybridizes to the guide sequence of the target sequence and binds to the Cas protein. The complex can identify and cleave target nucleotides that can hybridize with the guide RNA or mature crRNA.
[0354] Guide RNA (gRNA) The terms “guide RNA (gRNA),” “mature crRNA,” “crRNA,” “guide sequence,” and “guide RNA” are interchangeable and have meanings generally understood by those skilled in the art. Generally, guide RNA may include a direct repeat (DR) sequence and a spacer sequence, or be substantially composed of a direct repeat (DR) sequence and a spacer sequence, or may be composed of a direct repeat (DR) sequence and a spacer sequence.
[0355] In some cases, the spacer sequence is any polynucleotide sequence that has sufficient complementarity with the target sequence and hybridizes with the target sequence to induce specific binding between the CRISPR-Cas complex and the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the spacer sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. The guide sequence includes a sequence (e.g., a direct repeat (DR) sequence) that has sufficient complementarity with the target nucleic acid sequence and hybridizes with the target nucleic acid sequence to induce sequence-specific binding between the complex and the target nucleic acid sequence.
[0356] It is known in this field that perfect complementarity is not required if there is sufficient complementarity for it to function. Therefore, cleavage efficiency can be adjusted by introducing mismatches (e.g., one or more mismatches between the spacer sequence and the target nucleic acid, e.g., one or two nucleotide mismatches, including the mismatch locations along the spacer / target sequence) where necessary. For example, if a target cleavage rate of less than 100% (e.g., within a cell population) is expected, one or two mismatches between the spacer sequence and the target sequence may be introduced into the spacer sequence.
[0357] In one embodiment, the present invention provides a guide RNA comprising a direct repeat (DR) sequence capable of binding to the Cas protein and a spacer sequence capable of targeting a target sequence.
[0358] In one embodiment, the direct repeat sequence (DR) includes the sequence shown in sequence number 5.
[0359] In one embodiment, the 3' end of the direct repeat sequence includes a stem-loop structure, where the first stem nucleotide strand and the second stem nucleotide strand hybridize with each other to form the stem of the stem-loop structure, and the loop nucleotide strand forms the loop of the stem-loop structure.
[0360] In one embodiment, the direct repeat sequence includes a nucleotide sequence having at least 80% identity with the nucleotide sequence described in Sequence ID No. 5.
[0361] In one embodiment, the direct repeat sequence includes a nucleotide sequence having at least 85%, more preferably 90%, and even more preferably 95% identity with the nucleotide sequence described in Sequence ID No. 5.
[0362] In one embodiment, the direct repeat sequence includes the nucleotide sequence described in SEQ ID NO: 5.
[0363] In one embodiment, more than 80% of the spacer sequence is complementary to the target nucleic acid.
[0364] In one embodiment, 90% or more, more preferably 95% or more, even more preferably 99% or more, and even more preferably 100% of the spacer sequence is complementary to the target nucleic acid.
[0365] In one embodiment, the length of the spacer sequence is longer than the length of 15 nucleotides, preferably 15 to 100 nt, more preferably 15 to 50 nt (for example, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 nt), more preferably 15 to 25 nt, more preferably 15 to 23 nt, and more preferably 18 to 22 nt.
[0366] In one embodiment, the length of the spacer array is 15nt, 18-23nt, or 25nt, preferably selected from 15nt or 18-23nt. For example, the length of the spacer array can be selected from 15, 18, 19, 20, 21, 22, 23, or 25nt.
[0367] target nucleic acid In the present invention, the target nucleic acid refers to a specific nucleic acid that is interchangeable with the target sequence, target nucleic acid sequence, or target nucleic acid molecule and includes a nucleic acid sequence that is whole or partially complementary to the spacer sequence in the guide RNA. The "target sequence" refers to a polynucleotide targeted by the spacer sequence in the guide RNA, for example, a sequence complementary to the spacer sequence, where hybridization between the target sequence and the spacer sequence promotes the formation of a CRISPR-Cas complex (including the Cas protein and the guide RNA). Complete complementarity is not required, as long as there is sufficient complementarity to induce hybridization and promote the formation of the CRISPR-Cas complex. In some examples, the target nucleic acid includes a non-coding region (e.g., a promoter or terminator). In some examples, the target nucleic acid is single-stranded or double-stranded.
[0368] The target sequence may include any polynucleotide, such as DNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located within the cell nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts).
[0369] The target nucleic acid may be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer-adjacent motif (PAM).
[0370] Donor template In the present invention, donor template nucleic acid or donor template refers to a nucleic acid molecule that is interchangeable and which, after the target nucleic acid has been modified by the Cas protein described herein, can be used by one or more cellular proteins to modify the structure of the target nucleic acid.
[0371] In some examples, the donor template nucleic acid is a double-stranded or single-stranded nucleic acid. In some examples, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some examples, the donor template nucleic acid is an exogenous nucleic acid molecule. In some examples, the donor template nucleic acid is an endogenous nucleic acid (e.g., a chromosome). In some examples, genetic recombination can be achieved using the donor template, and the recombination is homologous.
[0372] Cutting Cleavage refers to DNA cleavage in target nucleic acids by the Cas protein described herein. In some examples, cleavage is double-stranded DNA. In some examples, cleavage is single-stranded DNA.
[0373] In this invention, the cleavage of target nucleic acid and the modification of target nucleic acid may overlap in meaning. Modification of target nucleic acid includes not only the modification of a single nucleotide, but also the insertion or deletion of nucleic acid fragments.
[0374] Reporter nucleic acid A reporter nucleic acid refers to a molecule that can be cleaved by a protein in the activated CRISPR system described herein or otherwise inactivated. A reporter nucleic acid comprises a nucleic acid element that can be cleaved by a CRISPR protein (e.g., employing a single-stranded non-target nucleic acid molecule containing different reporter groups or labeling molecules at both ends). Cleavage of the nucleic acid element generates a detectable signal. Before cleavage, or while the reporter nucleic acid is in an "active" state, the reporter nucleic acid prevents the generation or detection of a positive detectable signal. It will be understood that in some exemplary embodiments, minimal background signal may be generated in the presence of an active reporter nucleic acid. The positive detectable signal may be any signal detectable by optical, fluorescent, chemiluminescent, electrochemical, or other detection methods known in the art. For example, in some embodiments, in the presence of a reporter nucleic acid, a first signal (i.e., a negative detectable signal) is detected, and then, after a target molecule is detected and cleaved or inactivated by an activated CRISPR protein, it is converted into a second signal (e.g., a positive detectable signal). The reporter nucleic acid may be a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.
[0375] The detection method described in the present invention can be used for the quantitative detection of a target nucleic acid. The quantitative detection index can be quantified based on the signal intensity of a reporter group, for example, based on the emission intensity of a fluorescent group, or based on the width of a color band.
[0376] Functional domain In this specification, the term "functional domain" is used in its broadest sense and includes proteins such as enzymes or factors themselves, or fragments / domains thereof having a specific function. A Cas protein (e.g., a dCas protein) is linked / bound to one or more functional domains, which are one or more selected from localization signals, reporter proteins, Cas protein targeting regions, DNA-binding domains, epitope tags, transcriptional activation domains, transcriptional repression domains, nucleases, deaminase domains, methylases, demethylases, transcription termination factors, HDACs, polypeptides with cleavage activity, and ligases. If a Cas protein contains two or more functional domains, these functional domains may be identical or different.
[0377] Deaminase Domain In the present invention, the deaminase domain comprises a deaminase (e.g., adenosine deaminase or cytidine deaminase) catalytic domain, and as used herein, “adenosine deaminase” or “adenosine deaminase protein” means a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze a hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) to hypoxanthine (or the hypoxanthine portion of a molecule).
[0378] In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0379] Adenosine deaminases include, but are not limited to, family members containing other adenosine deaminase domains (ADAD), including an enzyme family member called an adenosine deaminase that acts on RNA (ADAR), an enzyme family member called an adenosine deaminase that acts on tRNA (ADAT), and other adenosine deaminase domains (ADAD). According to this disclosure, adenosine deaminases can target adenine in RNA / DNA and RNA double strands. In certain embodiments, adenosine deaminases are modified to enhance their ability to edit DNA in RNA / DNA heteroduplexes of RNA double strands.
[0380] In some embodiments, the deaminase is cytidine deaminase. The terms “cytidine deaminase” or “cytidine deaminase protein” refer to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze a hydrolytic deamination reaction that converts cytosine (or the cytosine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0381] Cytidine deaminases include, but are not limited to, members of the enzyme family called apolipoprotein BmRNA editing complex (APOBEC) family deaminases, activated-induced deaminases (AIDs), or cytidine deaminase 1 (CDA1). In certain embodiments, APOBEC family deaminases are included.
[0382] In some embodiments, the cytidine deaminase comprises the wild-type amino acid sequence of cytosine deaminase. In some embodiments, the cytidine deaminase comprises one or more mutations in the cytosine deaminase sequence to modify the editing efficiency and / or substrate editing bias of cytosine deaminase according to specific needs.
[0383] In one embodiment, the deaminase domain includes an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with the amino acid sequence described in any of SEQ ID NOs: 28, 29, 54-55.
[0384] In one embodiment, the functional domain is the full length or functional fragment of TadA8e (SEQ ID NO: 55).
[0385] identity "Identity" refers to the degree of sequence matching between two polypeptides or two nucleic acids. "Identity" is expressed as the percentage of identical residues in the total number of residues between the polypeptide or nucleic acid sequences, and the calculation of the total number of residues is determined based on the type of mutation. Types of mutations include insertion (elongation) at any one or both ends of a sequence, deletion (truncation) from any one or both ends of a sequence, substitution / replacement of one or more amino acids / nucleotides, insertion within a sequence, and deletion from within a sequence.
[0386] Taking polypeptide sequences as an example, if the mutation type is one or more of the following: amino acid / nucleotide substitution / replacement, insertion into a sequence, and deletion from a sequence, the total number of residues is calculated as the larger number of residues in the molecules being compared. If the mutation type further includes insertion (elongation) at any one or both ends of the sequence, or deletion (truncation) from any one or both ends of the sequence, the number of amino acids inserted or deleted at any one or both ends (e.g., less than 20 insertions or deletions at both ends) is not included in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned to produce the greatest possible matching between them, and any gaps in the alignment (if any) are resolved using a specific algorithm. Nucleotide identity is calculated based on the same principle.
[0387] vector A vector is a nucleic acid molecule that can transport other nucleic acid molecules linked to it.
[0388] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules with or without free ends (e.g., circular); nucleic acid molecules containing DNA, RNA, or both; and various other polynucleotides known in the art. Vectors can be introduced into host cells by transformation, transduction, or transfection to cause the genetic material elements they support to be expressed in the host cells. Vectors can be introduced into host cells to produce transcripts, proteins, or peptides (including proteins and their variants, fusion proteins, isolated nucleic acid molecules, etc. (e.g., CRISPR transcripts, etc., nucleic acid transcripts, proteins, or enzymes)). Vectors may contain various expression regulatory elements (including, but are not limited to, promoter sequences, transcription start sequences, enhancer sequences, selection elements, and reporter genes). Vectors may also contain origins of replication.
[0389] Vectors include plasmids and viral vectors, where a plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. In the case of viral vectors, the vector used for viral packaging contains viral DNA or RNA sequences. Viruses include, for example, retroviruses, replication-deficient retroviruses, adenoviruses, replication-deficient adenoviruses, and adeno-associated viruses. Viral vectors also contain polynucleotides carried by the virus for transfection into host cells. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episomal mammalian vectors) can autonomously replicate in the host cells into which they are introduced.
[0390] Other vectors (e.g., non-episomal mammalian vectors) are replicated alongside the host cell's genome after being introduced into the cell. Some vectors can also direct the expression of genes to which they are operably linked. Such vectors are called "expression vectors."
[0391] In some embodiments, the vector (e.g., viral vector or non-viral vector, e.g., lentiviral vector or plasmid) may be delivered to the target tissue by means of, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. The delivery may be a single dose or multiple doses. Those skilled in the art will understand that the actual dose delivered herein can vary considerably depending on a variety of factors, including, but not limited to, the choice of vector, target cells, organism, tissue, general condition of the subject being treated, the degree of transformation / modification required, route of administration, mode of administration, and type of transformation / modification required.
[0392] Adjustment element In this specification, “regulatory elements” include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression regulatory elements (e.g., polyadenylation signals, transcription termination signals such as poly-U sequences). For a detailed explanation, see Goeddel, GENE EXPRESSIONTECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include sequences that direct the constitutive expression of a nucleotide sequence in many types of host cells, as well as sequences that direct the expression of said nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in desired target tissues, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific types of cells (e.g., lymphocytes). In other cases, the regulatory elements do not have to be tissue-specific or cell-type-specific, and their expression can be directed in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner).
[0393] The term "promoter" refers to a non-coding nucleotide sequence located upstream of a gene that can initiate the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide that codes for or defines a gene product, causes the cell to produce that gene product under almost all physiological conditions in a cell. An inductive promoter is a promoter that selectively expresses a coding sequence or functional RNA in the presence of endogenous or exogenous stimuli, for example, in response to a chemical compound (chemical inducer), or in response to the environment, hormones, chemicals, and / or developmental signals. Examples of inductive or regulatory promoters include promoters that are induced or regulated by light, heat, stress, submersion or drought, salt stress, osmotic stress, plant hormones, injury, or chemicals (e.g., ethanol, abscisic acid (ABA), jasmonic acid ester, salicylic acid, or safener).
[0394] host cell In this specification, “host cell” refers to a cell derived from a eukaryotic cell (e.g., animal cell, plant cell, fungal cell, etc.), a prokaryotic cell (e.g., certain microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a multicellular organism cultured as a single cell entity (e.g., a cell line). The cell includes offspring of the original cell that have been genetically modified by nucleic acid and used as a nucleic acid receptor (e.g., an expression vector).
[0395] It should be understood that offspring of a single cell may not necessarily have the exact same morphology or genome as the original parent cell due to natural, accidental, or intentional mutations. A "recombinant host cell" (also called a "genetically modified host cell") is a host cell into which a different nucleic acid, such as an expression vector, has been introduced.
[0396] Those skilled in the art will understand that the design of an expression vector may depend on factors such as the selection of the host cells to be transformed and the desired expression level.
[0397] In another embodiment, the present invention also provides a host cell or its offspring comprising the Cas protein, the fusion protein, the polynucleotide, the vector system, the CRISPR-Cas system, or the composition.
[0398] In one embodiment, the host cells include non-human mammals, humans, insects, birds, reptiles, amphibians, rodents, fish, worms, nematodes, or yeast cells.
[0399] In one embodiment, the present invention also provides a multicellular organism comprising the aforementioned cells or their offspring.
[0400] In one embodiment, the multicellular organism is an animal or plant model used for the related disease.
[0401] NLS NLS refers to a "nuclear localization sequence" or "nuclear localization signal," which is an amino acid sequence that facilitates the entry of a protein into the cell nucleus. Nuclear localization sequences are known in the art (for example, described in international PCT application PCT / EP2000 / 011690 filed by Plank et al. on November 23, 2000, and disclosed on May 31, 2001, as WO / 2001 / 038547), and this patent is incorporated herein by reference to the published material relating to its exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, such as described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS includes the following amino acid sequence:
[0402] KRTADGSEFESPKKKRKV (SEQ ID NO: 35), AVKRPAATKKAGQAKKKKLD (SEQ ID NO: 36), KRPAATKKAGQAKKKK (SEQ ID NO: 37), KKTELQTTNAENKTKKL (SEQ ID NO: 38), KRGINDRNFWRGENGRKTR (SEQ ID NO: 39), RKSGKIAAIVVKRPRK (SEQ ID NO: 40), PKKKRKV (SEQ ID NO: 41), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 42).
[0403] Operable connection "Operationally linked" means that the target nucleotide sequence is linked to a regulatory element in such a way that it enables the expression of the nucleotide sequence (for example, in an in vitro transcription / translation system, or in the host cell if the vector is introduced into the host cell). Suitable vectors include lentiviruses and adeno-associated viruses, and the type of these vectors can be selected to target specific types of cells.
[0404] Complementarity "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence, either through the traditional Watson-Crick method or other non-traditional methods. The complementarity percentage indicates the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairings) with the nucleic acid sequence of the other (for example, if 5, 6, 7, 8, 9, and 10 out of 10 residues are complementary, the complementarity percentages are 50%, 60%, 70%, 80%, 90%, and 100%). "Perfectly complementary" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in the other nucleic acid sequence. "Substantially complementary" means having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% complementarity in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or two nucleic acids that hybridize under stringent conditions.
[0405] The term "stringent conditions" in relation to hybridization refers to conditions under which nucleic acids complementary to the target sequence primarily hybridize with the target sequence, and substantially do not hybridize with non-target sequences. Stringent conditions are usually sequence-dependent and determined by many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes with its target sequence.
[0406] Hybridization refers to a reaction in which one or more polynucleotides react to form a complex stabilized by hydrogen bonds between the bases of these nucleotide residues. This complex may include two strands forming a double helix, three or more strands forming a multi-strand complex, a single strand undergoing self-hybridization, or any combination thereof. A hybridization reaction may constitute a single step in a broader process (e.g., initiating PCR or enzymatic cleavage of polynucleotides). A sequence capable of hybridizing with a given sequence is called its "complement."
[0407] Hybridization of a target sequence and gRNA means that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and gRNA can hybridize to form a complex, or that at least 12, 15, 16, 17, 18, 19, 20 or more bases of the nucleic acid sequences of the target sequence and gRNA can complementary pair and hybridize to form a complex.
[0408] Expression Nucleic acid expression includes one or more of the following: generation of an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5' cap formation, and / or 3' end processing), translation of RNA into a polypeptide or protein, or post-translational modification of a polypeptide or protein.
[0409] delivery "Delivery" refers to providing an entity (e.g., a pharmaceutical product) to a destination. For example, the components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as DNA / RNA, RNA / RNA, or protein-RNA combinations. For example, the Cas protein can be delivered as a polynucleotide encoding DNA or RNA, or as a protein.
[0410] In one embodiment, the present invention also provides a delivery system comprising the Cas protein, the fusion protein, the polynucleotide, or the CRISPR-Cas composition.
[0411] In one embodiment, the delivery system further includes a delivery medium comprising nanoparticles, liposomes, exosomes, microbubbles, gene guns, or electroporation devices.
[0412] Furthermore, when the target of delivery is plant cells, delivery methods such as the use of cell-permeable peptides (CPPs) are employed. For example, in one specific embodiment, a Cas protein and / or at least one guide RNA is coupled with one or more CPPs, thereby effectively transporting the coupled CPPs into plant cells (e.g., within protoplasts). CPPs are short peptides of less than 35 amino acids derived from proteins or chimeric sequences, and can transport biomolecules across cell membranes in a receptor-independent manner. CPPs can be cationic peptides, peptides with hydrophobic sequences, amphiphilic peptides, peptides with proline-rich and antimicrobial sequences, and chimeric or bipartite peptides. CPPs can permeate biological membranes, trigger transmembrane movement of various biomolecules into the cytoplasm, improve their intracellular pathways, and thereby facilitate the interaction between biomolecules and their targets.
[0413] For example, CPPs include Tat (a nuclear transcription-activating protein necessary for HIV type 1 virus replication), penetratin, Kaposi's fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, and sweet arrow peptide.
[0414] Linker In this specification, “linker” refers to a chemical group or molecule that links two molecules or parts, such as two domains of a fusion protein, such as a Cas protein and a deaminase. In some linking schemes, the linker is located between or to the side of the two groups, molecules, or other parts and links them together via a covalent bond.
[0415] In some embodiments, the linker is a linear polypeptide consisting of amino acids or multiple amino acid residues linked by peptide bonds. In some embodiments, the linker is an organic molecule, group, polymer, or chemical part. The length and type of the linker can be designed as needed. In some embodiments, the linker can be selected from an artificially synthesized amino acid sequence or a naturally occurring polypeptide sequence.
[0416] detection In one embodiment, the present invention also provides a method for targeting and editing a target nucleic acid, comprising contacting the target nucleic acid with any of the above-described CRISPR-Cas systems or compositions.
[0417] In one embodiment, the present invention also provides a method for nonspecifically degrading single-stranded DNA after recognizing a target nucleic acid, the method comprising contacting the target nucleic acid with the CRISPR-Cas composition.
[0418] In one embodiment, the present invention also provides a method for targeting and nicking a non-spacer complementary strand of a double-stranded target nucleic acid after recognizing the spacer complementary strand of the double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with the CRISPR-Cas system or composition.
[0419] In one embodiment, the present invention also provides a method for targeting and cleaving a double-stranded target nucleic acid, comprising contacting the double-stranded target nucleic acid with the CRISPR-Cas system or composition.
[0420] In one embodiment, the non-spacer sequence complementary strand of the double-stranded target nucleic acid is nicked before the spacer complementary strand of the double-stranded DNA is nicked.
[0421] In one embodiment, the present invention also provides a method for specifically editing double-stranded nucleic acids, comprising contacting (1) and (2) below under sufficient conditions for a sufficient period of time.
[0422] (1) the Cas protein or fusion protein, another enzyme having sequence-specific nicking activity, and the guide RNA; and (2) the double-stranded nucleic acid.
[0423] The guide RNA instructs the Cas protein variant or the fusion protein to nick the opposing strand with another enzyme having sequence-specific nicking activity. The method causes a double-strand break.
[0424] In one embodiment, the present invention also provides a method for editing double-stranded nucleic acids, comprising contacting (1) and (2) below under sufficient conditions for a sufficient period of time.
[0425] (1) the Cas protein or fusion protein, and the fusion protein of a protein domain having DNA modification activity, and the guide RNA that targets the double-stranded nucleic acid; and (2) the double-stranded nucleic acid.
[0426] The Cas protein of the fusion protein is modified to nicking the non-target strand of the double-stranded nucleic acid.
[0427] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at different sites, resulting in alternating cleavage.
[0428] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at the same site, resulting in a blunt-end double-strand break.
[0429] In one embodiment, the present invention also provides a method for targeting and cleaving a single-stranded target nucleic acid, comprising contacting the target nucleic acid with a CRISPR-Cas composition according to any of the above claims.
[0430] In one embodiment, the present invention also provides a method for inducing a change in cellular state, comprising contacting the CRISPR-Cas composition with the target nucleic acid in a cell.
[0431] In one embodiment, the cellular state includes apoptosis or dormancy.
[0432] In one embodiment, the cells include eukaryotic cells or prokaryotic cells.
[0433] In one embodiment, the cells include mammalian cells or plant disease cells.
[0434] In one embodiment, the cells include cancer cells.
[0435] In one embodiment, the cells include infectious cells or cells infected with a pathogen.
[0436] In one embodiment, the cells include virus-infected cells and prion-infected cells.
[0437] In one embodiment, the cells include fungal cells, protozoan cells, or parasitic cells.
[0438] In one embodiment, the present invention also provides a method for detecting a target nucleic acid in a sample, the method comprising contacting the sample with the Cas protein, guide RNA and a non-target sequence, and detecting the target nucleic acid by detecting a detectable signal produced when the non-target sequence is cleaved by the Cas protein, wherein the non-target sequence does not hybridize with the guide RNA.
[0439] kit In one embodiment, the present invention provides a kit comprising the use of the Cas protein, the fusion protein, the polynucleotide, the CRISPR-Cas composition, and the host cell in the production of the kit, wherein the components of the kit are arranged in the same or different containers.
[0440] In one embodiment, the present invention also provides a container comprising the kit.
[0441] In one embodiment, the container includes a sterile container.
[0442] In one embodiment, the container includes a syringe.
[0443] In some embodiments, the kit further includes instructions for using the kit, for example, instructions in multiple languages. The kit may further include one or more reagents used in the process of utilizing the one or more components. The reagents may be provided in any suitable container. The kit may provide, for example, one or more reaction or storage buffers. The reagents may be provided in a form that requires the addition of one or more other components before use (e.g., in a concentrated or lyophilized form), and the buffer may be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer and combinations thereof. The buffer may have an appropriate pH (pH value), for example, it may be alkaline. In some embodiments, the pH of the buffer is about 7 to 10.
[0444] treatment "Treatment" means treating or curing a disorder in a subject, delaying the onset of symptoms of a disorder, and / or delaying the severity of a disorder. The term "subject" includes, but is not limited to, various animals, plants, and microorganisms. Animals include mammals such as, for example, bovines, equids, sheep, pigs, canids, felines, rabbits, rodents (e.g., mice or rats), non-human primates (e.g., rhesus monkeys or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., human) has a disorder (e.g., a disorder due to a deficiency of a disease-related gene). "Plants" means any differentiated multicellular organism capable of photosynthesis, including crop plants at any mature or developing stage.
[0445] In one embodiment, the present invention also provides the use of the Cas protein, the fusion protein, the polynucleotide, the CRISPR-Cas composition, and the host cells in the manufacture of pharmaceuticals for treating disorders or diseases in subjects requiring them.
[0446] In one embodiment, the use includes administering the CRISPR-Cas composition to the subject or the subject's ex-vivo cells.
[0447] In one embodiment, the spacer sequence is complementary to at least 15 nucleotides of the target nucleic acid associated with the disorder or disease, and the Cas protein or the fusion protein cleaves the target nucleic acid.
[0448] In one embodiment, the disorder or disease includes cancer or an infectious disease.
[0449] In one embodiment, the cancer includes one or more of Wilms' tumor, Ewing's sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, multiple myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, or bladder cancer.
[0450] In one embodiment, the disorder or disease includes one or more of the following: cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell anemia, hypercholesterolemia, transthyretin amyloidosis, or β-thalassemia.
[0451] In one embodiment, the pathogen of the infectious disease includes one or more of the following: human immunodeficiency virus, herpes simplex virus-1, or herpes simplex virus-2.
[0452] The present invention has the following main advantages.
[0453] (a) The present invention has discovered a novel Cas protein for the first time. The Cas protein of the present invention has excellent gene editing activity and can effectively edit or cleave target genes and can effectively treat disorders or diseases in subjects in need (e.g., one or more of the following: cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell anemia, hypercholesterolemia, transthyretin amyloidosis, or β-thalassemia).
[0454] (b) The Cas protein of the present invention is a completely new Cas enzyme that exhibits good nuclease activity both in vivo and in vitro, and has a wide range of potential applications.
[0455] (c) Compared to Cas enzymes disclosed in the prior art, the Cas protein of the present invention has superior editing efficiency and provides more options for base editing tools.
[0456] (d) The base editor constructed with the Cas enzyme disclosed in this invention can effectively perform base editing and has potential application prospects.
[0457] (e) In this invention, a novel Cas protein mutant is unexpectedly obtained by mutating the wild-type Cas protein, and the Cas protein mutant has better gene editing activity than the wild-type Cas protein.
[0458] The present invention will be further described below with reference to specific examples. It should be understood that these examples are for illustrative purposes only and do not limit the scope of the present invention. Experimental methods in the following examples where specific conditions are not explicitly stated are usually carried out under general conditions, for example, those described in Sambrook et al., Molecular Clones: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or under conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts refer to weight percentages and weight parts.
[0459] Unless otherwise specified, the reagents and materials used in the examples of this invention are all commercially available products.
[0460] Example 1: Acquisition of Cas protein The inventors analyzed the metagenomics of uncultured tissue and identified a novel Cas protein through analyses such as redundancy removal and protein clustering. Blast analysis revealed that the Cas protein showed low sequence identity with reported Cas proteins, and it was named CasY6 in this invention.
[0461] The amino acid sequence of the CasY6 protein is shown in SEQ ID NO: 1, and its human codon-optimized nucleotide sequence is shown in SEQ ID NO: 2.
[0462] Analysis of the direct repeat (DR) sequence of the guide RNA corresponding to the CasY6 protein revealed that the DNA sequence encoding the direct repeat (DR) sequence of the guide RNA corresponding to the CasY6 protein is as follows:
[0463] TATCCATCGTGCCGCCTCTTGGCAC(Sequence ID: 5) The inventors further analyzed the RNA secondary structure of the DR sequence in pre-crRNA using RNAfold. The results of the analysis are shown in Figure 2. The analysis revealed that the PAM corresponding to CasY6 is TTN, and N is A / T / C / G. The gRNA (also called crRNA) sequence of the CasY6 protein consists of a spacer sequence and a direct repeat (DR) sequence.
[0464] Based on the identification results, the CasY6 of the present invention belongs to the Cas12 protein family.
[0465] Example 2: Verification of Cas protein cleavage activity 1. Plasmid construction (1) Targeting the TTR gene, a spacer sequence: GCATCTCCCCATTCCATGAG (Sequence ID: 6) was designed based on the target sequence of the TTR gene.
[0466] Based on the DR sequences of the CasY6 protein (amino acid sequence shown in SEQ ID NO: 1) and the LbCpf1 protein (amino acid sequence shown in SEQ ID NO: 12), sgRNA sequences targeting TTR target genes were designed; see the table below for details. [Table 5] Depending on the requirements for vector expression, the T7 promoter and rrnB T2 terminator were added to the 5' and 3' ends, respectively, of the sgRNA sequences of CasY6 and LbCpf1 to obtain CasY6-sgRNA expression fragment sequences.
[0467] [ka] Here, the underlined sequence is the CasY6 DR sequence, the double underlined sequence is the spacer sequence, the italicized sequence is the T7 promoter, the wavy underlined sequence is the rrnB T2 terminator sequence, the space between the spacer sequence and the rrnB T2 terminator sequence is the linker sequence, the dashed sequence is the MfeI enzyme cleavage site, the bolded sequence is the MluI enzyme cleavage site, and CACCG is the linker.
[0468] Similarly, LbCpf1-sgRNA expression fragment sequences were synthesized.
[0469] [ka] Here, the underlined sequence is the LbCpf1 DR sequence, the double underlined sequence is the spacer sequence, the italicized sequence is the T7 promoter, the wavy underlined sequence is the rrnB T2 terminator sequence, the space between the spacer sequence and the rrnB T2 terminator sequence is the linker sequence, the dashed sequence is the MfeI enzyme cleavage site, the bolded sequence is the MluI enzyme cleavage site, and CACCG is the linker.
[0470] To maintain sequence integrity, when synthesizing the CasY6-sgRNA expression fragment sequences and LbCpfl-sgRNA expression fragment sequences, AGC was introduced at the 5' end and ATA at the 3' end as protective bases.
[0471] (2) Suzhou Hongxun Biotechnology Co., Ltd. synthesized a nucleotide sequence fragment (SEQ ID NO: 2) encoding the CasY6 protein, and constructed the synthesized nucleotide sequence fragment encoding the CasY6 protein at positions 466 to 5160 of the ABE8e plasmid (Addgene, Plasmid#138489) to obtain a CasY6 recombinant expression plasmid (see Figure 1A for plasmid map).
[0472] Suzhou Hongxun Biotechnology Co., Ltd. synthesized an optimized nucleotide sequence encoding LbCpf1 (SEQ ID NO: 13) and constructed a recombinant LbCpf1 expression plasmid using the same method (see Figure 1B for plasmid map).
[0473] (3) Suzhou Hongxun Biotechnology Co., Ltd. synthesized the sgRNA expression sequence fragments (CasY6-sgRNA expression sequence and LbCpf1-sgRNA expression sequence) described in step (1), treated the sgRNA expression sequence with double enzyme digestion (MfeI / MluI), and inserted it into a CasY6 recombinant expression plasmid vector that had also been treated with double enzyme digestion (MfeI / MluI) to obtain a CasY6+sgRNA expression plasmid, which expresses CasY6 and sgRNA. The LbCpf1+sgRNA expression plasmid was constructed and obtained by the same method.
[0474] (4) Construct a Target plasmid containing a targeting sequence, and the construction process is as follows:
[0475] Suzhou Hongxun Biotechnology Co., Ltd. synthesized the araC-pBAD-CCDB fragment (SEQ ID NO: 11) containing the TTR gene target sequence (SEQ ID NO: 6). This araC-pBAD-CCDB fragment was inserted into the pKESK22 (Addgene, Plasmid #64857) plasmid at positions 1284-1300 to obtain the Target plasmid. Refer to SEQ ID NO: 4 for the sequence of the Target plasmid, and Figure 3 for the plasmid map.
[0476] 2. Preparation and transformation of competent E. coli cells The Target plasmid was introduced into DH5a competent cells, lines were drawn using inoculation loops, and isolated and inoculated into LB solid medium containing 50 μg / ml kanamycin sulfate. The cells were incubated overnight in a biochemical incubator at 37°C. The following day, single colonies were picked from the plates and inoculated into test tubes containing 4 ml of LB liquid medium containing 50 μg / ml kanamycin sulfate (Biological Engineering, A100408-0100). These were incubated overnight with shaking at 37°C and 200 rpm. The next day, 4 ml of bacterial suspension was inoculated into a 2 L shaking flask containing 400 ml of LB liquid medium containing 50 μg / ml kanamycin sulfate. These were incubated with shaking for 2-3 hours at 37°C and 200 rpm.
[0477] When the OD600nm value of the bacterial suspension reached 0.3-0.5, the shaking bottle was removed and placed on ice for 10-15 minutes. Under sterile conditions, the bacterial suspension was placed in a pre-cooled 500 ml centrifuge flask and centrifuged for 8 minutes at 4°C and 3000 rpm. The supernatant was discarded, approximately 200 ml of pre-cooled CaCl2 solution was added, and the bacterial cells were uniformly pipetted to suspend them. The mixture was then left in an ice bath for 30 minutes. Subsequently, the bacterial suspension was centrifuged for 8 minutes at 4°C and 3000 rpm, the supernatant was discarded, approximately 8 ml of pre-cooled CaCl2 solution was added, and the bacterial cells were resuspended. The resuspended bacterial cells were dispensed in 110 μl portions into 1.5 ml EP tubes and stored in an ultra-low temperature refrigerator at -80°C for use.
[0478] 3. Measurement of in-vivo editing efficiency of E. coli The CasY6+ sgRNA expression plasmid and the LbCpf1+ sgRNA expression plasmid were introduced into the competent cells prepared in Step 2, respectively. The specific procedure is as follows:
[0479] (1) Remove the competent cells from -80°C and quickly place them in ice. After about 5 minutes, the cell mass will thaw. Add the CasY6+ sgRNA expression plasmid, then gently tap the bottom of the centrifuge tube by hand to mix evenly, and let stand on ice for 25 minutes. Heat shock in a 42°C water bath for 45 seconds, then quickly return to ice and let stand for 2 minutes. Add 900 μl of antibiotic-free sterile LB medium to the centrifuge tube, mix evenly, and then resuscitate at 37°C and 220 rpm for 60 minutes. 100 μl of bacterial suspension was spread onto LB agar plates (abbreviated as C-LB medium) containing 30 μg / ml of carbenicillin-resistant bacteria (Biological Engineering, A100358-0001), and onto LB agar plates (abbreviated as CL-LB medium) containing both 30 μg / ml of carbenicillin-resistant bacteria and 10 mM L-arabinose (Biological Engineering, A610071-0100). The two LB agar plates were inverted and placed in an incubator, where they were incubated overnight at 37°C.
[0480] (2) Detection of in vivo editing efficiency of E. coli As shown in Figure 3, the Target plasmid contains a PBAD promoter that can be induced by L-arabinose and a CCDB gene whose expression is regulated by the PBAD promoter. The CCDB gene can express a CCDB toxic protein, which acts as a DNA gyrase inhibitor, locking DNA gyrase and cleaved double-stranded DNA complexes, thereby rendering DNA gyrase inoperable and ultimately leading to cell death.
[0481] Based on this, the inventors designed a method for detecting the in-vivo editing efficiency of E. coli.
[0482] Under conditions where L-arabinose is present in the culture medium, if the CasY6 protein or LbCpf1 protein can specifically target and cleave the target sequence (SEQ ID NO: 6) of the TTR gene on the Target plasmid via sgRNA guidance, the PBAD promoter's regulatory pathway for CCDB toxicity protein expression is cleaved, and the host cell survives without producing CCDB toxicity protein. Conversely, if the CasY6 protein or LbCpf1 protein cannot specifically target the TTR target sequence on the Target plasmid, the L-arabinose-induced PBAD promoter regulates the CCDB gene, causing the host cell E. coli to die from the expression of CCDB toxicity protein.
[0483] Therefore, the editing efficiency of the CasY6 protein in which it targets and cleaves TTR target genes within E. coli can be calculated from the ratio of the number of bacterial clones on CL-LB medium to the number of bacterial clones on C-LB medium in step (2).
[0484] The results are shown in Figure 4. When the number of E. coli clones was counted and the ratio was calculated, the editing efficiency of the CasY6 protein was 16.2%, while the editing efficiency of the LbCpf1 protein was 6.4%, indicating that the editing efficiency of the CasY6 protein is clearly higher than that of the LbCpf1 protein.
[0485] Example 3: Detection of editing efficiency in HEK293T cells 1. Construction of a TTR-sgRNA expression plasmid (1) Based on the target sequence of the TTR gene (Sequence ID: 6), a TTR-sgRNA sequence was designed and oligonucleotides were synthesized.
[0486] CasY6-TTR-sgRNA sequence: TATCCATCGTGCCGCCTCTTGGCAC GCATCTCCCCATTCCATGAG (Sequence ID: 8). The underlined part of the sequence is the DR sequence, and the remaining part is the spacer sequence.
[0487] LbCpf1-TTR-sgRNA sequence: TAATTTCTACTAAGTGTAGAT GCATCTCCCCATTCCATGAG (Sequence ID: 15). The underlined part of the sequence is the DR sequence, and the remaining part is the spacer sequence.
[0488] (2) A CACC sequence is added to the 5' end of the upstream sequence of TTR-sgRNA, and an AAAA sequence is added to the 5' end of the downstream sequence. Oligos are also synthesized, and the specific sequence is as follows. [Table 6] After synthesizing the upstream and downstream primers for the TTR-sgRNA, annealing was performed using a pre-set program (95°C for 5 min; -2°C / s at 95°C to 85°C; -0.1°C / s at 85°C to 25°C; and holding at 4°C). Subsequently, the annealing product was ligated into a linearized PHK09T vector using BsmBI (NEB, #R0580L). The sequence of the PHK09T vector is shown in Sequence ID No. 3, and the plasmid map is shown in Figure 5.
[0489] The linearization of the PHK09T vector and the ligation method with the TTR-sgRNA annealing product are as follows:
[0490] First, the PHK09T vector was linearized. The linearization reaction system consisted of 3 μg of PHK09T vector, 6 μL of buffer (NEB:R0539L), and 2 μL of BsmBI, which were combined with ddH2O to a total volume of 60 μL and enzymatically digested overnight at 50°C.
[0491] The ligation reaction system for TTR-gRNA annealing products and linearization vectors was prepared by concentrating 1 μL of T4 ligase buffer (NEB,#M0202L), 20 ng of linearization vector, 5 μL of annealed oligo fragment, and 0.5 μL of T4 ligase (NEB,#M0202L) in ddH2O to 10 μL, and ligating overnight at 16°C to obtain CasY6-TTR-sgRNA expression plasmids and LbCpf1-TTR-sgRNA expression plasmids.
[0492] (3) The CasY6-TTR-sgRNA expression plasmid and LbCpf1-TTR-sgRNA expression plasmid obtained in step (2) are introduced into E. coli DH5a competent cells (Yuiji Organisms, DL1001). The specific procedure is as follows:
[0493] Remove DH5α competent cells from a -80°C refrigerator, quickly place them in ice, and after 5 minutes, when the cell mass has thawed, add the ligation product, gently tap the bottom of the centrifuge tube by hand to mix evenly, and let stand on ice for 25 minutes. Heat shock in a 42°C water bath for 45 seconds, quickly return to ice, and let stand for 2 minutes. Add 700 μl of sterile LB medium without antibiotics to the centrifuge tube, mix evenly, and resuscitate at 37°C and 200 rpm for 60 minutes. Centrifuge at 3000 rpm for 1 minute to collect the cells, leave about 100 μl of supernatant, gently pipette to resuspend the cells, and spread on LB medium containing Amp antibiotic. Invert the plate and incubate overnight in a 37°C incubator. Single colonies were picked and confirmed by sequencing. The positive clones were then cultured with shaking to extract the plasmid (endotoxin-free plasmid mega kit, TIANGEN: DP120-01), the concentration was measured, and the plasmids were stored in a refrigerator at -20°C for use.
[0494] 2. Detection of editing efficiency at the cellular level (1) Culture of HEK293T cells HEK293T cells (purchased from ATCC) were inoculated into DMEM medium (Gibco, 11965092) containing 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) with 10% FBS (v / v) added, and cultured in a cell incubator at 37°C with 5% CO2. Cells to be used for transfection were inoculated into 24-well cell culture plates the day before and cultured. The following day, the cells were observed, and transfection was performed when the cell density reached approximately 80%.
[0495] (2) HEK293T cells were co-transfected with the EGFP-C1 (Addgene, Plasmid, #54759) plasmid each with the CasY6 recombinant expression plasmid (see plasmid map in Figure 1A), the CasY6-TTR-sgRNA expression plasmid, the LbCpf1 recombinant expression plasmid (see plasmid map in Figure 1B), and the LbCpf1-TTR-sgRNA expression plasmid.
[0496] The amount of plasmid used per well for cell transfection in a 24-well plate was 0.3 μg of nuclease expression plasmid (CasY6 recombinant expression plasmid or LbCpf1 recombinant expression plasmid), 0.3 μg of sgRNA expression plasmid (CasY6-TTR-sgRNA expression plasmid or LbCpf1-TTR-sgRNA expression plasmid), and 0.3 μg of EGFP-C1 plasmid, respectively. The specific transfection procedure is as follows.
[0497] The CasY6 expression plasmid, CasY6-TTR-sgRNA expression plasmid, and EGFP-C1 plasmid were each mixed, then diluted in 25 μl of UltraFectin® transfection-specific serum-reduced medium (Genbai Seibutsu, L530KJ), and 2 μl of Lipofectamine 3000 (Invitrogen, L3000015) reagent was added. The mixture was then uniformly pipetted to form Reagent A, and allowed to stand for 5 minutes. Simultaneously, 2 μl of Lipofectamine 3000 transfection reagent (Invitrogen, L3000015) was diluted in 25 μl of UltraFectin® transfection-specific serum-reduced medium (Genbai Seibutsu, L530KJ), uniformly mixed to form Reagent B, and allowed to stand for 5 minutes.
[0498] Reagents A and B were mixed and uniformly pipetted using a pipette, then allowed to stand for 20 minutes. After standing, the mixed reagent was added drop by drop to the cells to be transfected in a 24-well plate, and the mixture was returned to a 37°C, 5% CO2 incubator for incubation. Six hours after transfection, the culture medium was replaced with DMEM medium containing 10% FBS.
[0499] Similarly, HEK293T cells were transfected with LbCpf1 recombinant expression plasmid, LbCpf1-TTR-sgRNA expression plasmid, and EGFP-C1 plasmid.
[0500] (3) Detection of editing efficiency 48 hours after transfection, the expression of EGFP fluorescent protein indicated successful cell transfection, and EGFP-positive cells were selected for detection of editing efficiency. Genome extraction was performed on these cells (Genome DNA Extraction Kit, TIANGEN, DP304-03). Identification primers were designed according to experimental needs, and the identification primer sequences used are shown in the table below. [Table 7] Using the genome as a template, PCR amplification is performed on sequences near the target site using the primers shown in the table above. The PCR amplification system is as follows:
[0501] 25 μL of 2×Taq Master Mix (Vazyme, P112-03), 1 μL of Primer-F (TTR-F) (10 pmol / μL), 1 μL of Primer-R (TTR-R) (10 pmol / μL), and 1 μL of template were combined with ddH2O to a total volume of 50 μL.
[0502] The amplified PCR products were used for high-throughput deep quenching (Jin Weizhi Biotechnology Co., Ltd.) or Sanger quenching (Hakushang Biotechnology (Shanghai) Co., Ltd.) to identify their editing efficiency.
[0503] The results of detecting and identifying the editing efficiencies of CasY6 and LbCpf1, as shown in Figure 6, indicate that in 293T cells, the editing efficiency of CasY6 was 25%, while the editing efficiency of LbCpf1 was only 18%, indicating that the editing efficiency of the CasY6 protein is much higher than that of LbCpf1.
[0504] Example 4: Use of CasY6 for base editing (1) Acquisition of catalyst-inactive CasY6 To obtain catalytically inactive (i.e., lost cleavage activity) dCasY6, the inventors constructed CasY6 mutants having single point mutations D659A, D711A, E895A, and D1069A, respectively: D659A-dCasY6, D711A-dCasY6, E895A-dCasY6, and D1069A-dCasY6. The specific construction method is as follows.
[0505] In Example 2, the CasY6+ sgRNA expression plasmid obtained in Step 1(3) was subjected to point mutations, and amino acid modifications were made to CasY6 at four sites: aspartic acid at position 659 (Asp,D), aspartic acid at position 711 (Asp,D), glutamic acid at position 895 (Glu,E), and aspartic acid at position 1069 (Asp,D). These amino acids were mutated to alanine (Ala,A), and the codons before and after the mutation of each amino acid are shown in the table below. [Table 8] For each amino acid and its codon listed in the table above, forward and reverse primers were designed and synthesized, and PCR amplification was performed using the CasY6+ sgRNA expression plasmid as a template. After amplification, the amplification product was recovered and purified using a general-purpose DNA purification and recovery kit (Tiangen Biochemical Technology (Beijing) Co., Ltd., DP214). The purified product was transformed into E. coli Dh5a competent cells (Weidi Biotechnology, DL1001) and cultured overnight at 37°C. The following day, single colonies were picked and sequenced. After confirming the sequencing, the positive clones were cultured with shaking to extract the plasmid (TIANGEN, DP120-01). The concentration was then measured and stored in a refrigerator at -20°C for use.
[0506] The obtained point mutation recombinant plasmids were named D659A-dCasY6+sgRNA expression plasmid, D711A-dCasY6+sgRNA expression plasmid, E895A-dCasY6+sgRNA expression plasmid, and D1069A-dCasY6+sgRNA expression plasmid, respectively.
[0507] Subsequently, the in-vivo editing efficiency of E. coli was detected for each of the constructed D659A-dCasY6+sgRNA expression plasmids: D711A-dCasY6+sgRNA expression plasmid, E895A-dCasY6+sgRNA expression plasmid, and D1069A-dCasY6+sgRNA expression plasmid. The detection method and calculation method were the same as in steps 2 and 3 of Example 2, and the experimental results are shown in Figure 7. When the number of E. coli clones was counted and the ratio was calculated, it was found that D659A-dCasY6, D711A-dCasY6, E895A-dCasY6, and D1069A-dCasY6 lost their catalytic activity (cleavage activity), which suggests that the CasY6 protein lost its cleavage activity due to point mutations in D659A, D711A, E895A, and D1069A.
[0508] 2. Detection of base editing efficiency at the cellular level (1) Construction of CasY6-TTR-sgRNA'' and CasY6-TTR-sgRNA'' plasmids Based on the target sequence of the TTR gene, sgRNAs were designed to obtain the CasY6-TTR-sgRNA' and CasY6-TTR-sgRNA'' sequences, and oligonucleotides were also synthesized.
[0509] CasY6-TTR-sgRNA': TATCCATCGTGCCGCCTCTTGGCAC tatatcccttctacaaattc(Array:20); CasY6-TTR-sgRNA'': TATCCATCGTGCCGCCTCTTGGCAC gtgtctatttccactttgta (array number: 21). Here, the underlined part of the sequence is the DR sequence, and the remaining part is the spacer sequence.
[0510] (2) A CACC sequence was added to the 5' end of the upstream sequence of each sgRNA, and an AAAA sequence was added to the 5' end of the downstream sequence. The specific morphology is as follows: [Table 9] Following the method of Step 1 in Example 3, the upstream and downstream sequences of CasY6-TTR-sgRNA' and CasY6-TTR-sgRNA'' were annealed, then ligated into a PHK09T vector to obtain CasY6-TTR-sgRNA' expression plasmids and CasY6-TTR-sgRNA'' expression plasmids. The plasmids were introduced into E. coli DH5a competent cells for amplification culture, accuracy was confirmed by sequencing, and the concentrations were measured before storage for use.
[0511] (3) Construction of a base editor plasmid (005V1-10-3 will be used as an example) The adenosine deaminase catalytic domain selected by the inventors was chosen from a mutant of the amino acid sequence shown in SEQ ID NO: 28: Q148G+Q149M+P150R (named 005V1-10-3). The amino acid sequence of this mutant is shown in SEQ ID NO: 29, and the nucleotide sequence encoding deaminase 005V1-10-3 is shown in SEQ ID NO: 30. A base editor fusion protein consisting of deaminase 005V10-3 and dCasY6 protein (amino acid sequence shown in SEQ ID NO: 43) was constructed by homologous recombination. The specific procedure is as follows.
[0512] First, Suzhou Hongxun Biotechnology Co., Ltd. synthesized the 005V1-10-3 nucleotide fragment containing homologous arm sequences and linkers.
[0513] [ka] Here, the bolded portion is the 005V1-10-3 nucleotide sequence, the italicized portion is the homologous arm region on the left and right, and the wavy portion is the linker sequence.
[0514] The D1069A-dCasY6+sgRNA expression plasmid obtained in step (1) was amplified by PCR and linearized to obtain a linearized expression vector. The primers used are shown in the table below. [Table 10] A nucleotide fragment of 005V1-10-3 containing homologous arm sequences and linkers (SEQ ID NO: 31) and a D1069A-dCasY6+sgRNA linearization expression vector were homologously recombined and reacted using Gibson Assembly Master Mix (NEB, E2611S). After the reaction was complete, the ligation product was transformed into E. coli DH5a competent cells (Yuiji Seibutsu, DL1001). The specific procedure is as follows.
[0515] Remove DH5α competent cells from a -80°C refrigerator, quickly place them in ice, and after 5 minutes, when the cell mass has thawed, add the ligation product, gently tap the bottom of the centrifuge tube by hand to mix evenly, and let stand on ice for 25 minutes. Heat shock in a 42°C water bath for 45 seconds, quickly return to ice and let stand for 2 minutes. Add 700 μl of sterile LB medium to the centrifuge tube, mix evenly, and resuscitate at 37°C and 200 rpm for 60 minutes. Centrifuge at 5000 rpm for 1 minute to collect the cells, leave about 100 μl of supernatant, gently pipette to resuspend the cells, and spread on LB medium containing Amp antibiotic. Invert the plate and incubate overnight in a 37°C incubator. Single colonies were picked and confirmed by sequencing. The positive clones were then cultured with shaking, and the base editor plasmid was extracted using an endotoxin-free plasmid mega kit (TIANGEN:DP120-01). The concentration was measured, and the plasmid was stored in a -20°C refrigerator for use. The nucleotide sequence encoding the base editor fusion protein 005V1-10-3-D1069A-dCasY6 is shown in SEQ ID NO: 34.
[0516] (4) Following the method of Step 2 of Example 3, the base editor plasmid and the EGFP-C1 (Addgene, Plasmid #54759) plasmid were co-transfected into 293T cells with CasY6-TTR-sgRNA' and CasY6-TTR-sgRNA'' expression plasmids, respectively.
[0517] Forty-eight hours after transfection, the expression of the EGFP fluorescent protein indicated successful cell transfection, and EGFP-positive cells were selected to detect editing efficiency. The genome of the aforementioned 293T cells was extracted using the kit (TIANGEN, DP304-03).
[0518] (5) Base editing efficiency was detected according to the method of step (3) of Example 3.
[0519] Primers were designed according to experimental needs, and the identification primer sequences used are shown in the table below. [Table 11] As a result, as shown in Figures 8A and 8B, the base editor consisting of D1069A-dCasY6 can achieve effective editing at multiple sites. As can be seen from Figure 8A, effective editing was observed at positions +2, +4, +13, +15, +16, and +17 of the target site mediated by CasY6-TTR-sgRNA, with editing efficiencies of approximately 10% to 30% at positions +13, +15, +16, and +17.
[0520] As can be seen in Figure 8B, effective editing was observed at positions +7, +13, and +20 of the target site mediated by CasY6-TTR-sgRNA'', with a base editing efficiency of over 10% at position +13 and over 15% at position +20.
[0521] Example 5: Obtaining CasY6 protein mutants For known Cas proteins (amino acid sequence number: SEQ ID NO: 1, nucleotide coding sequence: SEQ ID NO: 2, corresponding direct repeat (DR) sequence: SEQ ID NO: 5), site-directed mutagenesis and screening were used to obtain CasY6 mutant proteins with improved editing activity in order to improve the cleavage activity of the CasY6 protein.
[0522] Specifically, the methods for obtaining each CasY6 mutant are as follows:
[0523] Suzhou Hongxun Biotechnology Co., Ltd. synthesized a nucleotide sequence fragment (SEQ ID NO: 2) encoding the CasY6 protein. This synthesized CasY6 protein encoding nucleotide sequence fragment was then constructed at positions 466-5160 of the ABE8e plasmid (Addgene, Plasmid#138489) to obtain a CasY6 recombinant expression plasmid (see plasmid map in Figure 1A). Subsequently, PCR primers containing mutation sites were designed based on the CasY6 nucleotide encoding sequence, and expression plasmids for each mutant were obtained by homologous recombination.
[0524] The construction method for the CasY6-V1.1 recombinant expression plasmid is as follows:
[0525] 1. The designed primer is as follows: [Table 12] 2. The CasY6 recombinant expression plasmid was amplified using primers. The steps were as follows:
[0526] Amplification was performed using 2×Phanta Flash Master Mix (Dye Plus) high-fidelity enzyme, and the amplification system is as follows: [Table 13] 3. After purifying the amplification product, 100 ng was taken and added to Dh5a chemically competent cells, then heat-shocked and spread onto plates. The following day, single colonies were picked and sequenced, and those with correct sequences were designated as CasY6-V1.1 recombinant expression plasmids.
[0527] A total of 52 mutants were obtained using the method described above, and these mutants are shown in the table below. [Table 14-1] [Table 14-2] [Table 14-3] [Table 14-4] Example 6: Evaluation of the cleavage activity of CasY6 mutants 1. Construction of a TTR-sgRNA expression plasmid (1) Based on the target sequence of the TTR-3 locus of the TTR gene, tagaagggatatacaaagtg (SEQ ID NO: 48), a TTR-3-sgRNA sequence was designed and oligonucleotides were synthesized.
[0528] CasY6-TTR-3-sgRNA sequence: TATCCATCGTGCCGCCTCTTGGCAC tagaagggatatacaaagtg (Sequence ID: 49). The underlined part of the sequence is the DR sequence, and the remaining part is the spacer sequence.
[0529] (2) A CACC sequence is added to the 5' end of the upstream sequence of TTR-sgRNA, and an AAAA sequence is added to the 5' end of the downstream sequence to synthesize oligos. The specific sequence is as follows: [Table 15] After synthesizing the upstream and downstream primers for the TTR-sgRNA, annealing was performed according to a pre-set program (95°C for 5 min; -2°C / s at 95°C to 85°C; -0.1°C / s at 85°C to 25°C; and holding at 4°C). Subsequently, the annealing product was ligated into a linearized PHK09T vector using BsmBI (NEB, # R0739L). The sequence of the PHK09T vector is shown in Sequence ID: 3, and the plasmid map is shown in Figure 5.
[0530] The linearization of the PHK09T vector and the ligation method with the TTR-3-sgRNA annealing product are as follows:
[0531] First, the PHK09T vector was linearized. The linearized form is as follows:
[0532] 3 μg of PHK09T vector, 6 μL of buffer (NEB:R0739L), and 2 μL of BsmBI were combined with ddH2O to a volume of 60 μL and enzymatically digested overnight at 55°C.
[0533] The ligation reaction system between the TTR-gRNA annealing product and the linearization vector was prepared by concentrating 1 μL of T4 ligase buffer (NEB,#M0202L), 20 ng of linearization vector, 5 μL of annealed oligo fragment, and 0.5 μL of T4 ligase (NEB,#M0202L) in ddH2O to 10 μL, and ligating overnight at 16°C to obtain the CasY6-TTR-3-sgRNA expression plasmid.
[0534] (3) The CasY6-TTR-3-sgRNA expression plasmid obtained in step (2) is introduced into E. coli DH5a competent cells (Yuiji Seibutsu, DL1001). The specific procedure is as follows:
[0535] Remove DH5α competent cells from a -80°C refrigerator, quickly place them in ice, and after 5 minutes, when the cell mass has thawed, add the ligation product, gently tap the bottom of the centrifuge tube by hand to mix evenly, and let stand on ice for 25 minutes. Heat shock in a 42°C water bath for 45 seconds, quickly return to ice, and let stand for 2 minutes. Add 700 μl of sterile LB medium without antibiotics to the centrifuge tube, mix evenly, and resuscitate at 37°C and 200 rpm for 60 minutes. Centrifuge at 3000 rpm for 1 minute to collect the cells, leave about 100 μl of supernatant, gently pipette to resuspend the cells, and spread on LB medium containing Amp antibiotic. Invert the plate and incubate overnight in a 37°C incubator. Single colonies were picked and confirmed by sequencing. The positive clones were then cultured with shaking to extract the plasmid (endotoxin-free plasmid mega kit, TIANGEN: DP120-01), the concentration was measured, and the plasmids were stored in a refrigerator at -20°C for use.
[0536] 2. Detection of cleavage activity of each mutant using a dual luciferase reporter system To detect the cleavage activity of CasY6 mutants with high sensitivity, a dual luciferase reporter system was constructed (see plasmid map in Figure 9, plasmid sequence shown in SEQ ID NO: 52). Firefly luciferase (Fluc) was used as the internal standard gene, and a small luciferase (NanoLuc, nLuc) was used as the reporter gene. Luciferase can catalyze the oxidation of luciferin to oxyluciferin, and bioluminescence is released during the oxidation process of luciferin.
[0537] The successful catalysis and emission of light by Fluc in response to a substrate indicates successful introduction of the reporter vector. The LU and UC sequences in nLu×uc are the 119nt N-terminal sequence and the 469nt C-terminal sequence encoding nLuc, respectively, with a 60nt overlap between the two sequences. The insertion fragment (SEQ ID NO: 53) is located in the center of nLu×uc and contains the target sequence (SEQ ID NO: 48) targeted by the CRISPR / Cas system. A TAG early stop codon is present in the center of the target sequence, and upon cleavage, nLu×uc generates the precise nLuc coding frame via a recombination mechanism, expressing the nLuc gene, changing the cell from a NanoLuc non-expressing state to an expressing state, and further catalyzing a substrate to emit biofluorescence (Figure 10).
[0538] Using a 96-well plate, TTR-3-sgRNA expression plasmid (50 ng) and dual luciferase reporter system plasmid (50 ng) were co-transfected into HEK293 cells (purchased from ATCC) in each well, along with the respective CasY6 mutant recombinant plasmids (50 ng) obtained in Example 5, by PEI. After 24 hours of culture, the expression intensities of Fluc and nLuc were detected using a microplate reader with a dual luciferase reporter gene detection kit (Beyotime, RG028), and the cleavage activity of each CasY6 mutant was characterized by the nLuc fluorescence intensity.
[0539] The editing efficiencies of CasY6 wild-type and each mutant are as follows: [Table 16] As the analysis shows, each mutant exhibits some degree of improved cleavage activity compared to CasY6, with CasY6-V3.1, CasY6-V3.3, and CasY6-V3.12 having cleavage activity close to or exceeding 40%, and CasY6-V3.3 having the highest cleavage activity.
[0540] Example 7: Detection of the effective spacer array length of CasY6 Using the dual plasmid fluorescence reporter system of Example 6, spacer sequences of different lengths (15–25 nt, SEQ ID NO: 48, 58–64) targeting the TTR locus were designed to measure the effective spacer sequence length of CasY6, and the LUxUC intermediate insertion fragment (SEQ ID NO: 53) of the dual luciferase reporter plasmid of Example 6 was targeted. The LUxUC intermediate insertion fragment contains the same target sequence as the 25 nt spacer sequence (SEQ ID NO: 64). Following the construction method of Example 6, crRNA expression vectors containing either the spacer sequence of different lengths (SEQ ID NO: 48 or SEQ ID NO: 58–64) were constructed, and the other elements of the fluorescence reporter system were not changed.
[0541] Following the experimental procedure in Example 6, crRNA expression plasmids and CasY6-V3.3 mutant plasmids with spacer sequences of different lengths were prepared. Using a 96-well plate, HEK293 cells in each well were co-transfected with the CasY6-V3.3 mutant plasmid (50 ng) along with the crRNA expression plasmid (50 ng) and dual luciferase reporter system plasmid (50 ng) of different lengths using the PEI method. After 24 hours of culture, the expression intensities of Fluc and nLuc were detected using a microplate reader with a dual luciferase reporter gene detection kit (Beyotime, RG028), and the cleavage activity of each CasY6 mutant was characterized by the nLuc fluorescence intensity. Spacer sequences within the range of 15-25 nucleotide lengths were all effective for the spacer sequence-specific cleavage activity of CasY6-V3.3, and among them, 18-20 nt showed the best cleavage activity. The following table shows spacer sequences of different lengths and their corresponding cleavage activities (the 20nt spacer sequence in the table is the spacer sequence (sequence number: 48) that targets the TTR-3 locus in Example 6). [Table 17] The above examples demonstrate that the CRISPR-CasY6 system can achieve specific and efficient gene editing in prokaryotic and mammalian cells, and is superior to other Cas12 systems (e.g., the LbCpf1 system). The CasY6 protein of this disclosure has a small volume and crRNA length, making it suitable for multiple gene editing applications in the subject's body, and can be delivered in various ways (e.g., LNP, AAV), showing great potential for therapeutic gene editing applications.
[0542] The engineering design of the CasY6 protein has confirmed that the CasY6 system can be used for base editing applications. Regarding base editing, the D1069A-dCasY6 system effectively edits multiple sites at the TTR locus. At target sites mediated by CasY6-TTR-sgRNA', editing efficiencies at positions +13, +15, +16, and +17 reach approximately 10% to 30%. At target sites mediated by CasY6-TTR-sgRNA'', base editing efficiencies at position +13 reach over 10%, and at position +20 reach over 15%. This indicates greater potential as a base editor.
[0543] Site-directed mutagenesis and screening yielded 52 mutants, including CasY6-V1.1-CasY6-V3.13, several of which exhibited specific and more efficient editing activity.
[0544] In short, the CasY6 system of this disclosure possesses robust editing activity and high specificity, functioning as a versatile platform for genome editing or base editing in mammalian cells, and potentially useful for in-vivo or ex-vivo therapeutic applications in the future.
[0545] Sequence information: [Table 18-1] [Table 18-2] [Table 18-3] [Table 18-4] [Table 18-5] [Table 18-6] Table 18-7 Table 18-8 Table 18-9 Table 18-10 Table 18-11 Table 18-12 Table 18-13 Table 18-14 Table 18-15 Table 18-16 Table 18-17 Table 18-18 Table 18-19 Table 18-20 Table 18-21 Table 18-22 Table 18-23 [Table 18-24] [Table 18-25] [Table 18-26] All documents referenced herein are incorporated herein by reference as if each document were cited individually. Furthermore, it should be understood that a person skilled in the art may make various changes or modifications to the invention after reviewing the above, and that these equivalent forms are also limited in scope by the claims appended to this application.
Claims
1. A Cas protein selected from the following group, (a) Polypeptide having the amino acid sequence shown in Sequence ID No. 1; (b) Polypeptides having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% or more homology (or identity) with the amino acid sequence shown in Sequence ID No. 1, and having the biological function of Sequence ID No. 1; (c) Inducible polypeptides having one or more (preferably 1 to 20, more preferably 1 to 10, more preferably 1 to 5) amino acid residues substituted, deleted, or added in the amino acid sequence shown in Sequence ID No. 1, and retaining the biological function of Sequence ID No. 1; Preferably, the Cas protein is E118, G119, D120, V121, Q122, Y123, K241, S242, S243, S246, L272, S275, C277, T285, S286, C287, D288, E289, A290, I292, K299, R300, W301, L394, D395 in SEQ ID NO: 1 , selected from the group consisting of R396, D397, R398, K399, E400, D401, E402, D403, S404, C405, R406, F407, E408, K502, K557, L657, C672, P758, L789, K827 and V860, containing mutations in one or more important amino acid sites related to cleavage activity, Preferably, the Cas protein contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of L272, S275, and C277 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in two or more important amino acid sites related to cleavage activity, selected from the group consisting of E289, A290, and I292 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of K399, E400, D401, E402, and D403 in SEQ ID NO:
1. Preferably, the Cas protein further includes mutations in three or more important amino acid sites related to cleavage activity, selected from the group consisting of D403, S404, C405, R406, F407, and E408 in SEQ ID NO: 1, and preferably, the Cas protein further includes mutations in four or more important amino acid sites related to cleavage activity, selected from the group consisting of L394, D395, R396, D397, and R398 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in five or more important amino acid sites related to cleavage activity, selected from the group consisting of D397, R398, K399, E400, D401, and E402 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of S242, S243, and S246 in SEQ ID NO: 1, and preferably, the Cas protein further contains mutations in four or more important amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, and Y123 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of D397, R398, E400R, D401, E402, and L789 in SEQ ID NO:
1. Preferably, the Cas protein further includes mutations in important amino acid sites related to cleavage activity, selected from the group consisting of T285, S286, C287, D288, D403, C405, R406, F407, E408, K502, and K557 in SEQ ID NO: 1, and preferably, the Cas protein further includes mutations in important amino acid sites related to cleavage activity, selected from the group consisting of E289, I292, K399, E400, D401, E402, and D403 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of E289, A290, I292, P758 and V860 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of K299, R300, W301, C672, and P758 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of D397, R398, E400, D401, E402, L657, and L789 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, C672, and P758 in SEQ ID NO:
1. Preferably, the Cas protein further includes mutations in important amino acid sites related to cleavage activity, selected from the group consisting of K399, E400, D401, E402, D403, P758, and K827 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in important amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, L394, D395, R396, D397, R398, C672, and P758 in SEQ ID NO:
1. Preferably, the Cas protein further contains mutations in 12 or more important amino acid sites related to cleavage activity, selected from the group consisting of E118, G119, D120, V121, Q122, D403, S404, C405, R406, F407, E408, C672 and P758 in SEQ ID NO:
1. Preferably, the Cas protein further includes mutations in important amino acid sites related to cleavage activity, selected from the group consisting of K241, S242, S243, S246, D397, R398, E400, D401, E402, and L789 in SEQ ID NO:
1. A Cas protein characterized by the following features.
2. It is a protein variant that is a non-natural protein, The wild-type protein corresponds to Sequence ID No. 1 and contains a mutation in one or more important amino acid sites related to cleavage activity, selected from the group consisting of the following: The 659th aspartic acid (D) site; and / or The 711th aspartic acid (D) site; and / or Glutamate (E) site 895; and / or The 1069th aspartic acid (D) site; Preferably, the amino acid substitution is one or more selected from D659A, D711A, E895A and D1069A. A protein variant characterized by the following features.
3. A fusion protein comprising the Cas protein described in claim 1 or the protein variant described in claim 2, and one or more functional domains.
4. An isolated polynucleotide characterized by encoding the Cas protein described in claim 1, the protein variant described in claim 2, or the fusion protein described in claim 3.
5. An isolated nucleic acid molecule comprising or composed of a sequence selected from the following: (i) Sequence ID: the sequence shown in 5; (ii) Sequences obtained by substituting, deleting, or adding one or more bases to the sequence shown in Sequence ID No. 5 (for example, one, two, three, four, five, six, seven, eight, nine, or ten bases being substituted, deleted, or added); (iii) Sequences having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, and at least 95% sequence identity with the sequence shown in sequence number 5; (iv) A sequence that hybridizes with any one of the sequences described in (i) to (iii) under stringent conditions; or (v) A sequence complementary to any one of the sequences described in (i) to (iii); The sequence described in any one of the above paragraphs (ii) to (v) substantially retains the biological function of the sequence from which it is derived, The isolated nucleic acid molecule is, for example, RNA. The isolated nucleic acid molecule includes, for example, a direct repeat sequence in the CRISPR / Cas system. An isolated nucleic acid molecule characterized by the following.
6. A guide RNA (gRNA) comprising a direct repeat (DR) sequence capable of binding to the Cas protein described in claim 1 or the protein variant described in claim 2, and a spacer sequence capable of targeting a target sequence.
7. (i) a protein component selected from the group consisting of the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, or a combination thereof; and (ii) A nucleic acid component selected from the group consisting of the guide RNA described in claim 6, the nucleic acid encoding the guide RNA described in claim 6, the precursor RNA of the guide RNA described in claim 6, the nucleic acid encoding the precursor RNA of the guide RNA described in claim 6, or a combination thereof. Includes, The protein component and nucleic acid component are bonded to each other, A composite device characterized by the following features.
8. A vector characterized by comprising the polynucleotide described in claim 4.
9. (i) A first component selected from the group consisting of the Cas protein according to claim 1, the protein variant according to claim 2, the fusion protein according to claim 3, a nucleotide sequence encoding the Cas protein according to claim 1 or the protein variant according to claim 2 or the fusion protein according to claim 3, and any combination thereof; and (ii) A second component which is a nucleotide sequence comprising one or more guide RNAs according to claim 6, or a nucleotide sequence encoding the guide RNAs according to one or more claims 6. Includes, The guide RNA can form a complex with the protein, protein variant, or fusion protein described in (i). A CRISPR-Cas composition characterized by the following features.
10. (i) a first nucleic acid which is a nucleotide sequence encoding the Cas protein according to claim 1, the protein variant according to claim 2, or the fusion protein according to claim 3, and which is optionally operably linked to the first regulatory element; and (ii) A second nucleic acid encoding a nucleotide sequence containing the guide RNA described in claim 6, which is optionally operably linked to a second regulatory element. Includes one or more vectors containing The first nucleic acid and the second nucleic acid are present in the same or different vectors. The guide RNA can form a complex with the Cas protein or fusion protein described in (i). A CRISPR-Cas system characterized by the following features.
11. A kit comprising one or more components selected from the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the polynucleotide described in claim 4, the complex described in claim 7, the vector described in claim 8, the CRISPR-Cas composition described in claim 9, or the system described in claim 10.
12. A delivery composition comprising a delivery vector and one or more selected from the following: the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the polynucleotide described in claim 4, the complex described in claim 7, the vector described in claim 8, the CRISPR-Cas composition described in claim 9, or the system described in claim 10.
13. A host cell comprising the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the polynucleotide described in claim 4, the complex described in claim 7, the vector described in claim 8, the composition described in claim 9, the system described in claim 10, or the delivery composition described in claim 12.
14. An enzyme preparation comprising the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the complex described in claim 7, the CRISPR-Cas composition described in claim 9, the system described in claim 10, or the delivery composition described in claim 12.
15. A pharmaceutical kit comprising a first container, and a complex according to claim 7, a composition according to claim 9, or a system according to claim 10, which are disposed within the first container, or a pharmaceutical product containing the complex according to claim 7, the composition according to claim 9, or the system according to claim 10.
16. (a1) A first container, and a pharmaceutical product comprising the Cas protein according to claim 1, the protein variant according to claim 2, the fusion protein according to claim 3, or their coding genes or expression vectors, disposed within the first container; or the Cas protein according to claim 1, the protein variant according to claim 2, the fusion protein according to claim 3, or their coding genes or expression vectors; (b1) A second container of any choice, and a guide RNA or expression vector thereof according to claim 6, or a pharmaceutical product containing the guide RNA or expression vector thereof according to claim 6, which is placed in the second container. A pharmaceutical kit characterized by containing the following.
17. The method comprises contacting a target gene with a Cas protein according to claim 1, a protein variant according to claim 2, a fusion protein according to claim 3, a complex according to claim 7, a composition according to claim 9, a system according to claim 10, a delivery composition according to claim 12, an enzyme preparation according to claim 14, or a drug kit according to claim 15 or 16, or delivering the said target gene to a cell containing the said target gene. The aforementioned target gene contains a target sequence. A method for targeting and editing a target gene or cleaving a target gene, characterized by the above.
18. A method for inducing a change in cellular state, characterized by bringing into contact with a target gene in a cell the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the complex described in claim 7, the composition described in claim 9, the system described in claim 10, the delivery composition described in claim 12, the enzyme preparation described in claim 14, or the pharmaceutical kit described in claim 15 or 16.
19. The method comprises contacting a nucleic acid molecule encoding a gene product with the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the complex described in claim 7, the composition described in claim 9, the system described in claim 10, the delivery composition described in claim 12, the enzyme preparation described in claim 14, or the drug kit described in claim 15 or 16, with a nucleic acid molecule encoding the gene product, or delivering the nucleic acid molecule to a cell containing the nucleic acid molecule. The nucleic acid molecule contains a target sequence. A method for modifying the expression of a gene product, characterized by the following:
20. Uses of the Cas protein according to claim 1, the protein variant according to claim 2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9, the system according to claim 10, the kit according to claim 11, the delivery composition according to claim 12, the enzyme preparation according to claim 14, or the pharmaceutical kit according to claim 15 or 16, characterized in that they are used in the manufacture of pharmaceuticals or preparations for nucleic acid editing (e.g., gene or genome editing).
21. Uses of the Cas protein according to claim 1, the protein variant according to claim 2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9, the system according to claim 10, the kit according to claim 11, the delivery composition according to claim 12, the enzyme preparation according to claim 14, or the pharmaceutical kit according to claim 15 or 16, characterized in that they are used in the manufacture of pharmaceuticals or preparations for one or more items selected from the following group. (i) ex vivo gene or genome editing; (ii) Detection of single-stranded DNA in ex-vivo; (iii) Modification of a living or non-human organism by editing a target sequence at a target gene locus; (iv) Treatment of disorders caused by deletion of target sequences at target gene loci; (v) Treatment of a disability or disease of a subject that is in need.
22. The method involves contacting a sample with the Cas protein described in claim 1, the protein variant described in claim 2, the fusion protein described in claim 3, the complex described in claim 7, the CRISPR-Cas composition described in claim 9, the system described in claim 10, the kit described in claim 11, the delivery composition described in claim 12, or the enzyme preparation described in claim 14, and a non-target sequence, and detecting a detectable signal resulting from the cleavage of the non-target sequence to detect a target nucleic acid molecule. A method for detecting the presence of a target nucleic acid molecule in a sample, characterized in that the non-target sequence does not hybridize with the guide RNA.