Cas protein, variant of Cas protein, corresponding gene editing system of Cas protein and application of Cas protein and variant of Cas protein

By mutation of Cas protein at specific sites, a new Cas protein variant was developed, which improved the cleavage activity and specificity of the target molecule, solved the differential problem of the existing CRISPR/Cas system in the target recognition and cleavage process, and achieved efficient gene editing applications.

CN120330161APending Publication Date: 2025-07-18YOLTECH THERAPEUTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311497160.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are differences in the existing CRISPR/Cas systems in targeted recognition and cleavage, which is difficult to meet diverse application needs, and the cleavage activity and specificity need to be improved.

Method used

A new Cas protein variant is developed to increase the cleavage activity and specificity of target molecules by mutations at specific amino acid sites, combining guide RNA to form complexes for efficient gene editing.

Benefits of technology

It improves the cleavage activity and specificity of Cas protein on target molecules, enhances the gene editing efficacy of the CRISPR/Cas system, and is suitable for the treatment of various diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004543082030000181
    Figure BDA0004543082030000181
  • Figure BDA0004543082030000191
    Figure BDA0004543082030000191
  • Figure BDA0004543082030000192
    Figure BDA0004543082030000192
Patent Text Reader

Abstract

The invention provides a Cas protein as well as a variant, a corresponding gene editing system and application thereof, and particularly, the Cas protein variant disclosed by the invention has better gene editing activity compared with a wild type Cas protein, can be used for effectively editing or cutting a target gene and can be used for effectively treating diseases or diseases of a subject in need.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gene editing, and specifically, to a Cas protein and its variants, its corresponding gene editing system and applications. Background Art

[0002] The Clustered regularly interspaced short palindromic repeats (CRISPR) system is formed by bacteria and archaea to defend against the DNA of invading phages. The most common one is the CRISPR / Cas9 system. The Cas9 protein can process pre-crRNA into mature crRNA that binds to tracrRNA with the assistance of trans-encoded small RNA (tracrRNA). Subsequently, it was found that by artificially constructing a single-stranded chimeric guide RNA (gRNA) that mimics the crRNA-tracrRNA complex, the Cas9 protein can be effectively mediated to recognize and cleave the target site. The three bases immediately adjacent to the 3' end of the target site must be in the form of 5'-NGG-3', thus forming the PAM (protospacer adjacent motif) structure required for the Cas / crRNA complex to recognize the target site.

[0003] However, the existing different CRISPR / Cas systems have different advantages and disadvantages. For example, there are differences in the sizes of different Cas proteins, guide RNAs, and PAMs.

[0004] Therefore, there is still a need to develop new Cas proteins and CRISPR-Cas systems to meet diverse application requirements. Summary of the Invention

[0005] The main object of the present invention is to provide a new Cas protein, its gene editing system and applications to meet the above application requirements. Based on this, the present invention also provides a new CRISPR-Cas composition, as well as a gene editing method and a nucleic acid detection method based on this system.

[0006] The first aspect of the present invention provides a Cas protein, which is selected from the following groups:

[0007] (a) a polypeptide having the amino acid sequence shown in SEQ ID NO:1;

[0008] (b) a polypeptide having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% homology (or identity) with the amino acid sequence shown in SEQ ID NO:1, and the polypeptide has the biological function of SEQ ID NO:1;

[0009] (c) a derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1 - 20, more preferably 1 - 10, still more preferably 1 - 5) amino acid residues in any of the amino acid sequences shown in SEQ ID NO:1, and retaining the biological function of SEQ ID NO:1.

[0010] In another preferred embodiment, the Cas protein includes wild - type and mutant Cas proteins (i.e., orthologs, homologs, variants or functional fragments of the Cas protein, wherein the orthologs, homologs, variants or functional fragments substantially retain the biological function of the sequence from which they are derived).

[0011] In certain embodiments, the Cas protein is a variant of the protein shown by the amino acid sequence of SEQ ID NO:1.

[0012] In a preferred embodiment, the Cas protein is non - naturally occurring.

[0013] In certain embodiments, relative to the cleavage activity of the wild - type Cas protein, the cleavage activity of the Cas protein variant against the target sequence of the target molecule complementary to the guide RNA sequence is increased or equivalent to the cleavage activity of the wild - type Cas protein, or the binding of the CRISPR composition containing the Cas protein variant to the binding site is enhanced or the editing preference is changed.

[0014] In certain embodiments, the Cas protein variant significantly increases the spacer - specific endonuclease cleavage activity against the target sequence of the target DNA complementary to the spacer (e.g., increased by at least 5%, preferably at least 10%, more preferably at least 20%, such as at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%, for example, at least doubled, at least tripled, at least quadrupled, at least quintupled, at least sextupled, at least octupled, at least nonupled, at least decupled) compared to the wild - type Cas protein.

[0015] In certain embodiments, the Cas protein relative to the polypeptide of the amino acid sequence shown in SEQ ID NO:1:

[0016] (1) Comprising mutations at one or more sites such that the cleavage activity or preference is altered;

[0017] In certain embodiments, the mutation occurs at any site of the amino acid sequence shown in SEQ ID NO:1;

[0018] In certain embodiments, the Cas protein relative to the polypeptide of the amino acid sequence shown in SEQ ID NO:1: (2) Mutations occur at one or more core amino acid sites corresponding to SEQ ID NO:1 in the wild-type Cas protein and selected from the following group:

[0019] V7, T11, S12, Y165, S166, S196, K197, S198, A199, S240, S241, Q242, E243, I244, N276, Y282, D283, A285, I304, L307, Y308, S309, R315, E316, T317, I318, I319, V348, I349, E350, P351, G416, I417, E418, F419, D420, E648, L650, A651, Y652, T681, N682, E683, S684, E748, G749, S751, K865, P866, Y867, N868, I872, M957, F958, Q960, W961.

[0020] In certain embodiments, the amino acid substitutions at V7, T11, S12, Y165, S166, S196, K197, S198, A199, S240, S241, Q242, E243, I244, N276, Y282, D283, A285, I304, L307, Y308, S309, R315, E316, T317, I318, I319, V348, I349, E350, P351, G416, I417, E418, F419, D420, E648, L650, A651, Y652, T681, N682, E683, S684, E748, G749, S751, K865, P866, Y867, N868, I872, M957, F958, Q960, W961 are selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C.

[0021] In certain embodiments,

[0022] (a) The amino acid substitution at D283 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from D283N, D283Q, D283S, D283E; further preferably, from D283N, D283Q, D283S;

[0023] (b) The amino acid substitution at A285 is selected from A285V, A285L, A285I, A285S, A285T, A285G; preferably, from A285V, A285T, A285P; more preferably, from A285T, A285P;

[0024] (c) The amino acid substitution at Y308 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from Y308W, Y308F, Y308T, Y308S, Y308A; even more preferably, from Y308A, Y308F;

[0025] (d) The amino acid substitution at R315 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; preferably, R315K, R315Q, R315N, R315H, R315A, R315S; preferably, from R315K, R315A, R315S; more preferably, from R315A, R315S;

[0026] (e) The amino acid substitution at E316 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from E316D, E316R, E316Q; preferably, from E316R, E316Q;

[0027] (f) The amino acid substitution at T317 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from T317S, T317K, T317R, T317N, T317Q, T317A; even more preferably, from T317S, T317K, T317R; further preferably, from T317K, T317R;

[0028] (g) The amino acid substitution at I318 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; preferably, from I318L, I318V, I318M, I318A, I318F, I318G, I318K; more preferably, from I318L, I318K;

[0029] (h) The amino acid substitution at I319 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; preferably, from I319L, I319V, I319M, I319A, I319F, I319Q, I319W, I319G, preferably, from I319L, I319Q, I319W, more preferably, from I319Q, I319W;

[0030] (i) The amino acid substitution at G416 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from G416P, G416A, G416V, G416L, G416I, G416S, G416R; more preferably, from G416A, G416S, G416R; even more preferably, from G416S, G416R;

[0031] (j) The amino acid substitution at I417 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; preferably, from I417L, I417V, I417M, I417A, I417F, I417G, I417H; more preferably, from I417L, I417H, I417V; further preferably, from I417H, I417V;

[0032] (k) The amino acid substitution at F419 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from F419L, F419V, F419I, F419A, F419Y, F419W, F419T; more preferably, from F419L, F419T, F419V; further preferably, from F419T, F419V;

[0033] (l) The amino acid substitution at D420 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from D420E, D420T, D420V, further preferably, from D420T, D420V;

[0034] (m) The amino acid substitution at E648 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from E648D, E648L, E648R, E648S, E648T, preferably, from E648L, E648R, E648S, E648T;

[0035] (n) The amino acid substitution at A651 is selected from any amino acid; preferably, from the amino acid types shown in Table A, Table B, and Table C; more preferably, from A651V, A651L, A651I, A651G, A651P, A651S, A651T; even more preferably, from A651V, A651L, A651G, A651P, A651S; still further preferably, from A651L, A651G, A651P, A651S.

[0036] In certain embodiments, the mutations are selected from the group consisting of: V7I, T11V, S12M, Y165W, S166R, S196N, K197F, S198N, A199T, S240L, S241A, Q242S, E243H, I244L, N276S, Y282F, D283N, D283Q, D283S, A285T, A285P, I304A, L307R, Y308A, Y308F, S309A, R315A, R315S, E316R, E316Q, T317K, T317R, I318K, I318L, I319Q, I319W, V348I, I349V, E350L, P351T, G416S, G416R, I417H, I417V, E418Q, F419T, F419V, D420T, D420V, E648L, E648R, E648S, E648T, L650I, A651L, A651G, A651P, A651S, Y652F, T681P, N682P, E683Q, S684C, E748A, G749S, S751G, K865S, P866S, Y867H, N868F, I872L, M957V, F958L, Q960V, W961V or combinations thereof.

[0037] In certain embodiments, the mutations are selected from the group consisting of:

[0038] Y282F + D283Q + A285T + G416S + I417H + E418Q + F419T + D420V;

[0039] Y165W + S166R + G416R + I417V + E418Q + F419V + D420T;

[0040] Y282F + D283Q + A285T + V348I + I349V + E350L + P351T;

[0041] D283S + A285P + L307R + Y308F + S309A;

[0042] L307R + Y308F + S309A + E648R + L650I + A651P + Y652F;

[0043] Y282F + D283Q + A285T + E648T + A651S + Y652F;

[0044] Y282F + D283Q + A285T + E648L + A651G;

[0045] Y282F + D283Q + A285T + T681P + N682P + E683Q + S684C;

[0046] Y282F + D283Q + A285T + E748A + G749S + S751G;

[0047] N276S + D283N + E648S + A651L;

[0048] Y282F + D283Q + A285T + R315A + E316Q + T317R + I318L + I319Q;

[0049] Y282F + D283Q + A285T + R315S + E316R + T317K + I318K + I319W + E648S + A651L;

[0050] V7I + T11V + S12M + Y282F + D283Q + A285T;

[0051] S240L + S241A + Q242S + E243H + I244L + L307R + Y308F + S309A;

[0052] S196N + K197F + S198N + A199T + I304A + L307R + Y308A + S309A;

[0053] Y282F + D283Q + A285T + K865S + P866S + Y867H + N868F + I872L;

[0054] Y282F + D283Q + A285T + M957V + F958L + Q960V + W961V (numbered as in SEQ ID NO:1).

[0055] In another preferred example, except for the above mutation sites, the amino acid sequence of the described Cas protein variant is the same as or substantially the same as the sequence of the wild-type Cas protein.

[0056] In certain embodiments, "substantially the same" means that there are at most 50 (preferably 1 - 20, more preferably 1 - 10, still more preferably 1 - 5) amino acids that are different, where the differences include amino acid substitutions, deletions, or additions, and the gene editing activity of the Cas protein variant is enhanced.

[0057] In certain embodiments, the Cas protein variant has at least 80% homology with the wild - type Cas protein, preferably at least 85% or 90%, more preferably at least 95%, and most preferably at least 98% or 99%.

[0058] In certain embodiments, the Cas protein variant is formed by mutating the wild - type Cas protein.

[0059] A second aspect of the present invention provides a fusion protein comprising the Cas protein according to the first aspect of the present invention; and one or more functional domains.

[0060] In certain embodiments, the functional domain is selected from a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA - binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage - active polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, and a base excision repair inhibitor (such as uracil - DNA glycosylase inhibitor (UGI)).

[0061] In certain embodiments, the functional domain includes one or more of the following enzymatic activities towards a target sequence: methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O - GlcNAc transferase), and deglycosylation activity.

[0062] In certain embodiments, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0063] In certain embodiments, the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0064] In some embodiments, the adenosine deaminase catalytic domain comprises an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% identity to the amino acid sequence shown in SEQ ID NO: 31 (005V1 deaminase selected from CN114634923A, the amino acid sequence in this application is SEQ ID NO: 2), and it retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 31.

[0065] In some embodiments, the amino acid sequence of the adenosine deaminase catalytic domain has amino acid additions, insertions, deletions and substitutions relative to the amino acid sequence shown in SEQ ID NO: 31.

[0066] In some embodiments, the adenosine deaminase catalytic domain comprises a mutant of the amino acid sequence shown in SEQ ID NO: 31: Q148G+Q149M+P150R, named deaminase005V1-10-3.

[0067] In some embodiments, the functional domain is the full length or a functional fragment of TadA8e.

[0068] In some embodiments, the localization signal comprises a nuclear localization signal (NLS) and / or a nuclear export signal (NES).

[0069] In some embodiments, the sequence of the nuclear localization signal is as shown in any of SEQ ID NOs: 38-45.

[0070] In some embodiments, the sequence of the nuclear localization signal is located at, near or close to the end (e.g., N-terminus or C-terminus) of the variant recited in claim 1.

[0071] In some embodiments, the nuclear export signal comprises protein tyrosine kinase 2 (such as human protein tyrosine kinase 2).

[0072] In some embodiments, the reporter protein comprises glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, a self-fluorescent protein.

[0073] In certain embodiments, the autofluorescent protein includes green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyan1, etc.), yellow fluorescent protein (e.g., (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire).

[0074] In certain embodiments, the DNA binding domain includes a methylation binding protein, LexA DBD, Gal4 DBD.

[0075] In certain embodiments, the epitope tag includes a histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, thioredoxin tag, streptavidin tag.

[0076] In certain embodiments, the transcriptional activation domain includes VP64 and / or VPR.

[0077] In certain embodiments, the transcriptional repression domain includes KRAB and / or SID.

[0078] In certain embodiments, the nuclease includes FokI.

[0079] In certain embodiments, the cleavage-active polypeptide includes a polypeptide with single-stranded RNA cleavage activity, a polypeptide with double-stranded RNA cleavage activity, a polypeptide with single-stranded DNA cleavage activity, or a polypeptide with double-stranded DNA cleavage activity.

[0080] In certain embodiments, the ligase includes DNA ligase and / or RNA ligase.

[0081] In certain embodiments, the functional domain is connected to the N-terminus and / or C-terminus of the Cas protein variant.

[0082] In certain embodiments, the functional domain is inserted between the N-terminus and C-terminus of the Cas protein variant.

[0083] In certain embodiments, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the Cas protein variant via a linker.

[0084] In some embodiments, the functional domain is inserted between the N-terminus and C-terminus of the Cas protein variant via a linker.

[0085] In some embodiments, the fusion protein has the following structure from the N-terminus to the C-terminus:

[0086] Z1-Z2 (I); or

[0087] Z2-Z1 (II); or

[0088] Z3-Z1-Z4 (III);

[0089] wherein, Z1 is a cytosine deaminase or an adenosine deaminase;

[0090] Z2 is the Cas protein described in the first aspect of the present invention;

[0091] Z3 is an N-terminal fragment of the Cas protein described in the first aspect of the present invention;

[0092] Z4 is a C-terminal fragment of the Cas protein described in the first aspect of the present invention;

[0093] and each "-" is independently a bond or a linker.

[0094] In another preferred example, the fusion protein has the amino acid sequence shown in SEQ ID NO. 46.

[0095] The third aspect of the present invention provides an isolated polynucleotide encoding the Cas protein described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention.

[0096] In some embodiments, the isolated nucleotide includes a sequence optimized by humanization.

[0097] In some embodiments, the polynucleotide further contains an auxiliary element selected from the following group flanking the ORF of the variant: a signal peptide, a secretion peptide, a tag sequence (such as 6His), or a combination thereof.

[0098] In some embodiments, the polynucleotide is selected from the following group: a genomic sequence, a cDNA sequence, an RNA sequence, or a combination thereof.

[0099] In some embodiments, the polynucleotide further contains a promoter operably linked to the ORF sequence of the variant.

[0100] In some embodiments, the promoter is selected from the following group: a constitutive promoter, a tissue-specific promoter, an inducible promoter, or a strong promoter.

[0101] In certain embodiments, the polynucleotide is a polynucleotide codon-optimized according to the codon preference of the host cell.

[0102] In certain embodiments, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0103] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).

[0104] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.

[0105] In certain embodiments, the yeast cell is selected from yeast of one or more sources in the following group: Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes: Kluyveromyces, more preferably Kluyveromyces marxianus and / or Kluyveromyces lactis.

[0106] In certain embodiments, the host cell is selected from the following group: Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.

[0107] The fourth aspect of the present invention provides an isolated nucleic acid molecule, comprising a sequence selected from the following, or consisting of a sequence selected from the following:

[0108] (i) the sequence shown in SEQ ID NO: 5 or 71;

[0109] (ii) a sequence having one or more base substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base substitutions, deletions, or additions) compared to the sequence shown in SEQ ID NO: 5 or 71;

[0110] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to the sequence shown in SEQ ID NO: 5 or 71;

[0111] (iv) a sequence that hybridizes to the sequence described in any one of (i)-(iii) under stringent conditions; or

[0112] (v) the complementary sequence of the sequence described in any one of (i)-(iii);

[0113] And, the sequence described in any one of (ii)-(v) substantially retains the biological function of the sequence from which it is derived;

[0114] For example, the isolated nucleic acid molecule is RNA;

[0115] For example, the isolated nucleic acid molecule comprises a direct repeat (DR) sequence in the CRISPR / Cas system.

[0116] In another preferred embodiment, the nucleic acid molecule comprises one or more stem-loops or optimized secondary structures;

[0117] For example, the sequence of any one of (ii)-(v) retains the secondary structure of the sequence from which it is derived.

[0118] In another preferred embodiment, the DR sequence comprises a stem-loop structure near the 3' end of the DR sequence, and the stem-loop structure comprises 5'-X1X2X3X4X5NNNNNNNX6X7X8X9X 10 -3'); X1, X2, X3, X4, X5, X6, X7, X8, X9, X 10 is any base comprising A, T, C or G, and N is any base comprising A, T, C or G; wherein X1, X2, X3, X4, X5 and X6, X7, X8, X9, X 10 can hybridize to each other to form a stem and such that NNNNNNN forms a loop; more preferably, wherein the DR sequence comprises a stem-loop structure near the 3' end of the DR sequence as shown below:

[0119] 5'-CCGTCNNNNNNNGACGG-3' (SEQ ID NO.80);

[0120] wherein, N is any base comprising A, T, C or G.

[0121] In another preferred embodiment, the nucleic acid molecule comprises a sequence selected from the following, or consists of a sequence selected from the following:

[0122] (a) the nucleotide sequence shown in SEQ ID NO: 5 or 71;

[0123] (b) a sequence that hybridizes to the sequence described in (a) under stringent conditions; or

[0124] (c) a complementary sequence of the nucleotide sequence shown in SEQ ID NO: 5 or 71.

[0125] The fifth aspect of the present invention provides a guide RNA (gRNA), which comprises a direct repeat (DR) sequence capable of binding to the Cas protein described in the first aspect of the present invention and a spacer sequence capable of targeting a target sequence.

[0126] The sixth aspect of the present invention provides a complex, comprising:

[0127] (i) A protein component selected from the group consisting of: the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, or a combination thereof; and

[0128] (ii) A nucleic acid component selected from the group consisting of: the guide RNA described in the fifth aspect of the present invention, the nucleic acid encoding the guide RNA described in the fifth aspect of the present invention, the precursor RNA of the guide RNA described in the fifth aspect of the present invention, the nucleic acid of the precursor RNA encoding the guide RNA described in the fifth aspect of the present invention, or a combination thereof;

[0129] Wherein, the protein component and the nucleic acid component bind to each other to form a complex.

[0130] In certain embodiments, the direct repeat (DR) sequence in the guide RNA (gRNA) is linked to the 3'-end or 5'-end of the nucleic acid molecule.

[0131] In certain embodiments, the spacer sequence in the guide RNA (gRNA) contains a complementary sequence of the target sequence.

[0132] The seventh aspect of the present invention provides a vector comprising the polynucleotide described in the third aspect of the present invention.

[0133] In certain embodiments, the vector comprises:

[0134] (1) A first regulatory element, which is operably linked to a nucleotide sequence encoding the Cas protein described in the first aspect of the present invention or a nucleotide sequence encoding the fusion protein described in the second aspect of the present invention; and

[0135] (2) A second regulatory element, which is operably linked to a nucleotide sequence encoding a guide RNA, and the guide RNA comprises:

[0136] (a) A spacer sequence capable of hybridizing with a target sequence, and

[0137] (b) A direct repeat (DR) sequence, which is linked to the spacer sequence and can guide the Cas protein to bind to the guide RNA to form the complex described in the sixth aspect of the present invention targeting the target sequence.

[0138] In certain embodiments, the first regulatory element and the second regulatory element are located on the same or different vectors.

[0139] In certain embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.

[0140] In certain embodiments, the vector comprises one or more promoters, which are operably linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, origin of replication, selectable marker, nucleic acid restriction site, and / or homologous recombination site.

[0141] In certain embodiments, the vector includes a plasmid, a viral vector.

[0142] In certain embodiments, the viral vector is selected from the group consisting of: adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or a combination thereof.

[0143] In certain embodiments, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, a multifunctional vector.

[0144] The eighth aspect of the present invention provides a CRISPR-Cas composition, comprising:

[0145] (i) a first component selected from the group consisting of: the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, a nucleotide sequence encoding the Cas protein described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention, and any combination thereof; and

[0146] (ii) a second component, which is a guide RNA comprising one or more of the fifth aspect of the present invention, or a nucleotide sequence encoding the guide RNA comprising one or more of the fifth aspect of the present invention;

[0147] The guide RNA is capable of forming a complex with the protein or protein variant or fusion protein described in (i).

[0148] In certain embodiments, the guide RNA comprises a direct repeat sequence and a spacer sequence from the 5' to 3' direction, and the spacer sequence is capable of hybridizing with a target sequence.

[0149] In certain embodiments, the composition further comprises a pharmaceutically acceptable carrier.

[0150] In certain embodiments, the composition includes a pharmaceutical composition.

[0151] In certain embodiments, the dosage form of the composition is selected from the group consisting of: a freeze-dried preparation, a liquid preparation, or a combination thereof.

[0152] In certain embodiments, the dosage form of the composition is a liquid preparation.

[0153] In some embodiments, the dosage form of the composition is an injectable dosage form.

[0154] In some embodiments, the composition is a cell preparation.

[0155] The ninth aspect of the present invention provides a CRISPR-Cas system, comprising one or more vectors, and the one or more vectors comprise:

[0156] (i) a first nucleic acid, which is a nucleotide sequence encoding the Cas protein described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention; optionally, the first nucleic acid is operably linked to a first regulatory element; and

[0157] (ii) a second nucleic acid, which encodes a nucleotide sequence comprising the guide RNA described in the fifth aspect of the present invention; optionally, the second nucleic acid is operably linked to a second regulatory element;

[0158] Wherein:

[0159] the first nucleic acid and the second nucleic acid are present on the same or different vectors;

[0160] the guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

[0161] In some embodiments, the vector includes a plasmid, a viral vector.

[0162] In some embodiments, the guide RNA includes a spacer sequence capable of hybridizing with a target sequence; and a direct repeat (DR) sequence that is linked to the spacer sequence and is capable of guiding the protein to bind to the guide RNA, thereby forming a CRISPR-Cas composition or complex targeting the target sequence.

[0163] In some embodiments, the guide RNA includes unmodified and modified guide RNAs.

[0164] In some embodiments, the modified guide RNA includes chemical modifications of bases.

[0165] In some embodiments, the chemical modification includes methylation modification, methoxy modification, fluorination modification or thiolation modification.

[0166] In some embodiments, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.

[0167] In some embodiments, at least one component in the composition is non-naturally occurring or modified.

[0168] In certain embodiments, the spacer sequence is linked to the 3' end of the direct repeat (DR) sequence.

[0169] In certain embodiments, the spacer sequence comprises a complementary sequence of the target sequence.

[0170] In certain embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence of 5'-PAM being TTN, where N is A, T, C, or G.

[0171] In certain embodiments, the target sequence is DNA from a prokaryotic cell or a eukaryotic cell or a DNA sequence formed by reverse transcription based on RNA; alternatively, the target sequence is a non-naturally occurring DNA or a DNA sequence formed by reverse transcription based on RNA.

[0172] In certain embodiments, the target sequence includes a cDNA sequence.

[0173] In certain embodiments, the target sequence includes single-stranded DNA or double-stranded DNA sequences.

[0174] In certain embodiments, the target sequence is present within a cell.

[0175] In certain embodiments, the target sequence is present within the nucleus or within the cytoplasm (e.g., an organelle) of the cell.

[0176] In certain embodiments, the cell is a eukaryotic cell.

[0177] In certain embodiments, the cell is a prokaryotic cell.

[0178] In certain embodiments, the target sequence is present outside the cell.

[0179] In certain embodiments, the Cas protein according to the first aspect of the present invention is linked to one or more NLS sequences, or the fusion protein comprises one or more NLS sequences.

[0180] In certain embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the Cas protein according to the first aspect of the present invention.

[0181] In certain embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the Cas protein according to the first aspect of the present invention.

[0182] The tenth aspect of the present invention provides a kit, comprising one or more components selected from the following: the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the complex described in the sixth aspect of the present invention, the vector described in the seventh aspect of the present invention, the CRISPR-Cas composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention.

[0183] In certain embodiments, the kit further comprises a label or instructions.

[0184] In certain embodiments, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, cleaving a target gene or a non-target gene.

[0185] The eleventh aspect of the present invention provides a delivery composition, comprising a delivery vector and one or more selected from the following: the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the complex described in the sixth aspect of the present invention, the vector described in the seventh aspect of the present invention, the CRISPR-Cas composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention.

[0186] In certain embodiments, the delivery vector is a particle.

[0187] In certain embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0188] The twelfth aspect of the present invention provides a host cell, comprising the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, the polynucleotide described in the third aspect of the present invention, the complex described in the sixth aspect of the present invention, the vector described in the seventh aspect of the present invention, the composition described in the eighth aspect of the present invention, the system described in the ninth aspect of the present invention, or the delivery composition described in the eleventh aspect of the present invention.

[0189] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell or a mammalian cell (including human and non-human mammals).

[0190] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.

[0191] In certain embodiments, the yeast cell is yeast from one or more sources selected from the group consisting of Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell comprises Kluyveromyces, more preferably Kluyveromyces marxianus and / or Kluyveromyces lactis.

[0192] In certain embodiments, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.

[0193] The thirteenth aspect of the present invention provides an enzyme preparation, which comprises the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, the complex described in the sixth aspect of the present invention, the CRISPR-Cas composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention, or the delivery composition described in the eleventh aspect of the present invention.

[0194] In certain embodiments, the enzyme preparation comprises an injection and / or a freeze-dried preparation.

[0195] The fourteenth aspect of the present invention provides a kit, comprising:

[0196] A first container, and the complex described in the sixth aspect of the present invention, or the composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention, or a drug containing the complex described in the sixth aspect of the present invention, or the composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention, located in the first container.

[0197] In certain embodiments, the drug in the first container is a single-agent preparation containing the complex described in the sixth aspect of the present invention, or the composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention.

[0198] In certain embodiments, the dosage form of the drug is selected from the group consisting of a freeze-dried preparation, a liquid preparation, or a combination thereof.

[0199] In certain embodiments, the dosage form of the drug is an oral dosage form or an injection dosage form.

[0200] In certain embodiments, the kit further contains an instruction manual.

[0201] The fifteenth aspect of the present invention provides a kit, comprising:

[0202] (a1) A first container, and the Cas protein according to the first aspect of the present invention, or the fusion protein according to the second aspect of the present invention, or its coding gene or its expression vector, which are located in the first container, or a drug containing the Cas protein according to the first aspect of the present invention, or the fusion protein according to the second aspect of the present invention, or its coding gene or its expression vector;

[0203] (b1) Optionally, a second container, and the guide RNA according to the fifth aspect of the present invention or its expression vector, which are located in the second container, or a drug containing the guide RNA according to the fifth aspect of the present invention or its expression vector.

[0204] In certain embodiments, the first container and the second container are different containers.

[0205] In certain embodiments, the drug in the first container is a single formulation containing the Cas protein according to the first aspect of the present invention, or the fusion protein according to the second aspect of the present invention, or its coding gene or its expression vector.

[0206] In certain embodiments, the drug in the second container is a single formulation containing the guide RNA according to the fifth aspect of the present invention or its expression vector.

[0207] In certain embodiments, the dosage form of the drug is selected from the group consisting of: freeze-dried preparation, liquid preparation, or a combination thereof.

[0208] In certain embodiments, the dosage form of the drug is an oral dosage form or an injection dosage form.

[0209] In certain embodiments, the kit further contains an instruction manual.

[0210] The sixteenth aspect of the present invention provides a method for targeting and editing a target gene or cleaving a target gene, including: contacting the Cas protein according to the first aspect of the present invention, or the fusion protein according to the second aspect of the present invention, or the complex according to the sixth aspect of the present invention, or the composition according to the eighth aspect of the present invention, or the system according to the ninth aspect of the present invention, or the delivery composition according to the eleventh aspect of the present invention, or the enzyme preparation according to the thirteenth aspect of the present invention, or the kit according to the fourteenth or fifteenth aspect of the present invention with the target gene, or delivering it into a cell containing the target gene, and the target sequence exists in the target gene.

[0211] In certain embodiments, the target gene exists inside a cell.

[0212] In certain embodiments, the cell is a prokaryotic cell.

[0213] In certain embodiments, the cell is a eukaryotic cell, such as a mammalian cell (such as a human cell) or a plant cell.

[0214] In certain embodiments, the target gene is present in a nucleic acid molecule (e.g., plasmid) in vitro.

[0215] In certain embodiments, editing or cleaving the target gene includes breaking of the target sequence, such as double-strand break of DNA or single-strand break of RNA, or inserting an exogenous nucleic acid into the break.

[0216] In certain embodiments, the target gene comprises DNA.

[0217] In certain embodiments, the DNA comprises single-stranded DNA, double-stranded DNA.

[0218] The seventeenth aspect of the present invention provides a method for inducing a change in cell state, the method comprising contacting a target gene in a cell with the Cas protein described in the first aspect of the present invention, or the fusion protein described in the second aspect of the present invention, or the complex described in the sixth aspect of the present invention, or the composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention, or the delivery composition described in the eleventh aspect of the present invention, or the enzyme preparation described in the thirteenth aspect of the present invention, or the kit described in the fourteenth or fifteenth aspect of the present invention.

[0219] The eighteenth aspect of the present invention provides a method for altering the expression of a gene product, comprising: contacting the Cas protein described in the first aspect of the present invention, or the fusion protein described in the second aspect of the present invention, or the complex described in the sixth aspect of the present invention, or the composition described in the eighth aspect of the present invention, or the system described in the ninth aspect of the present invention, or the delivery composition described in the eleventh aspect of the present invention, or the enzyme preparation described in the thirteenth aspect of the present invention, or the kit described in the fourteenth or fifteenth aspect of the present invention with a nucleic acid molecule encoding the gene product, or delivering it into a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0220] In certain embodiments, the nucleic acid molecule is present in a nucleic acid molecule (e.g., plasmid) in vitro.

[0221] In certain embodiments, the expression of the gene product is altered (e.g., enhanced or reduced).

[0222] In certain embodiments, the gene product is a protein.

[0223] In certain embodiments, the protein, fusion protein, polynucleotide, isolated nucleic acid molecule, complex, vector or composition is comprised in a delivery vehicle.

[0224] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0225] In certain embodiments, one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product are altered to modify a cell, cell line or organism.

[0226] The nineteenth aspect of the present invention provides a cell or its progeny obtained by the method according to any one of the sixteenth to eighteenth aspects of the present invention, wherein the cell comprises a modification that does not exist in its wild type.

[0227] The twentieth aspect of the present invention provides a cell product of the cell or its progeny according to the nineteenth aspect of the present invention.

[0228] The twenty-first aspect of the present invention provides a cell or cell line or their progeny in vitro, ex vivo or in vivo, the cell or cell line or their progeny comprising: the Cas protein according to the first aspect of the present invention, the fusion protein according to the second aspect of the present invention, the polynucleotide according to the third aspect of the present invention, the complex according to the sixth aspect of the present invention, the vector according to the seventh aspect of the present invention, the CRISPR-Cas composition according to the eighth aspect of the present invention, or the system according to the ninth aspect of the present invention, or the delivery composition according to the eleventh aspect of the present invention.

[0229] In certain embodiments, the cell is a prokaryotic cell.

[0230] In certain embodiments, the cell is a eukaryotic cell, such as a mammalian cell (such as a human cell) or a plant cell.

[0231] In certain embodiments, the cell is a stem cell or a stem cell line.

[0232] The twenty-second aspect of the present invention provides the use of the Cas protein according to the first aspect of the present invention, the fusion protein according to the second aspect of the present invention, the polynucleotide according to the third aspect of the present invention, the complex according to the sixth aspect of the present invention, the vector according to the seventh aspect of the present invention, the CRISPR-Cas composition according to the eighth aspect of the present invention, or the system according to the ninth aspect of the present invention, or the kit according to the tenth aspect of the present invention, or the delivery composition according to the eleventh aspect of the present invention, or the enzyme preparation according to the thirteenth aspect of the present invention, or the kit according to the fourteenth or fifteenth aspect of the present invention, for the preparation of a drug or preparation for nucleic acid editing (such as gene or genome editing).

[0233] In certain embodiments, the gene or genome editing includes modifying a gene, knocking out a gene, altering the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide.

[0234] The twenty-third aspect of the present invention provides the use of the Cas protein according to the first aspect of the present invention, the fusion protein according to the second aspect of the present invention, the polynucleotide according to the third aspect of the present invention, the complex according to the sixth aspect of the present invention, the vector according to the seventh aspect of the present invention, the CRISPR-Cas composition according to the eighth aspect of the present invention, or the system according to the ninth aspect of the present invention, or the kit according to the tenth aspect of the present invention, or the delivery composition according to the eleventh aspect of the present invention, or the enzyme preparation according to the thirteenth aspect of the present invention, or the kit according to the fourteenth or fifteenth aspect of the present invention, for preparing a drug or preparation, and the drug or preparation is used for one or more selected from the following groups:

[0235] (i) Ex vivo gene or genome editing;

[0236] (ii) Detection of single-stranded DNA ex vivo;

[0237] (iii) Editing a target sequence at a target locus to modify a biological or non-human organism;

[0238] (iv) Treating a disorder caused by a defect in a target sequence at a target locus;

[0239] (v) Treating a disorder or disease of a subject in need thereof.

[0240] In certain embodiments, the disorder or disease includes cancer, infectious diseases, neurological diseases, ophthalmic diseases, and hearing diseases.

[0241] In certain embodiments, the disease or disorder includes cystic fibrosis, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy, alpha-1-antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell disease, beta-thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, hypercholesterolemia, transthyretin amyloidosis, retinal diseases, macular degeneration, Wilms tumor, Ewing sarcoma, neuroendocrine tumors, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and urinary bladder cancer.

[0242] In certain embodiments, the disorder or disease is caused by a pathogenic point mutation.

[0243] A twenty-fourth aspect of the present invention provides a method for detecting the presence of a target nucleic acid molecule in a sample, the method comprising contacting the sample with the Cas protein described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, or the complex described in the sixth aspect of the present invention, the CRISPR-Cas composition described in the eighth aspect of the present invention or the system described in the ninth aspect of the present invention, the kit described in the tenth aspect of the present invention or the delivery composition described in the eleventh aspect of the present invention or the enzyme preparation described in the thirteenth aspect of the present invention and a non-target sequence, and detecting a detectable signal generated by cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA.

[0244] In certain embodiments, if the non-target sequence is cleaved by the protein in the complex or CRISPR-Cas composition or system or delivery composition, it indicates the presence of the target nucleic acid molecule in the sample; while if the non-target sequence is not cleaved by the protein in the complex or CRISPR-Cas composition or system or delivery composition, it indicates the absence of the target nucleic acid molecule in the sample.

[0245] In certain embodiments, the target nucleic acid molecule is a target DNA.

[0246] In certain embodiments, the target DNA includes DNA formed by reverse transcription based on RNA.

[0247] In certain embodiments, the target DNA includes cDNA.

[0248] In some embodiments, the target DNA is selected from the group consisting of: single-stranded DNA, double-stranded DNA, or a combination thereof.

[0249] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as in the examples) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be elaborated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS

[0250] Figure 1 Shows the CasY7 recombinant expression plasmid ( Figure 1 A) and the map of the LbCpf1 recombinant expression plasmid ( Figure 1 B).

[0251] Figure 2 Shows the prediction of the secondary structure of the direct repeat (DR) sequence corresponding to CasY7.

[0252] Figure 3 Shows the map of the Target plasmid.

[0253] Figure 4 Shows the comparison of the editing efficiency of CasY7 and LbCpf1 in E. coli.

[0254] Figure 5 Shows the map of the PHK09T plasmid.

[0255] Figure 6 Shows the comparison of the editing efficiency of CasY7 and LbCpf1 in HEK293T cells.

[0256] Figure 7 Shows four mutants of CasY7, which eliminate the catalytic activity (cleavage activity) compared to the wild-type CasY7 protein.

[0257] Figure 8 Shows the activity detection of dCasY7 in base editing, where Figure 8 A shows the single-base editing (A>G) efficiency of the base editor composed of dCasY7 at the A2, A4, A13, A15-A17 sites of the TTR gene target; Figure 8 B shows the single-base editing (A>G) efficiency of the base editor composed of dCasY7 at the A7, A13, A20 of another target of the TTR gene.

[0258] Figure 9 Shows the comparison of the cleavage activity of different mutants of CasY7 on another target of the TTR gene.

[0259] Figure 10 The comparison of the cleavage activities of different mutants of CasY7 on the PD-1 gene is shown, where WT represents the cleavage activity of wild-type CasY7.

[0260] Figure 11 The comparison of the cleavage activities of different mutants of CasY7 on the Trac-1 target of the TRAC gene is shown.

[0261] Figure 12 The comparison of the cleavage activities of different mutants of CasY7 on the Trac-2 target of the TRAC gene is shown.

[0262] Figure 13 The comparison of the cleavage activities of different mutants of CasY7 on the Trac-3 target of the TRAC gene is shown.

[0263] Figure 14 The comparison of the cleavage activities of different mutants of CasY7 on the Trac-4 target of the TRAC gene is shown.

[0264] Figure 15 The secondary structure of the direct repeat (DR) sequence variant DR-1 is shown.

[0265] Figure 16 The comparison of the cleavage activities of CasY7 targeting the TTR gene mediated by different direct repeat (DR) sequences is shown.

[0266] Figure 17 shows the vector structure of the dual luciferase reporter system, where Figure 17A The LU×UC structure is shown. Figure 17B The map of the dual luciferase reporter plasmid is shown. Detailed implementation

[0267] Through extensive and in-depth research, the inventors of the present invention unexpectedly discovered a new Cas protein variant for the first time. The Cas protein variant of the present invention has equivalent or better gene editing activity compared to the wild-type Cas protein, can effectively edit or cleave the target gene, and can effectively treat the diseases or disorders of subjects in need. On this basis, the inventors of the present invention completed the present invention.

[0268] Terms

[0269] The following examples are only used to describe the present invention and do not limit the present invention. Unless otherwise specified, the experiments and methods described in the examples are basically carried out according to the conventional methods well-known in the art and described in various references.

[0270] In addition, for those conditions not specified in the examples, they are carried out according to conventional conditions or conditions recommended by the manufacturer. For reagents or instruments whose manufacturers are not indicated, they are all conventional products that can be obtained commercially. Those skilled in the art will understand that the examples describe the present invention by way of illustration and are not intended to limit the scope claimed by the present invention. All the publications and other reference materials mentioned herein are incorporated herein by reference in their entirety.

[0271] To make the present disclosure more readily understandable, certain terms are first defined. As used in this application, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.

[0272] The term "about" can refer to a value or a component within an acceptable error range of a specific value or component determined by one of ordinary skill in the art, which will depend in part on how the value or component is measured or determined. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0273] As used herein, the term "comprising" or "including (containing)" can be open-ended, semi-closed, and closed. In other words, the term also includes "consisting essentially of...", or "consisting of...".

[0274] Sequence identity (or homology) is determined by comparing two aligned sequences along a predefined comparison window, which can be 50%, 60%, 70%, 80%, 90%, 95% or 100% of the length of the reference nucleotide sequence or protein, and determining the number of positions at which identical residues occur. Generally, this is expressed as a percentage. Methods for measuring sequence identity of nucleotide sequences are well known to those skilled in the art.

[0275] Cas protein

[0276] The present invention provides a protein having the amino acid sequence shown in SEQ ID NO:1 or its orthologs, homologs, variants or functional fragments, wherein the orthologs, homologs, variants or functional fragments substantially retain the biological functions of the sequences from which they are derived (including the activity of binding to guide RNA, endonuclease activity, and the activity of specifically binding to and cleaving a target sequence at a specific site under the guidance of guide RNA).

[0277] In a preferred embodiment, the ortholog, homolog, variant, or functional fragment has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 99.5% sequence identity compared to the sequence from which it is derived (such as SEQ ID NO:1), and substantially retains the Cas protein activity of the sequence from which it is derived (e.g., the activity of binding to a guide RNA, endonuclease activity, the activity of binding to a specific site on a target sequence and cleaving under the guidance of a guide RNA).

[0278] Wild-type Cas protein

[0279] As used herein, "wild-type Cas protein" refers to a naturally occurring, unmodified Cas protein, the nucleotide of which can be obtained by genetic engineering techniques such as genomic sequencing, polymerase chain reaction (PCR), etc., and the amino acid sequence of which can be deduced from the nucleotide sequence.

[0280] In a preferred example of the present invention, the wild-type Cas protein is CasY7, and the sequence is as shown in SEQ ID NO.1.

[0281] Cas protein variants and their encoding nucleic acids

[0282] As used herein, the terms "Cas protein variant", "variant of the present invention", "gene editing mutant protein of the present invention", and "mutant protein" can be used interchangeably, and all refer to a non-naturally occurring mutant Cas protein, and the mutant protein has a mutation at one or more core amino acid sites related to cleavage activity corresponding to SEQ ID NO.1 of the wild-type Cas protein selected from the following group:

[0283] V7, T11, S12, Y165, S166, S196, K197, S198, A199, S240, S241, Q242, E243, I244, N276, Y282, D283, A285, I304, L307, Y308, S309, R315, E316, T317, I318, I319, V348, I349, E350, P351, G416, I417, E418, F419, D420, E648, L650, A651, Y652, T681, N682, E683, S684, E748, G749, S751, K865, P866, Y867, N868, I872, M957, F958, Q960, W961.

[0284] The term "core amino acids" refers to sequences that are based on wild-type Cas proteins and have at least 80% homology with wild-type Cas proteins, such as 84%, 85%, 90%, 92%, 95%, 98% or 99%, and at the corresponding positions are the specific amino acids described herein. For example, for gene editing proteins based on wild-type, the core amino acids are:

[0285] V7, T11, S12, Y165, S166, S196, K197, S198, A199, S240, S241, Q242, E243, I244, N276, Y282, D283, A285, I304, L307, Y308, S309, R315, E316, T317, I318, I319, V348, I349, E350, P351, G416, I417, E418, F419, D420, E648, L650, A651, Y652, T681, N682, E683, S684, E748, G749, S751, K865, P866, Y867, N868, I872, M957, F958, Q960, W961.

[0286] Moreover, the mutant proteins obtained by mutating the above core amino acids have higher cleavage activity relative to the wild-type Cas protein (SEQ ID NO. 1); or, the cleavage activity of the mutant proteins is equivalent to that of the wild-type Cas protein, or the binding of the CRISPR composition containing the Cas protein variant to the binding site is enhanced or the editing preference is changed.

[0287] Preferably, in the present invention, the following mutations are made to the core amino acids of the present invention:

[0288] V7I, T11V, S12M, Y165W, S166R, S196N, K197F, S198N, A199T, S240L, S241A, Q242S, E243H, I244L, N276S, Y282F, D283N, D283Q, D283S, A285T, A285P, I304A, L307R, Y308A, Y308F, S309A, R315A, R315S, E316R, E316Q, T317K, T317R, I318K, I318L, I319Q, I319W, V348I, I349V, E350L, P351T, G416S, G416R, I417H, I417V, E418Q, F419T, F419V, D420T, D420V, E648L, E648R, E648S, E648T, L650I, A651L, A651G, A651P, A651S, Y652F, T681P, N682P, E683Q, S684C, E748A, G749S, S751G, K865S, P866S, Y867H, N868F, I872L, M957V, F958L, Q960V, W961V.

[0289] It should be understood that the amino acid numbering in the mutant proteins of the present invention is based on the wild-type Cas protein. When the sequence homology of a specific mutant protein with the wild-type Cas protein reaches 80% or more, the amino acid numbering of the mutant protein may be misaligned relative to the amino acid numbering of the wild-type Cas protein, such as being misaligned 1-100 positions towards the N-terminus or C-terminus of the amino acid. By using conventional sequence alignment techniques in the art, those skilled in the art can generally understand that such misalignment is within a reasonable range, and mutant proteins with a homology of 80% (such as 90%, 95%, 98%) to the wild-type Cas protein (SEQ ID NO.1) that have the same or similar gene editing activities (such as cleavage activities) relative to the wild-type Cas protein and are equivalent or higher should not be excluded from the scope of the mutant proteins of the present invention due to the misalignment of amino acid numbering.

[0290] The mutant proteins of the present invention are synthetic proteins or recombinant proteins, that is, they can be products of chemical synthesis or produced using recombinant techniques from prokaryotic or eukaryotic hosts (such as bacteria, yeast, plants). Depending on the host used in the recombinant production protocol, the mutant proteins of the present invention can be glycosylated or non-glycosylated. The mutant proteins of the present invention may also include or not include the starting methionine residue.

[0291] The present invention also includes fragments, derivatives, and analogs of the mutant protein. As used herein, the terms "fragment", "derivative", and "analog" refer to proteins that substantially retain the same biological function or activity as the mutant protein.

[0292] The mutant protein fragments, derivatives, or analogs of the present invention can be (i) mutant proteins in which one or more conservative or non-conservative amino acid residues (preferably conservative amino acid residues) are substituted, and such substituted amino acid residues can or cannot be encoded by the genetic code, or (ii) mutant proteins having a substituent group in one or more amino acid residues, or (iii) mutant proteins formed by fusion of the mature mutant protein with another compound (such as a compound that prolongs the half-life of the mutant protein, for example, polyethylene glycol), or (iv) mutant proteins formed by fusion of an additional amino acid sequence to this mutant protein sequence (such as a leader sequence or a secretion sequence or a sequence used to purify this mutant protein or a proprotein sequence, or a fusion protein formed with an antigen IgG fragment). According to the teachings herein, these fragments, derivatives, and analogs are within the scope well known to those skilled in the art.

[0293] In certain embodiments, other selected groups of amino acids that are considered to be conservative substitutions of each other.

[0294] Table A

[0295]

[0296]

[0297] In certain embodiments, other selected groups of amino acids that are considered to be conservative substitutions of each other (see, for example, Creighton, Proteins (1984)):

[0298] Table B

[0299]

[0300] In certain embodiments, other selected groups of amino acids that are considered to be conservative substitutions of each other:

[0301] Table C

[0302] Group 1 Ala (A), Ser (S), Thr (T) Group 2 Asp (D), Glu (E) Group 3 Asn (N), Gln (Q) Group 4 Arg (R), Lys (K) Group 5 Ile (I), Leu (L), Met (M) Group 6 Phe (F), Tyr (Y), Trp (W)

[0303] The active mutant protein of the present invention has comparable or higher gene editing activity (such as cleavage activity) relative to the wild-type Cas protein (SEQ ID NO.1).

[0304] In addition, the mutant protein of the present invention can also be modified. The modified (usually without changing the primary structure) forms include: chemically derived forms of the mutant protein in vivo or in vitro, such as acetylation or carboxylation. Modification also includes glycosylation, such as those mutant proteins that are glycosylated during the synthesis and processing of the mutant protein or in further processing steps. Such modification can be accomplished by exposing the mutant protein to enzymes that perform glycosylation (such as mammalian glycosylating enzymes or deglycosylating enzymes). Modified forms also include sequences having phosphorylated amino acid residues (such as phosphotyrosine, phosphoserine, phosphothreonine). Also included are mutant proteins that are modified to enhance their proteolytic resistance or optimize their solubility properties.

[0305] The term "polynucleotide encoding a mutant protein" can be a polynucleotide that includes the polynucleotide encoding the mutant protein of the present invention, or can also be a polynucleotide that further includes additional coding and / or non-coding sequences.

[0306] The present invention also relates to variants of the above polynucleotides, which encode polypeptides having the same amino acid sequence as the present invention or fragments, analogs and derivatives of the mutant protein. These nucleotide variants include substitution variants, deletion variants and insertion variants. As is known in the art, allelic variants are alternative forms of a polynucleotide, which may be substitutions, deletions or insertions of one or more nucleotides, but do not substantially change the function of the mutant protein encoded thereby.

[0307] The present invention also relates to polynucleotides that hybridize with the above sequences and have at least 50%, preferably at least 70%, more preferably at least 80% identity between the two sequences. The present invention particularly relates to polynucleotides that can hybridize with the polynucleotides described in the present invention under stringent conditions (or high stringency conditions). In the present invention, "stringent conditions" refer to: (1) hybridization and washing at lower ionic strength and higher temperature, such as 0.2×SSC, 0.1% SDS, 60°C; or (2) adding a denaturing agent during hybridization, such as 50% (v / v) formamide, 0.1% calf serum / 0.1% Ficoll, 42°C, etc.; or (3) hybridization occurs only when the identity between the two sequences is at least 90% or more, preferably 95% or more.

[0308] The mutant protein and polynucleotide of the present invention are preferably provided in isolated form, and more preferably, are purified to homogeneity.

[0309] The full-length sequence of the polynucleotide of the present invention can generally be obtained by PCR amplification, recombination, or artificial synthesis. For PCR amplification, primers can be designed according to the nucleotide sequences disclosed in the present invention, especially the open reading frame sequences, and a commercially available cDNA library or a cDNA library prepared by conventional methods known to those skilled in the art can be used as a template for amplification to obtain the relevant sequence. When the sequence is relatively long, it is often necessary to perform PCR amplification two or more times, and then splice the amplified fragments together in the correct order.

[0310] Once the relevant sequence is obtained, the relevant sequence can be obtained in large quantities by recombination. This is usually to clone it into a vector, then transfer it into cells, and then isolate the relevant sequence from the proliferated host cells by conventional methods.

[0311] In addition, the relevant sequence can also be synthesized by artificial synthesis, especially when the fragment length is relatively short. Usually, a very long fragment can be obtained by first synthesizing multiple small fragments and then ligating them.

[0312] Currently, it is already possible to obtain the DNA sequence encoding the protein (or its fragment, or its derivative) of the present invention entirely by chemical synthesis. Then this DNA sequence can be introduced into various existing DNA molecules (such as vectors) and cells known in the art. In addition, mutations can also be introduced into the protein sequence of the present invention by chemical synthesis. The method of using PCR technology to amplify DNA / RNA is preferably used to obtain the polynucleotide of the present invention. Especially when it is difficult to obtain full-length cDNA from the library, the RACE method (rapid amplification of cDNA ends) can be preferably used. The primers for PCR can be appropriately selected according to the sequence information of the present invention disclosed herein and can be synthesized by conventional methods. The amplified DNA / RNA fragments can be separated and purified by conventional methods such as gel electrophoresis.

[0313] Fusion protein

[0314] In one aspect, the present invention provides a fusion protein, which comprises the Cas protein or its variant as described in any one of the foregoing and one or more functional domains.

[0315] In one embodiment, the functional domain includes one or more of a localization signal, a reporter protein, a targeting portion of the Cas protein or its variant, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage-active polypeptide, a ligase;

[0316] In one embodiment, "methylase", by way of example, such as HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3, ZMET2, CMT1, CMT2, etc.

[0317] "Demethylase" refers to an enzyme that removes a methyl (CH3-) group from nucleic acids, proteins (e.g., histones), and other molecules. Demethylases are important in epigenetic modification mechanisms. Demethylase proteins alter the transcriptional regulation of the genome by controlling the methylation levels that occur on DNA and histones, and thereby regulate the chromatin state at specific loci within an organism, such as TET1 (ten-eleven translocation 1), ten-eleven translocation (TET) dioxygenase 1 (TET1CD), DME, DML1, DML2, ROS1, etc.

[0318] In certain embodiments, the transcriptional release factor, by way of example, such as eukaryotic release factor 1 (ERF1) activity, eukaryotic release factor 3 (ERF3).

[0319] In one embodiment, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0320] In one embodiment, the localization signal includes a nuclear localization signal and / or a nuclear export signal;

[0321] Preferably, the nuclear export signal includes human protein tyrosine kinase 2;

[0322] Preferably, the reporter protein includes one or more of glutathione-S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, or a self-fluorescent protein;

[0323] Preferably, the self-fluorescent protein includes one or more of green fluorescent protein, HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, or blue fluorescent protein;

[0324] Preferably, the DNA binding domain includes one or more of a methylation binding protein, LexA DBD, or Gal4 DBD;

[0325] Preferably, the epitope tag includes one or more of a histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, or thioredoxin tag;

[0326] Preferably, the transcriptional activation domain comprises VP64 and / or VPR;

[0327] Preferably, the transcriptional repression domain comprises KRAB and / or SID;

[0328] Preferably, the nuclease comprises FokI;

[0329] Preferably, the deamination domain comprises one or more of ADAR1, ADAR2, APOBEC, AID or TAD;

[0330] In one embodiment, the deamination domain is an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with the amino acid sequence set forth in any one of SEQ ID NO: 31-32, 70.

[0331] In one embodiment, the functional domain is the full length or a functional fragment of TadA8e.

[0332] Preferably, the cleavage active polypeptide comprises a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity or a polypeptide having double-stranded DNA cleavage activity;

[0333] Preferably, the ligase comprises DNA ligase and / or RNA ligase.

[0334] Polynucleotide

[0335] In one aspect, the present invention provides a polynucleotide which is a polynucleotide sequence encoding the Cas protein or its variant, or a polynucleotide sequence encoding the aforementioned fusion protein.

[0336] In one embodiment, the polynucleotide is a DNA molecule codon-optimized according to the codon preference of the host cell;

[0337] The optimizations described in the present disclosure may require mutations to the nucleotide sequence encoding the protein (e.g., the Cas protein or its variant of the present disclosure) to mimic the codon preferences of the expected host organism or cell when simultaneously encoding the same protein. Thus, the codons may be altered, but the encoded protein remains the same. For example, if the expected target cell is a human cell, a nucleotide sequence encoding the protein optimized with human codons may be used. As another non-limiting example, if the expected host cell is an animal cell (e.g., a mouse cell, an insect cell), a nucleotide sequence encoding the protein optimized with the codons of that animal may be generated. As another non-limiting example, if the expected host cell is a plant cell, a nucleotide sequence encoding the protein optimized with plant codons may be generated.

[0338] Lists of codon choices are readily available, for example, the "Codon Usage Database" available at www.kazusa.or.jp / codon. In some cases, the nucleic acids of the present disclosure comprise a nucleotide sequence encoding CasY7 or its variant, or a fusion protein thereof, which nucleotide sequence is codon-optimized for expression in eukaryotic cells. In some cases, the nucleic acids of the present disclosure comprise a nucleotide sequence encoding CasY7 or its variant, or a fusion protein thereof, which nucleotide sequence is codon-optimized for expression in animal cells. In some cases, the nucleic acids of the present disclosure comprise a nucleotide sequence encoding CasY7 or its variant, or a fusion protein thereof, which nucleotide sequence is codon-optimized for expression in fungal cells. In some cases, the nucleic acids of the present disclosure comprise a nucleotide sequence encoding CasY7 or its variant, or a fusion protein thereof, which nucleotide sequence is codon-optimized for expression in plant cells.

[0339] In one embodiment, the host cell comprises a prokaryotic cell or a eukaryotic cell;

[0340] CRISPR system

[0341] The term "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" is used interchangeably and has the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of directing the activity of the Cas genes.

[0342] CRISPR-Cas composition

[0343] In one aspect, the present invention also provides a CRISPR-Cas composition, the composition comprising:

[0344] (1) Protein component: the aforementioned Cas protein or its variant, or the aforementioned fusion protein; or a nucleic acid molecule encoding the Cas protein or its variant or the fusion protein;

[0345] (2) RNA component: guide RNA, or one or more nucleic acids encoding the guide RNA, or precursor RNA of the guide RNA, or a nucleic acid encoding the precursor RNA of the guide RNA;

[0346] The protein component and the nucleic acid component bind to each other to form a complex.

[0347] In one embodiment, the composition is an activated CRISPR complex, and the activated CRISPR complex further comprises: a target sequence of a target nucleic acid bound to the guide RNA.

[0348] In one embodiment, the CRISPR-Cas composition comprises one or more vectors, and the one or more vectors comprise:

[0349] (1) A first regulatory element, which is operably linked to a nucleotide sequence encoding the Cas protein or its variant or a nucleotide sequence encoding the fusion protein; and

[0350] (2) A second regulatory element, which is operably linked to a nucleotide sequence encoding the guide RNA, and the guide RNA comprises:

[0351] (a) A spacer sequence capable of hybridizing with a target sequence of a target nucleic acid, and

[0352] (b) A direct repeat (DR) sequence linked to the spacer sequence and capable of guiding the Cas protein or its variant to bind to the guide RNA to form a CRISPR-Cas complex targeting the target sequence;

[0353] Wherein the first regulatory element and the second regulatory element are located on the same or different vectors of the CRISPR-Cas vector system.

[0354] In one embodiment, the first regulatory element or the second regulatory element comprises a promoter, and the promoter comprises one or more of an inducible promoter, a constitutive promoter or a tissue-specific promoter;

[0355] In one embodiment, the promoter comprises one or more of T7, SP6, T3, CMV, EF1a, SV40, PGK1, human β-actin, CAG, U6, H1, T7, T7lac, araBAD, trp, lac or Ptac;

[0356] In one embodiment, the first regulatory element and the second regulatory element are located on the same or different vectors.

[0357] In one embodiment, the vector comprises a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral vector, a herpes simplex vector or a phagemid vector;

[0358] In one embodiment, the vector comprises a plasmid vector.

[0359] In one embodiment, the target nucleic acid comprises DNA derived from a eukaryote or DNA derived from a prokaryote;

[0360] In one embodiment, the eukaryote comprises an animal or a plant;

[0361] In one embodiment, the target nucleic acid comprises non-human mammalian DNA, human DNA, insect DNA, avian DNA, reptilian DNA, amphibian DNA, rodent DNA, fish DNA, worm DNA, nematode DNA or yeast DNA;

[0362] In one embodiment, the non-human mammalian DNA comprises non-human primate DNA.

[0363] CRISPR / Cas complex

[0364] The term "CRISPR / Cas complex" refers to a complex formed by the binding of a guide RNA, gRNA (guide RNA) or mature crRNA (or guide RNA) to a Cas protein, which comprises a direct repeat sequence hybridized to a guide sequence of a target sequence and bound to the Cas protein, and the complex is capable of recognizing and cleaving a target nucleotide that can hybridize to the guide RNA or mature crRNA.

[0365] Guide RNA (gRNA, guide RNA)

[0366] The terms "guide RNA (gRNA)", "mature crRNA", "crRNA", "guide sequence", "guide RNA" are used interchangeably and have the meaning commonly understood by those skilled in the art. Generally, a guide RNA can comprise a direct repeat (DR) sequence and a spacer sequence, or consist essentially of or consist of a direct repeat (DR) sequence and a spacer sequence.

[0367] In some cases, the spacer sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize with the target sequence and direct the specific binding of the CRISPR-Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the spacer sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. A guide sequence comprises a sequence (such as a direct repeat (DR) sequence) that has sufficient complementarity to a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct the sequence-specific binding of the complex to the target nucleic acid sequence.

[0368] It is known in the art that, based on sufficient complementarity functioning, complete complementarity is not required, and thus, if desired, the cleavage efficiency can be regulated by introducing mismatches (e.g., one or more mismatches between the spacer sequence and the target nucleic acid, such as a mismatch of 1 or 2 nucleotides (including the position of the mismatch along the spacer sequence / target sequence)). For example, if a cleavage rate of less than 100% of the target (e.g., in a cell population) is desired, then 1 or 2 mismatches between the spacer sequence and the target sequence can be introduced into the spacer sequence.

[0369] In one aspect, the present invention provides a guide RNA comprising a direct repeat (DR) sequence capable of binding the Cas protein or its variant and a spacer sequence capable of targeting a target sequence.

[0370] In one embodiment, the direct repeat (DR) of the direct repeat sequence comprises the sequence shown in SEQ ID NO.5 or 71.

[0371] In one embodiment, the 3'-end of the direct repeat sequence comprises a stem-loop structure, further comprising a stem formed by hybridization of a first stem nucleotide chain and a second stem nucleotide chain with each other, and a loop nucleotide chain forming the loop of the stem-loop structure.

[0372] In one embodiment, the DR sequence has one or more nucleotide changes including nucleotide additions, insertions, deletions, and substitutions as compared to the DR sequence shown in SEQ ID NO: 5 or 7.

[0373] In one embodiment, the DR sequence contains a stem-loop structure near the 3'-end of the DR sequence, and the stem-loop structure contains 5'-X1X2X3X4X5NNNNNNNX6X7X8X9X 10 -3'; X1, X2, X3, X4, X5, X6, X7, X8, X9, X 10 is any base including A, T, C, or G, and N is any base including A, T, C, or G; wherein X1, X2, X3, X4, X5 and X6, X7, X8, X9, X 10 can hybridize with each other to form a stem and such that NNNNNNN forms a loop; more preferably, wherein the DR sequence contains a stem-loop structure near the 3'-end of the DR sequence as shown below:

[0374] 5'-CCGTCNNNNNNNGACGG-3' (SEQ ID NO. 80); wherein, N is any base including A, T, C, or G.

[0375] In one embodiment, the direct repeat sequence includes a nucleotide sequence having at least 80% identity with the nucleotide sequence set forth in SEQ ID NO. 5 or 71. In one embodiment, the direct repeat sequence includes a nucleotide sequence having at least 85% or more, more preferably 90% or more, further preferably 95% or more identity with the nucleotide sequence set forth in SEQ ID NO. 5 or 71.

[0376] In one embodiment, the direct repeat sequence includes the nucleotide sequence set forth in SEQ ID NO. 5 or 71.

[0377] In one embodiment, more than 80% of the spacer sequence is complementary to the target nucleic acid;

[0378] In one embodiment, more than 90%, more preferably more than 95%, further preferably more than 99%, and even more preferably 100% of the spacer sequence is complementary to the target nucleic acid;

[0379] In one embodiment, the length of the spacer sequence is 18-41 nt, such as 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 nt, more preferably 18 to 27 nucleotides, more preferably 18 to 24 nucleotides, and most preferably 18 to 22 nucleotides.

[0380] In one embodiment, the length of the spacer sequence is 20 nt.

[0381] Target nucleic acid

[0382] In the present invention, the target nucleic acid is used interchangeably with the target sequence, the target nucleic acid sequence, or the target nucleic acid molecule, and refers to a specific nucleic acid that contains a nucleic acid sequence that is fully or partially complementary to the spacer sequence in the guide RNA. The "target sequence" refers to a polynucleotide targeted by the spacer sequence in the guide RNA, such as a sequence that is complementary to the spacer sequence, wherein hybridization between the target sequence and the spacer sequence will promote the formation of a CRISPR-Cas complex (including the Cas protein and the guide RNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR-Cas complex. In some embodiments, the target nucleic acid contains non-coding regions (e.g., promoters or terminators). In some embodiments, the target nucleic acid is single-stranded or double-stranded.

[0383] The target sequence can contain any polynucleotide, such as DNA. In certain cases, the target sequence is located intracellularly or extracellularly. In certain cases, the target sequence is located within the nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts) of a cell.

[0384] The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In certain cases, the target sequence should be related to the protospacer adjacent motif (PAM).

[0385] Donor template

[0386] In the present invention, the donor template nucleic acid or the donor template is used interchangeably, and refers to a nucleic acid molecule that one or more cellular proteins can use to alter the structure of the target nucleic acid after the Cas protein or its variant described herein has altered the target nucleic acid.

[0387] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid or a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some instances, the donor template nucleic acid is an exogenous nucleic acid molecule. In some instances, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, gene recombination can be achieved using the donor template, and the recombination is homologous recombination.

[0388] Cleavage

[0389] Cleavage refers to a DNA break in the target nucleic acid generated by the Cas proteins or variants thereof described herein. In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break.

[0390] In the present invention, the meanings of cleaving the target nucleic acid or modifying the target nucleic acid may overlap. Modifying the target nucleic acid includes not only modification of a single nucleotide, but also insertion or deletion of nucleic acid fragments.

[0391] Reporter nucleic acid

[0392] A reporter nucleic acid refers to a molecule that can be cleaved by the activated CRISPR system proteins described herein or otherwise inactivated. The reporter nucleic acid contains a nucleic acid element that can be cleaved by a CRISPR protein (e.g., using a single-stranded non-targeting nucleic acid molecule with different reporter groups or labeling molecules at both ends). Cleavage of the nucleic acid element generates a detectable signal. Before cleavage, or when the reporter nucleic acid is in an "active" state, the reporter nucleic acid prevents the generation or detection of a positive detectable signal. It will be understood that in certain exemplary embodiments, a minimal background signal may be generated in the presence of an active reporter nucleic acid. The positive detectable signal can be any signal that can be detected using optical, fluorescence, chemiluminescence, electrochemistry, or other detection methods known in the art. For example, in certain embodiments, when a reporter nucleic acid is present, a first signal (i.e., a negative detectable signal) can be detected, and then it is converted to a second signal (e.g., a positive detectable signal) after detection of the target molecule and cleavage or inactivation by the activated CRISPR protein. The reporter nucleic acid can be a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.

[0393] The detection method described in the present invention can be used for quantitative detection of the target nucleic acid to be detected. The quantitative detection index can be quantified according to the signal strength of the reporter group, such as according to the luminescence intensity of the fluorescent group, or according to the width of the color development band, etc.

[0394] Functional domain

[0395] As used herein, the functional domain is taken in its broadest sense and includes a protein such as an enzyme or factor itself or a fragment / domain thereof having a specific function. A Cas protein (such as a dCas protein) is linked / associated with one or more functional domains selected from a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage active polypeptide, a ligase, or more than one of these. When more than one functional domain is included, the functional domains may be the same or different.

[0396] Deamination domain

[0397] In the present invention, the deamination domain includes a catalytic domain of a deaminase (such as adenosine deaminase or cytidine deaminase). As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts adenine (or the adenine moiety of a molecule) to inosine (or the inosine moiety of a molecule).

[0398] In some embodiments, the adenine-containing molecule is adenosine (A), and the inosine-containing molecule is inosine (I). The adenine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0399] Adenosine deaminases that can be used in conjunction with the Cas proteins or variants thereof, or Cas proteins or variants substantially lacking catalytic activity of the present invention include, but are not limited to, members of the enzyme family called adenosine deaminases acting on RNA (ADAR), members of the enzyme family called adenosine deaminases acting on tRNA (ADAT), and other family members containing an adenosine deaminase domain (ADAD). According to the present disclosure, adenosine deaminases are capable of targeting adenine in RNA / DNA and RNA duplexes. In certain embodiments, the adenosine deaminase has been modified to increase its ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.

[0400] In some embodiments, the adenosine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squids, fish, flies, and worms. In some embodiments, the adenosine deaminase is human, squid, or Drosophila adenosine deaminase. In some embodiments, the adenosine deaminase is human ADAR, including hADAR1, hADAR2, hADAR3. In some embodiments, the adenosine deaminase is the Caenorhabditis elegans ADAR protein, including ADR-1 and ADR-2. In some embodiments, the adenosine deaminase is the Drosophila ADAR protein, including dAdar. In some embodiments, the adenosine deaminase is the squid (Loligo pealeii) ADAR protein, including sqADAR2a and sqADAR2b. In some embodiments, the adenosine deaminase is human ADAT protein. In some embodiments, the adenosine deaminase is Drosophila ADAT protein. In some embodiments, the adenosine deaminase is human ADAD protein, including TENR (hADAD1) and TENRL (hADAD2).

[0401] In some embodiments, the adenosine deaminase comprises the wild-type amino acid sequence of hADAR2-D. In some embodiments, the adenosine deaminase comprises one or more mutations in the hADAR2-D sequence such that the editing efficiency and / or substrate editing preference of hADAR2-D is altered according to specific needs.

[0402] In some embodiments, the adenosine deaminase is deaminase 005V1 (SEQ ID NO.31) or deaminase 004V1 (SEQ ID NO.70). In some embodiments, the adenosine deaminase is a variant of deaminase 005V1 or deaminase 004V1, and the variant comprises one or more mutations in the deaminase 005V1 or deaminase 004V1 sequence such that the editing efficiency and / or substrate editing preference of deaminase 005V1 or deaminase 004V1 is altered according to specific needs.

[0403] In one embodiment, the adenosine deaminase domain comprises an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with the amino acid sequence set forth in any of SEQ ID NO:31-32, 70.

[0404] In some embodiments, the adenosine deaminase is the full-length or functional fragment of TadA8e, such as TadA8e whose amino acid sequence comprises SEQ ID NO:79.

[0405] In some embodiments, the deaminase is a cytidine deaminase. The term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts cytosine (or the cytosine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0406] Cytidine deaminases include, but are not limited to, members of the enzyme family known as the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases, activation-induced deaminase (AID), or cytidine deaminase 1 (CDA1). In certain embodiments, the APOBEC family of deaminases is included.

[0407] In some embodiments, the cytidine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squid, fish, flies, and worms. In some embodiments, the cytidine deaminase is a human, primate, bovine, canine, rat, or murine cytidine deaminase.

[0408] In some embodiments, the cytidine deaminase is a human APOBEC, including hAPOBEC1 or hAPOBEC3. In some embodiments, the cytidine deaminase is a human AID.

[0409] In some embodiments, the cytidine deaminase comprises the wild-type amino acid sequence of a cytosine deaminase. In some embodiments, the cytidine deaminase contains one or more mutations in the cytosine deaminase sequence such that the editing efficiency and / or substrate editing preference of the cytosine deaminase is altered according to specific needs.

[0410] Identity

[0411] "Identity" is used to refer to the sequence match between two polypeptides or two nucleic acids. "Identity" represents the percentage of the number of identical residues between the polypeptide or nucleic acid sequences, and the calculation of the total number of residues is determined based on the type of mutation. Types of mutations include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitution / replacement of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence.

[0412] Taking a polypeptide sequence as an example, if the mutation type is one or more of the following: substitution / replacement of one or more amino acids / nucleotides, insertion within the sequence, and deletion within the sequence, the total number of residues is calculated based on the larger of the two molecules being compared. If the mutation type also includes insertion (extension) at either or both ends of the sequence or deletion (truncation) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (e.g., the number inserted or deleted at both ends is less than 20) is not included in the total number of residues. When calculating the percentage identity, the sequences being compared are aligned in a way that produces the maximum match between the sequences, and gaps in the alignment (if any) are resolved by a specific algorithm. The calculation of nucleotide identity is the same.

[0413] vector

[0414] A vector is a nucleic acid molecule capable of transporting another nucleic acid molecule linked thereto.

[0415] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, no free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and other diverse polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection, enabling the genetic material element it carries to be expressed in the host cell. A vector can be introduced into a host cell and thereby produce transcripts, proteins, or peptides, including those derived from protein variants, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts such as nucleic acid transcripts, proteins, or enzymes). A vector can contain various elements controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. A vector can also contain an origin of replication.

[0416] Vectors include plasmids and viral vectors. A plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. A viral vector, in which virus-derived DNA or RNA sequences are present in the vector used for packaging the virus, and the virus includes, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. A viral vector also contains the polynucleotide carried by the virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.

[0417] Other vectors (e.g., non-integrating mammalian vectors) integrate into the genome of a host cell upon introduction into the host cell and are thereby replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to as "expression vectors".

[0418] In some embodiments, a vector (e.g., a viral vector or a non-viral vector, such as a lentiviral vector or a plasmid) can be delivered to a target tissue by, for example, intramuscular injection, intravenous administration, percutaneous administration, intranasal administration, oral administration, or mucosal administration. The above delivery can be carried out via a single dose or multiple doses. Those skilled in the art will understand that the actual dose to be delivered herein can vary to a large extent depending on a variety of factors, including but not limited to vector selection, target cells, organisms, tissues, the general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the mode of administration, and the type of transformation / modification sought.

[0419] Regulatory elements

[0420] As used herein, "regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals, poly-U sequences), the detailed description of which can be found in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or particular cell types (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type specific.

[0421] The term "promoter" refers to a non-coding nucleotide sequence located upstream of a gene that can initiate the expression of downstream genes. A constitutive promoter is such a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, will result in the production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, such as in response to a chemical compound (chemical inducer), or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulatable promoters include, for example, promoters induced or regulated by light, heat, stress, waterlogging or drought, salt stress, osmotic stress, plant hormones, wounds, or chemicals (such as ethanol, abscisic acid (ABA), jasmonates, salicylic acid, or safeners).

[0422] host cell

[0423] As used herein, "host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism (e.g., a cell line) cultured as a single cell entity, which cell is used as a recipient of a nucleic acid (e.g., an expression vector), and includes the progeny of the original cell that has been genetically modified with the nucleic acid.

[0424] It should be understood that the progeny of a single cell may, due to natural, accidental, or deliberate mutations, not necessarily have exactly the same morphology or genome as the original parental cell. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.

[0425] Those skilled in the art will understand that the design of an expression vector can depend on factors such as the choice of host cell to be transformed, the desired level of expression, etc.

[0426] In another aspect, the present invention also provides a host cell or its progeny, which host cell comprises the foregoing Cas protein or its variant, or the foregoing fusion protein, or the foregoing polynucleotide, or the foregoing vector system, or the foregoing CRISPR-Cas system, or the foregoing composition.

[0427] In one embodiment, the host cell includes a non-human mammal, a human, an insect, a bird, a reptile, an amphibian, a rodent, a fish, a worm, a nematode, or a yeast cell.

[0428] In one aspect, the present invention also provides a multicellular organism, which multicellular organism comprises the foregoing cell or its progeny.

[0429] In one embodiment, the multicellular organism is an animal model or a plant model for a related disease.

[0430] NLS

[0431] NLS refers to "nuclear localization sequence" or "nuclear localization signal", which refers to the amino acid sequence that promotes the entry of a protein into the nucleus. Nuclear localization sequences are known in the art (for example, as described in International PCT Application PCT / EP2000 / 011690 filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001), and this patent is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the following amino acid sequences: KRTADGSEFESPKKKRKV (SEQ ID NO.38), AVKRPAATKKAGQAKKKKLD (SEQ ID NO.39), KRPAATKKAGQAKKKK (SEQ ID NO.40), KKTELQTTNAENKTKKL (SEQ ID NO.41), KRGINDRNFWRGENGRKTR (SEQ ID NO.42), RKSGKIAAIVVKRPRK (SEQ ID NO.43), PKKKRKV (SEQ ID NO.44), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO.45).

[0432] Operably linked

[0433] "Operably linked" means that a target nucleotide sequence is linked to a regulatory element in a manner that allows the nucleotide sequence to be expressed (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the types of these vectors can also be selected to target specific types of cells.

[0434] Complementary

[0435] "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. The percentage of complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., if 5, 6, 7, 8, 9, or 10 out of 10 are complementary, the percentage of complementarity is 50%, 60%, 70%, 80%, 90%, and 100%). "Fully complementary" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.

[0436] The term "stringent conditions" in relation to hybridization refers to conditions under which a nucleic acid that is complementary to a target sequence hybridizes predominantly to the target sequence and essentially not to non-target sequences. Stringent conditions are usually sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence hybridizes specifically to its target sequence.

[0437] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding of the bases between these nucleotide residues. The complex can comprise two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. The hybridization reaction can be a step in a broader process (such as the initiation of PCR, or the cleavage of a polynucleotide by an enzyme). A sequence that can hybridize to a given sequence is called the "complement" of the given sequence.

[0438] Hybridization of the target sequence with the gRNA means that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or it represents that at least 12, 15, 16, 17, 18, 19, 20, or more bases of the nucleic acid sequences of the target sequence and the gRNA can be complementary and hybridize to form a complex.

[0439] Expression

[0440] Nucleic acid expression includes one or more of generating an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5′ cap formation, and / or 3′ end processing), translation of the RNA into a polypeptide or protein, or post-translational modification of the polypeptide or protein.

[0441] Delivery

[0442] "Delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as combinations of DNA / RNA or RNA / RNA or protein-RNA. For example, the Cas protein or its variant can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.

[0443] In one aspect, the present invention also provides a delivery system, which includes the Cas protein or its variant or the fusion protein, or the polynucleotide, or the CRISPR-Cas composition described above.

[0444] In one embodiment, the delivery system further includes a delivery vehicle, and the delivery vehicle includes nanoparticles, liposomes, exosomes, microbubbles, gene guns, or electroporation devices.

[0445] In addition, when the delivery target is a plant cell, a delivery method such as using a cell-penetrating peptide (CPP) is also adopted. For example, in a specific embodiment, the Cas protein or its variant and / or at least one guide RNA are conjugated with one or more CPPs, so as to effectively transport the CPP conjugated with the Cas protein or its variant and / or guide RNA into the plant cell (e.g., into the protoplast). The CPP is a short peptide with less than 35 amino acids, which is derived from a protein or a chimeric sequence and can transport biomolecules across the cell membrane in a non-receptor-dependent manner. The CPP can be a cationic peptide, a peptide with a hydrophobic sequence, an amphiphilic peptide, a peptide rich in proline and antimicrobial sequences, and a chimeric or bipartite peptide. The CPP can penetrate biological membranes, and thus trigger the movement of different biomolecules across the cell membrane into the cytoplasm, and can improve their intracellular pathways, and thus promote the interaction between the biomolecules and the target.

[0446] Exemplarily, the CPPs include Tat (a nuclear transcriptional activator protein required for viral replication by human immunodeficiency virus type 1), penetratin, Kaposi fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, sweet arrow peptide, etc.

[0447] Linker

[0448] As used herein, "linker" refers to a chemical group or molecule that connects two molecules or moieties, such as two domains of a fusion protein, such as a Cas protein or its variant and a deaminase. In some linking modes, the linker is located between or flanks two groups, molecules or other moieties and covalently links the two.

[0449] In some embodiments, the linker is a linear polypeptide formed by amino acids or multiple amino acid residues linked by peptide bonds. In some embodiments, the linker is an organic molecule, group, polymer or chemical moiety. The length and type of the linker can be designed as needed. In some embodiments, the linker can be selected from synthetic amino acid sequences or naturally occurring polypeptide sequences.

[0450] Detection

[0451] In one aspect, the present invention also provides a method for targeting and editing a target nucleic acid, the method comprising contacting the target nucleic acid with any one of the foregoing CRISPR-Cas systems or compositions.

[0452] In one aspect, the present invention also provides a method for non-specifically degrading single-stranded DNA after identifying a target nucleic acid, the method comprising contacting the target nucleic acid with the foregoing CRISPR-Cas composition.

[0453] In one aspect, the present invention also provides a method for targeting and nicking the non-Spacer complementary strand of a double-stranded target nucleic acid after identifying the Spacer complementary strand of the double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with any one of the foregoing CRISPR-Cas systems or compositions.

[0454] In one aspect, the present invention also provides a method for targeting and cleaving a double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with any one of the foregoing CRISPR-Cas systems or compositions.

[0455] In one embodiment, before nicking the Spacer complementary strand of the double-stranded DNA, the non-Spacer sequence complementary strand of the double-stranded target nucleic acid is nicked.

[0456] In one aspect, the present invention also provides a method for specifically editing a double-stranded nucleic acid, the method comprising contacting the following under sufficient conditions for a sufficient amount of time

[0457] (1) The aforementioned Cas protein or its variant, or a fusion protein, another enzyme with sequence-specific nicking activity, and the guide RNA, where the guide RNA directs the Cas protein or its variant or the fusion protein to nick the opposing strand relative to the activity of the other sequence-specific nicking enzyme; and (2) the double-stranded nucleic acid; the method results in the formation of a double-strand break.

[0458] In one aspect, the present invention also provides a method for editing a double-stranded nucleic acid, the method comprising contacting the following for a sufficient amount of time under sufficient conditions:

[0459] (1) The aforementioned Cas protein or its variant, or a fusion protein, a fusion protein with a protein domain having DNA modification activity, and the RNA guide targeting the double-stranded nucleic acid; and (2) the double-stranded nucleic acid;

[0460] The Cas protein or its variant of the fusion protein is modified to nick the non-target strand of the double-stranded nucleic acid.

[0461] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at different sites, resulting in staggered cleavage.

[0462] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at the same site, resulting in a blunt double-strand break.

[0463] In one aspect, the present invention also provides a method for targeting and cleaving a single-stranded target nucleic acid, the method comprising contacting the target nucleic acid with the CRISPR-Cas composition according to any one of the preceding claims.

[0464] In one aspect, the present invention also provides a method for inducing a change in cell state, the method comprising contacting the aforementioned CRISPR-Cas composition with the target nucleic acid in a cell.

[0465] In one embodiment, the cell state includes apoptosis or dormancy;

[0466] In one embodiment, the cell includes a eukaryotic cell or a prokaryotic cell;

[0467] In one embodiment, the cell includes a mammalian cell or a plant lesion cell;

[0468] In one embodiment, the cell includes a cancer cell;

[0469] In one embodiment, the cell includes an infectious cell or a cell infected with an infectious agent;

[0470] In one embodiment, the cell includes a cell infected with a virus, a cell infected with a prion;

[0471] In one embodiment, the cell includes a fungal cell, a protozoan or a parasite cell.

[0472] In one aspect, the present invention also provides a method for detecting a target nucleic acid in a sample, the method comprising contacting the sample with the foregoing Cas protein or its variant, a guide RNA, and a non-target sequence; detecting a detectable signal generated by cleavage of the non-target sequence by the Cas protein or its variant, thereby detecting the target nucleic acid; the non-target sequence not hybridizing with the guide RNA.

[0473] Kit

[0474] In one aspect, the present invention provides a kit, the kit comprising the foregoing Cas protein or its variant, the foregoing fusion protein, the foregoing polynucleotide, the foregoing CRISPR-Cas composition, the use of the foregoing host cell in preparing the kit, and the components of the kit being in the same or different containers.

[0475] In one aspect, the present invention also provides a container containing the foregoing kit.

[0476] In one embodiment, the container includes a sterile container;

[0477] In one embodiment, the container includes a syringe.

[0478] In some embodiments, the kit further includes instructions for using the kit, such as instructions in more than one language. The kit may also contain one or more reagents for use in the process using one or more of the foregoing components. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The foregoing reagents may be provided in a form that requires addition of one or more other components before use (e.g., in concentrated or lyophilized form); the buffer may be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. The buffer may have a suitable pH value (pH), for example, it may be alkaline. In some embodiments, the pH of the buffer is between about 7 and 10.

[0479] Treatment

[0480] "Treatment" refers to treating or curing a subject's disorder, delaying the onset of symptoms of the disorder, and / or delaying the severity of the disorder. The term "subject" includes, but is not limited to, various animals, plants, and microorganisms. Animals include mammals such as bovines, equines, ovines, porcines, canines, felines, lagomorphs (e.g., rabbits), rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder caused by a disease-related gene defect). "Plant" means any differentiated multicellular organism capable of photosynthesis, including crop plants at any mature or developmental stage.

[0481] In one aspect, the present invention also provides the use of the foregoing Cas protein or its variant, the foregoing fusion protein, the foregoing polynucleotide, the foregoing CRISPR-Cas composition, and the foregoing host cell in the preparation of a medicament for treating a disorder or disease of a subject in need thereof.

[0482] In one embodiment, the use includes administering the CRISPR-Cas composition to the subject or an ex vivo cell of the subject;

[0483] In one embodiment, the spacer sequence is complementary to at least 15 nucleotides of the target nucleic acid related to the disorder or disease, and the Cas protein or its variant or the fusion protein cleaves the target nucleic acid;

[0484] In one embodiment, the disorder or disease includes cancer or an infectious disease;

[0485] In one embodiment, the cancer includes one or more of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, or urinary bladder cancer;

[0486] In one embodiment, the disorder or disease includes one or more of cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell disease, hypercholesterolemia, transthyretin amyloidosis, or β-thalassemia;

[0487] In one embodiment, the infectious agent of the infectious disease includes one or more of human immunodeficiency virus, herpes simplex virus-1, or herpes simplex virus-2.

[0488] The main advantages of the present invention include:

[0489] By mutating the wild-type Cas protein, the present invention obtains a new Cas protein variant, which has better gene editing activity compared to the wild-type Cas protein.

[0490] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions in the following embodiments are usually carried out under conventional conditions, such as the conditions described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or the conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are weight percentages and weight parts.

[0491] Unless otherwise specified, the reagents and materials in the embodiments of the present invention are all commercially available products.

[0492] Example 1. Obtaining of Cas Protein

[0493] The inventors analyzed the metagenome of uncultured samples, and through analysis such as dereplication and protein clustering, identified 1 new Cas protein. The Blast analysis results showed that the sequence identity of the Cas protein with the reported Cas proteins was relatively low, and it was named CasY7 in the present invention.

[0494] The amino acid sequence of the above CasY7 protein is shown in SEQ ID NO.1, and its nucleotide coding sequence optimized for human codons is shown in SEQ ID NO.2.

[0495] Analysis of the direct repeat (DR) sequence of the guide RNA corresponding to the above CasY7 protein showed that:

[0496] The DNA sequence encoding the direct repeat (DR) sequence of the guide RNA corresponding to the CasY7 protein is as follows:

[0497] CAAGTTGAATCCGTCTATAACTGACGG (SEQ ID NO: 5).

[0498] The inventors further analyzed the RNA secondary structure of the DR sequence in pre-crRNA using RNAfold. The analysis results are as Figure 2 shown. It was found that the PAM corresponding to CasY7 is 5'-TTN, where N is A / T / C / G. The sgRNA (also known as crRNA) sequence of the CasY7 protein consists of a spacer sequence and a direct repeat (DR) sequence.

[0499] After identification, CasY7 of the present invention belongs to the Cas12 protein family.

[0500] Example 2. Verification of the cleavage activity of Cas protein

[0501] 1. Plasmid construction

[0502] (1) Using the TTR gene as the target, a spacer1 sequence: GCATCTCCCCATTCCATGAG (SEQ ID NO.6) was designed according to the target sequence of the target gene TTR gene.

[0503] According to the DR sequences of the CasY7 protein and the LbCpf1 protein, sgRNA sequences targeting the TTR target gene were designed, as shown in the following table:

[0504]

[0505] According to the requirements of vector expression, a T7 promoter and an rrnB T2 terminator were added to the 5' end and 3' end of the respective sgRNA sequences of CasY7 and LbCpf1, resulting in the CasY7-sgRNA1 expression fragment sequence:

[0506] Among them, the single-underlined sequence part is the CasY7 DR sequence, the double-underlined sequence part is the spacer sequence, the italic sequence part is the T7 promoter, the wavy-underlined sequence part is the rrnB T2 terminator sequence, the linker sequence is between the spacer sequence and the rrnB T2 terminator sequence, the dotted sequence part is the MfeI cleavage site, the bold sequence part is the MluI cleavage site, and CACCG is the linker.

[0507] In the same way, synthesize the LbCpf1-sgRNA1 expression fragment sequence:

[0508]

[0509] The single underlined sequence is the LbCpf1DR sequence, the double underlined sequence is the spacer sequence, the italic sequence is the T7 promoter, the wavy underlined sequence is the rrnB T2 terminator sequence, the linker sequence is between the spacer sequence and the rrnB T2 terminator sequence, the dotted sequence is the MfeI restriction site, the bold sequence is the MluI restriction site, and CACCG is the linker.

[0510] To protect the integrity of the sequence, AGC was introduced at the 5' end and ATA was introduced at the 3' end as protective bases when synthesizing the CasY7-sgRNA expression fragment sequence and the LbCpfl-sgRNA expression fragment sequence.

[0511] (2) The CasY7 protein encoding nucleotide sequence (as shown in SEQ ID NO. 2) fragment was synthesized by Suzhou Hongxun Biotechnology Co., Ltd., and the synthesized CasY7 protein encoding nucleotide sequence fragment was constructed into the ABE8e plasmid (Addgene, Plasmid #138489) 466-5160 position to construct a CasY7 recombinant expression plasmid (see the plasmid map) Figure 1 of Figure 1 A).

[0512] The optimized coding nucleotide sequence of LbCpf1 (amino acid sequence such as SEQ ID NO.19) synthesized by Suzhou Hongxun Biotechnology Co., Ltd. (SEQ ID NO.20) was constructed in the same way to obtain the LbCpf1 recombinant expression plasmid (see the plasmid map for details). Figure 1 of Figure 1 B).

[0513] (3) Suzhou Hongxun Biotechnology Co., Ltd. synthesized the sgRNA expression sequence fragments as described in step (1) (CasY7-sgRNA expression fragment sequence and LbCpf1-sgRNA expression fragment sequence), and the sgRNA expression fragment sequence was double-digested (MfeI / MluI) and inserted into the CasY7 recombinant expression plasmid vector that was also double-digested (MfeI / MluI) to obtain the recombinant expression plasmid CasY7+sgRNA expression plasmid expressing CasY7 and sgRNA; the LbCpf1+sgRNA expression plasmid was constructed in the same way.

[0514] (4) Construct the Target plasmid with the targeting sequence, and the construction process is as follows:

[0515] The araC-pBAD-CCDB fragment (SEQ ID NO.18) with the TTR target sequence (SEQ ID NO.6) was synthesized by Suzhou Huaxun Biotechnology Co., Ltd. and inserted into the 1284-1300 site of the pKESK22 (Addgene, Plasmid#64857) plasmid to obtain the Target plasmid. The sequence of the Target plasmid is shown in SEQ ID NO.4, and the plasmid map is shown in Figure 3 。

[0516] 2. Preparation and transformation of Escherichia coli competent cells

[0517] Transfer the Target plasmid into DH5a competent cells, and use an inoculation loop to streak and separate and inoculate on an LB solid medium containing 50 μg / ml kanamycin sulfate. Incubate overnight at 37 °C in a biochemical incubator. The next day, pick a single colony from the plate and inoculate it into an LB liquid medium test tube containing 4 ml of 50 μg / ml kanamycin sulfate (Sangon Biotech, A100408-0100). Incubate with shaking overnight at 37 °C and 200 rpm. The next day, take 4 ml of the bacterial solution and inoculate it into a 2 L Erlenmeyer flask containing 400 ml of LB liquid medium with 50 μg / ml kanamycin sulfate, and culture it at 37 °C and 200 rpm with shaking for 2-3 hours.

[0518] When the OD600nm value of the bacterial solution reaches 0.3-0.5, take out the Erlenmeyer flask and place it on ice for 10-15 min. Pour the bacterial solution into a pre-cooled 500 ml centrifuge bottle under sterile conditions, centrifuge at 4 °C and 3000 rpm for 8 min, discard the supernatant, add about 200 ml of pre-cooled CaCl2 solution, pipette and mix well to suspend the bacteria, and place it on ice for 30 min. Then centrifuge the bacterial solution at 4 °C and 3000 rpm for 8 min, discard the supernatant, add about 8 ml of pre-cooled CaCl2 solution, resuspend the bacteria, and aliquot the resuspended bacteria into 1.5 ml EP tubes, 110 μl per tube, and store them in a -80 °C ultra-low temperature freezer for later use.

[0519] 3. Determination of in vivo editing efficiency of Escherichia coli

[0520] Transfer the CasY7+sgRNA expression plasmid and the LbCpf1+sgRNA expression plasmid into the competent cells prepared in step 2 respectively. The specific process is as follows:

[0521] (1) Take the competent cells out of -80 °C, quickly insert them into ice. After about 5 minutes, the bacterial block melts. Add the CasY7+sgRNA expression plasmid, and then gently mix by flicking the bottom of the centrifuge tube with your finger. Let it stand in ice for 25 minutes. Heat shock in a 42 °C water bath for 45 seconds, quickly put it back into ice and let it stand for 2 minutes. Add 900 μl of sterile LB medium without antibiotics to the centrifuge tube, mix well, and recover at 37 °C and 220 rpm for 60 min. Take 100 μl of the bacterial liquid respectively and spread them on LB agar plates containing 30 μg / ml carbenicillin resistance (Sangon Biotech, A100358-0001) (abbreviated as C-LB medium) and LB agar plates containing both 30 μg / ml carbenicillin resistance and 10 mM L-arabinose (Sangon Biotech, A610071-0100) (abbreviated as CL-LB medium). Invert the two LB agar plates and place them in an incubator, and culture overnight at 37 °C.

[0522] (2) Detection of in vivo editing efficiency in Escherichia coli

[0523] As Figure 3 shown, the Target plasmid carries a PBAD promoter that can be induced by L-arabinose and a CCDB gene regulated by the PBAD promoter. The CCDB gene can express the CCDB toxic protein. The CCDB toxic protein, as a DNA gyrase inhibitor, can lock the DNA gyrase and the broken double-stranded DNA complex, making the DNA gyrase unable to function, and ultimately leading to cell death.

[0524] Based on this, the inventor designed a method for detecting the in vivo editing efficiency in Escherichia coli:

[0525] Under the condition that L-arabinose exists in the medium, if the CasY7 protein or LbCpf1 protein can specifically target the target sequence (SEQ ID NO.6) of the TTR gene on the Target plasmid under the guidance of the sgRNA and play a cleavage role, then the regulatory expression pathway of the PBAD promoter for the CCDB toxic protein is cut off, and the host cell survives because it does not produce the ccdB toxic protein; conversely, if the CasY7 protein or LbCpf1 protein cannot specifically target the TTR target sequence on the Target plasmid, the host cell Escherichia coli will die due to the PBAD promoter induced by L-arabinose regulating the expression of the CCDB gene to produce the CCDB toxic protein.

[0526] Therefore, the editing efficiency of the CasY7 protein targeting and cleaving the TTR target gene in Escherichia coli can be calculated according to the ratio of the number of bacterial colonies on the CL-LB medium to the number of bacterial colonies on the C-LB medium in step (2).

[0527] The results are asFigure 4 As shown in the figure, after counting the number of Escherichia coli clones and calculating the ratio, the editing efficiency of CasY7 protein was found to be 45.5%, and the editing efficiency of LbCpf1 protein was 11.1%. The editing efficiency of CasY7 protein was significantly higher than that of LbCpf1 protein.

[0528] Example 3. Detection of editing efficiency in HEK293T cells

[0529] 1. Construction of TTR-sgRNA expression plasmid

[0530] (1) Design TTR-sgRNA sequences according to the target sequence (Target-TTR-spacer2) tagaagggatatacaaagtg (SEQ ID NO.11) of the TTR gene, and synthesize oligonucleotides (oligos):

[0531] CasY7-TTR-sgRNA2 sequence: CAAGTTGAATCCGTCTATAACTGACGG tagaagggatatacaaagtg (SEQ ID NO.12), the underlined part is the DR sequence, and the rest is the spacer sequence.

[0532] LbCpf1-TTR-sgRNA2 sequence:

[0533] TAATTTCTACTAAGTGTAGAT tagaagggatatacaaagtg (SEQ ID NO.13), the underlined part is the DR sequence, and the rest is the spacer sequence.

[0534] (2) Add the CACC sequence to the 5' end of the upstream sequence of TTR-sgRNA and the AAAA sequence to the 5' end of the downstream sequence, and synthesize oligos. The specific sequences are as follows:

[0535]

[0536] After the upstream and downstream primers of the aforementioned TTR-sgRNA are synthesized, annealing is carried out through a preset program (95°C, 5 min; 95°C - 85°C at -2°C / s; 85°C - 25°C at -0.1°C / s; maintained at 4°C). Then, the annealed product is ligated to the PHK09T vector linearized by BsmBI (NEB, #R0580L). The sequence of the PHK09T vector is shown in SEQ ID NO.3. See the plasmid map Figure 5 .

[0537] The linearization of the PHK09T vector and its ligation method with the TTR-sgRNA annealed product are as follows:

[0538] First, linearize the PHK09T vector. The linearization system is as follows:

[0539] 3 μg of PHK09T vector; 6 μL of buffer (NEB: R0539L); 2 μL of BsmBI; Make up to 60 μL with ddH2O, and digest overnight at 50 °C.

[0540] Ligation system of TTR-gRNA annealing product and linearized vector:

[0541] 1 μL of T4 ligase buffer (NEB, #M0202L), 20 ng of linearized vector, 5 μL of annealed oligo fragment, 0.5 μL of T4 ligase (NEB, #M0202L), Make up to 10 μL with ddH2O, and ligate overnight at 16 °C to obtain the CasY7-TTR-sgRNA expression plasmid and the LbCpf1-TTR-sgRNA expression plasmid.

[0542] (3) Transfer the CasY7-TTR-sgRNA expression plasmid and the LbCpf1-TTR-sgRNA expression plasmid obtained in step (2) into Escherichia coli DH5a competent cells (Vidy Biotechnology, DL1001) respectively. The specific steps are as follows:

[0543] Take out the DH5α competent cells from the -80 °C refrigerator and quickly insert them into ice. After 5 minutes, wait for the bacterial mass to melt. Add the ligation product and gently mix by flicking the bottom of the centrifuge tube. Let it stand in ice for 25 minutes. Heat shock in a 42 °C water bath for 45 seconds, quickly put it back into ice and let it stand for 2 minutes. Add 700 μl of sterile LB medium without antibiotics to the centrifuge tube, mix well, and recover at 37 °C and 200 rpm for 60 minutes. Centrifuge at 3000 rpm for one minute to collect the bacteria. Leave about 100 μl of the supernatant, gently resuspend the bacterial mass by pipetting and spread it on the LB medium with Amp antibiotic. Invert the plate and incubate it overnight in a 37 °C incubator. Pick a single colony, after confirmation by sequencing, shake the positive clone and extract the plasmid (using an endotoxin-free plasmid large extraction kit, TIANGEN: DP120-01), then measure the concentration and store it at -20 °C in the refrigerator for later use.

[0544] 2. Detection of editing efficiency at the cell level

[0545] (1) HEK293T cell culture

[0546] HEK293T cells (purchased from ATCC) were inoculated into DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v), containing 1% Penicillin Streptomycin (v / v) (Gibco, 15140122), and cultured in a 37 °C cell incubator with 5% CO2. For the cells to be transfected, they were inoculated into a 24-well cell culture plate the day before for culture. The next day, the cells were observed, and when the cell density reached about 80%, transfection was carried out.

[0547] (2) The recombinant expression plasmid of CasY7 (see the plasmid map in Figure 1 of Figure 1 A), the CasY7-TTR-sgRNA expression plasmid, and the recombinant expression plasmid of LbCpf1 (see the plasmid map in Figure 1 of Figure 1 B), the LbCpf1-TTR-sgRNA expression plasmid and the EGFP-C1 (Addgene, Plasmid, #54759) plasmid were transfected into HEK293T cells respectively.

[0548] The dosage of plasmids transfected into each well of the 24-well plate was 0.3 μg of the nuclease expression plasmid (recombinant expression plasmid of CasY7 or recombinant expression plasmid of LbCpf1), 0.3 μg of the sgRNA expression plasmid (CasY7-TTR-sgRNA expression plasmid or LbCpf1-TTR-sgRNA expression plasmid), and 0.3 μg of the EGFP-C1 plasmid. The specific transfection operation is as follows:

[0549] The CasY7 expression plasmid, the CasY7-TTR-sgRNA expression plasmid, and the EGFP-C1 plasmid were mixed respectively and diluted with 25 μl of serum-free transfection medium (Yuanpei Biology, L530KJ), and then 2 μl of Lipofectamine 3000 (Invitrogen, L3000015) reagent was added and pipetted evenly to serve as reagent A, and left standing for 5 minutes. At the same time, 2 μl of Lipofectamine 3000 transfection reagent (Invitrogen, L3000015) was diluted and mixed evenly with 25 μl of serum-free transfection medium (Yuanpei Biology, L530KJ) to serve as reagent B, and left standing for 5 minutes.

[0550] The above-mentioned reagent A and reagent B were mixed and pipetted evenly, and left standing for 20 minutes. After standing, the mixed reagent was added dropwise to the cells in the 24-well plate to be transfected, and then placed back into the 37 °C, 5% CO2 incubator for culture. After 6 hours of transfection, the medium was changed to DMEM medium containing 10% FBS.

[0551] In the same way, the recombinant expression plasmid of LbCpf1, the LbCpf1-TTR-sgRNA expression plasmid, and the EGFP-C1 plasmid were transfected into HEK293T cells.

[0552] (3) Detection of editing efficiency

[0553] After 48 hours of transfection, the expression of EGFP fluorescent protein indicated successful cell transfection. Cells with positive EGFP expression were sorted for the detection of editing efficiency. The genomic DNA of the cells was extracted (using a genomic DNA extraction kit, TIANGEN, DP304-03). Identification primers were designed according to experimental requirements, and the sequences of the identification primers used are shown in the following table:

[0554]

[0555]

[0556] Using the genomic DNA as a template, the sequences near the target site were amplified by PCR with the primers in the above table. The PCR amplification system was as follows:

[0557] 2×Taq Master Mix (Vazyme, P112-03) 25 μL; Primer-F (TTR-F) (10 pmol / μL) 1 μL; Primer-R (TTR-R) (10 pmol / μL) 1 μL; template 1 μL; ddH2O was added to make up to 50 μL.

[0558] The amplified PCR products were used for high-throughput deep sequencing (Beijing Tsingke Biotechnology Co., Ltd.) or Sanger sequencing (Shanghai Biosciences Co., Ltd.) for the identification of editing efficiency.

[0559] By detecting and identifying the editing efficiency of CasY7 and LbCpf1, the results were as Figure 6 shown. In 293T cells, the editing efficiency of CasY7 was 33%, while that of LbCpf1 was only 22%. The editing efficiency of CasY7 was much higher than that of LbCpf1.

[0560] Example 4. Application of CasY7 in base editing

[0561] 1. Obtaining catalytically inactive CasY7

[0562] To obtain dCasY7 with no catalytic activity (i.e., loss of cleavage activity), the inventors constructed CasY7 mutants with single-point mutations of D592A, D643A, E820A, and D992A respectively: D592A-dCasY7, D643-dCasY7, E820A-dCasY7, D992A-dCasY7. The amino acid sequences of the above mutants are shown in SEQ ID NO.47-50 respectively. The specific construction method is as follows:

[0563] The CasY7+sgRNA expression plasmid obtained in step 1(3) of Example 2 was subjected to site-directed mutagenesis to modify the amino acids at 4 sites of CasY7, namely aspartic acid (Asp, D) at position 592, aspartic acid (Asp, D) at position 643, glutamic acid (Glu, E) at position 820, and aspartic acid (Asp, D) at position 992 of SEQ ID No.1, and the amino acids at the above sites were mutated to alanine (Ala, A). The codons before and after the mutation of each of the above amino acids are shown in the following table:

[0564] Amino acid before mutation Codon Amino acid after mutation Codon Aspartic acid (Asp, D) GAC Alanine (Ala, A) GCA Aspartic acid (Asp, D) GAC Alanine (Ala, A) GCA Glutamic acid (Glu, E) GAG Alanine (Ala, A) GCA Aspartic acid (Asp, D) GAC Alanine (Ala, A) GCA

[0565] Forward and reverse primers were designed and synthesized for the amino acids and their codons in the above table respectively, and then PCR amplification was carried out using the CasY7+sgRNA expression plasmid as a template. After amplification, the amplified products were recovered and purified using a universal DNA purification and recovery kit (Tiangen Biochemical Technology (Beijing) Co., Ltd., DP214). The purified products were transformed into Escherichia coli Dh5a competent cells (Vidy Biotechnology, DL1001), cultured overnight at 37°C, and single colonies were picked and sent for sequencing the next day. After confirmation by sequencing, the positive clones were shaken and the plasmids were extracted (TIANGEN, DP120-01), and then the concentration was measured and stored at -20°C in the refrigerator for later use.

[0566] The obtained site-directed mutagenesis recombinant plasmids were named respectively: D592A-dCasY7+sgRNA expression plasmid, D643A-dCasY7+sgRNA expression plasmid, E820A-dCasY7+sgRNA expression plasmid, D992A-dCasY7+sgRNA expression plasmid.

[0567] After that, the constructed D592A-dCasY7+sgRNA expression plasmid, D643A-dCasY7+sgRNA expression plasmid, E820A-dCasY7+sgRNA expression plasmid, and D992A-dCasY7+sgRNA expression plasmid were respectively subjected to in vivo editing efficiency detection in Escherichia coli. The detection method and calculation method were the same as those in steps 2 and 3 of Example 2. The experimental results are as Figure 7As shown in the figure, by counting the number of Escherichia coli clones and calculating the ratio, it is considered that D592A-dCasY7, D643A-dCasY7, E820A-dCasY7, and D992A-dCasY7 have lost their catalytic activity (cleavage activity). That is to say, the point mutations of D592A, D643A, E820A, and D992A have caused the CasY7 protein to lose its cleavage activity.

[0568] 2. Detection of base editing efficiency at the cellular level

[0569] (1) Construction of CasY7-TTR-sgRNA3 and CasY7-TTR-sgRNA4 plasmids

[0570] Design sgRNAs according to the target sequence of the TTR gene: CasY7-TTR-sgRNA3 and CasY7-TTR-sgRNA4 sequences, and synthesize oligonucleotides (oligos):

[0571] CasY7-TTR-sgRNA3:

[0572] CAAGTTGAATCCGTCTATAACTGACGG tatatcccttctacaaattc (SEQ ID NO.23);

[0573] CasY7-TTR-sgRNA4:

[0574] CAAGTTGAATCCGTCTATAACTGACGG gtgtctatttccactttgta (SEQ ID NO.24), where the underlined part of the sequence is the DR sequence, and the rest of the sequence is the spacer sequence.

[0575] (2) Add the CACC sequence to the 5' end of the upstream sequence of each sgRNA and the AAAA sequence to the 5' end of the downstream sequence. The specific form is as follows:

[0576]

[0577]

[0578] According to the method of step 1 of Example 3, anneal the upstream and downstream sequences of CasY7-TTR-sgRNA3 and CasY7-TTR-sgRNA4 and ligate them to the PHK09T vector to obtain the CasY7-TTR-sgRNA3 expression plasmid (named CasY7-TTR-sgRNA' expression plasmid) and the CasY7-TTR-sgRNA4 expression plasmid (named CasY7-TTR-sgRNA" expression plasmid), and transfer them into Escherichia coli DH5a competent cells for plasmid amplification culture. After correct sequencing and determination of the concentration, they are stored for later use.

[0579] (3) Construction of base editor plasmid (taking 005V1-10-3 as an example for illustration)

[0580] The adenosine deaminase catalytic domain selected by the inventors is a mutant of the amino acid sequence shown in SEQ ID NO: 31 (named 005V1-10-3): Q148G+Q149M+P150R. The amino acid sequence of this mutant is shown in SEQ ID NO. 32, and the nucleotide sequence encoding deaminase 005V1-10-3 is shown in SEQ ID NO. 33. A base editor fusion protein composed of deaminase 005V10-3 and CasY7 protein was constructed by homologous recombination. The specific operations are as follows:

[0581] First, a 005V1-10-3 nucleotide fragment with homologous arm sequences and linker was synthesized by Suzhou Huoxun Biotechnology Co., Ltd.:

[0582] The bold part is the 005V1-10-3 nucleotide sequence, the italic part is the left and right homologous arm regions, and the wavy line part is the linker sequence.

[0583] The D992A-dCasY7+sgRNA expression plasmid obtained in step 1 was PCR amplified for linearization to obtain a linearized expression vector. The primers used are shown in the following table:

[0584] Primer name Specific sequence D992A-dCasY7-F GCCACGAAGACCATCGTGAG (SEQ ID NO:35) D992A-dCasY7-R Gactttccgcttcttctttgg (SEQ ID NO:36)

[0585] The nucleotide fragment (SEQ ID NO: 34) of 005V1-10-3 with homologous arm sequences and linker and the linearized expression vector of D992A-dCasY7+sgRNA were subjected to homologous recombination, and the reaction was carried out using Gibson Assembly Master Mix (NEB, E2611S). After the reaction, the ligation product was transformed into Escherichia coli DH5a competent cells (Weidi Biology, DL1001). The specific process is as follows:

[0586] After the DH5α competent cells were taken out of the -80°C refrigerator, they were quickly inserted into ice. After 5 minutes, when the bacterial mass melted, the ligation product was added and the bottom of the centrifuge tube was gently tapped by hand to mix evenly, and then left standing in ice for 25 minutes. Heat shock at 42°C in a water bath for 45 seconds, quickly put it back into ice and let it stand for 2 minutes. Add 700 μl of sterile LB medium to the centrifuge tube, mix well, and resuscitate at 37°C and 200 rpm for 60 minutes. Centrifuge at 5000 rpm for 1 minute to collect the bacteria, leave about 100 μl of the supernatant, gently pipette to resuspend the bacterial mass and spread it on the LB medium with Amp antibiotic. Invert the plate and incubate it overnight in a 37°C incubator. Pick a single colony, after sequencing confirmation, shake the positive clone, extract the base editor plasmid using an endotoxin-free plasmid large extraction kit (TIANGEN: DP120-01), then measure the concentration and store it in a -20°C refrigerator for later use. Among them, the nucleotide sequence encoding the base editor fusion protein 005V1-10-3-D992A-dCasY7 is shown in SEQ ID NO:37, and the amino acid sequence is shown in SEQ ID NO:46.

[0587] (4) According to the method in step 2 of Example 3, co-transfect the base editor plasmid and the EGFP-C1 (Addgene, Plasmid#54759) plasmid into 293T cells with the CasY7-TTR-sgRNA’ expression plasmid and the CasY7-TTR-sgRNA” expression plasmid respectively.

[0588] 48 hours after transfection, the expression of EGFP fluorescent protein indicated successful cell transfection, and cells with positive EGFP expression were sorted for detection of editing efficiency. Extract the genome of the 293T cells using a kit (TIANGEN, DP304-03).

[0589] (5) Detect the base editing efficiency according to the method in step (3) of Example 3.

[0590] Design primers according to experimental requirements, and the sequences of the identification primers used are shown in the following table:

[0591]

[0592] The results are as Figure 8 in Figure 8 shown in A and 8B. The base editor composed of D992A-dCasY7 can achieve effective editing at multiple sites. As can be seen from Figure 8 A, under the guidance of CasY7-TTR-sgRNA’, effective editing exists at the +2, +4, +13, +15, +16, and +17 sites, and the editing efficiency at the +13, +15, +16, and +17 sites reaches nearly 10% to nearly 30%.

[0593] As can be seen fromFigure 8 It can be seen that under the guidance of "CasY7-TTR-sgRNA", effective editing exists at the +7, +13, and +20 sites. The base editing efficiency at the +13 site reaches more than 10%, and the editing efficiency at the +20 site reaches more than 15%.

[0594] Example 5. Screening of CasY7 mutants

[0595] To screen for CasY7 mutants with high cleavage activity, the region from the 7th to 684th positions of the CasY7 protein (SEQ ID NO.1) was mutated, and 31 mutants were constructed.

[0596] The mutants were obtained by using the CasY7 recombinant expression plasmid in Example 2 (see the plasmid map in Figure 1 of Figure 1 A) as a template, designing PCR primers centered on the mutation sites, introducing the mutated nucleotide sequences on the PCR primers, and then obtaining each CasY7 mutant plasmid through PCR site-directed mutagenesis.

[0597] The mutation methods of each mutant are shown in the following table:

[0598]

[0599]

[0600] According to the target sequences of the TTR gene, PD-1 gene, and Trac gene, gRNA targeting sequences were designed. The expression plasmids of each targeted sgRNA were constructed by the same method as in Example 3. The CasY7 recombinant expression plasmid or its mutant recombinant expression plasmid and the TTR-sgRNA2 expression plasmid were co-transfected into HEK293 cells by the PEI transfection method. A negative control group (using other targeted sgRNAs, the NT-sgRNA targeting sequence is shown in SEQ ID NO.52, labeled as "NT") and a blank control group were also co-transfected, and identification primers were designed according to each gRNA targeting.

[0601]

[0602] After culturing the cells for 48 hours, genomic DNA was extracted, and PCR amplification was performed using the identification primers. The obtained PCR products were used for high-throughput deep sequencing (Beijing Tsingke Biotechnology Co., Ltd.) to identify the editing efficiency.

[0603] The cleavage efficiencies of CasY7 wild-type (WT) and each mutant against the TTR-2 target, PD-1 target, Trac-1 target, Trac-2 target, Trac-3 target, and Trac-4 targets are respectively asFigures 9 - 14 As shown. Through analysis, it can be seen that, compared with the wild-type CasY7, the editing efficiency of each mutant of CasY7 for each target gene is much higher than that of CasY7. There is no or only background-level cleavage activity in the NT control group and the blank control group set for each target by the wild-type CasY7 (WT) and each mutant.

[0604] Example 6 Effect of Direct Repeat (DR) sequence on the cleavage activity of CasY7

[0605] To test the effect of different DR sequences on the activity of CasY7, deletions and mismatches were designed at different positions of the DR sequence, and a new DR variant: DR-1 (SEQ ID NO.71) was designed. The secondary structure of DR-1 ( Figure 15 ) is basically the same as that of the wild-type DR sequence (SEQ ID NO.5) ( Figure 2 ). The stem-loop regions of DR-1 and the wild-type DR are the same, while there are certain differences in other regions.

[0606] Taking the hTTR gene as the target, a spacer sequence (Target-TTR-spacer2, SEQ ID NO.11) targeting the target sequence of the TTR gene was designed, and DR-1-TTR-sgRNA (SEQ ID NO.72) with the DR-1 sequence was obtained. In the manner of Example 3, the synthesized DR-1-TTR-sgRNA (SEQ ID NO.72) sequence fragment was constructed into the PHK09T vector to obtain the DR-1-CasY7-TTR-sgRNA expression plasmid.

[0607] After that, cell transfection was carried out in the manner of steps (2) and (3) of Example 3, and the upstream and downstream primers TTR-F (SEQ ID NO.21) and TTR-R (SEQ ID NO.22) were used to detect the editing efficiency. Through analysis, it was found that the DR-1 sequence can mediate higher editing activity of CasY7 ( Figure 16 ). This also shows that CasY7 has adaptability to DR sequences with different structures.

[0608] Example 7. Screening of high-activity CasY7 mutants

[0609] To screen out CasY7 mutants with high cleavage activity, the CasY7 protein (SEQ ID NO.1) was subjected to site-directed mutagenesis and screening, and 17 mutants were constructed.

[0610] The mutants were obtained by using the CasY7 recombinant expression plasmid in Example 2 (for the plasmid map, see Figure 1 ofFigure 1 A) Using the mutation site as the center as a template, PCR primers were designed, and the nucleotide sequences after mutation were introduced onto the PCR primers. Subsequently, each CasY7 mutant plasmid was obtained through site-directed mutagenesis by PCR.

[0611]

[0612]

[0613] To efficiently and sensitively detect the cleavage activities of each mutant, a dual-luciferase reporter system vector (SEQ ID NO.73, Figure 17B ) was constructed, which contains the coding sequence of Fluc luciferase (Firefly luciferase), and the LUxUC coding sequence ( Figure 17A ). The LU and UC sequences in LUxUC are respectively the 119bp sequence at the 5'-end and the 469bp sequence at the 3'-end encoding NanoLuc luciferase (NanoLuc). There is an overlapping sequence of 60bp between the two sequences. An insertion sequence was designed in the middle of the LUxUC sequence, and the insertion sequence contains the target sequence for cleavage activity detection. In this example, the insertion sequence and the target sequence used are shown in the following table:

[0614] Target Inserted sequence Target sequence TTR SEQ ID NO.76 SEQ ID NO.6 PD-1’ SEQ ID NO.77 SEQ ID NO.75

[0615] There is a TAG premature termination codon in the middle of the target sequence. When it is cleaved, LU×UC uses the recombination mechanism to generate the correct NanoLuc coding frame, express NanoLuc luciferase, and then catalyze the substrate to emit fluorescence. The luminescence intensity indicates the cleavage activity. Among them, Fluc luciferase is used as an internal reference, and the fluorescence emitted is used to indicate whether the reporter vector has been successfully transfected into host cells.

[0616] Using the same method as in Example 1, recombinant expression plasmids of each mutant were constructed. sgRNA expression plasmids targeting TTR and targeting PD-1' were constructed in the manner of Example 3, and by using the transfection method in Example 3, the recombinant expression plasmids of CasY7 and each mutant were co-transfected into the HEK293 cell line with the dual-luciferase reporter system vector and the sgRNA expression plasmid respectively. After culturing for 24 h, the expression intensities of Fluc and NanoLuc were detected on an enzyme-labeled instrument using a dual-luciferase reporter gene detection kit (Beyotime, RG028).

[0617] In addition, negative control groups were also set up for CasY7 and each mutant. The negative control groups used sgRNA expression plasmids with non-target sequences (SEQ ID NO.78). Analysis found that the cleavage activities of CasY7 and each mutant were as shown in the following table, while no cleavage activity was detected in the negative control groups, and the cleavage activities of each mutant were much higher than that of CasY7.

[0618] Mutant Targeted TTR cleavage activity (%) Targeted PD-1’ cleavage activity (%) C05440 4.05 8.39 C03897 33.92 31.64 C06865 32.49 32.51 C04558 30.10 27.15 C08675 23.80 35.92 C01144 29.49 31.38 C06623 29.33 34.29 C03023 29.75 39.67 C07879 30.38 35.04 C03803 25.31 38.54 C06467 27.84 18.66 C08724 26.15 36.45 C03563 26.11 35.15 C05664 28.69 40.12 C02742 21.18 35.02 C02009 27.09 32.39 C09752 20.31 35.04

[0619] Sequence information:

[0620]

[0621]

[0622]

[0623]

[0624]

[0625]

[0626]

[0627]

[0628]

[0629]

[0630]

[0631] All documents mentioned in the present invention are incorporated herein by reference as if each document was individually incorporated by reference. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of the present application.

Claims

1. A Cas protein, characterized in that, The protein is selected from the following group: (a) A polypeptide having the amino acid sequence shown in SEQ ID NO:1; (b) A polypeptide having a homology (or identity) of ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% with the amino acid sequence shown in SEQ ID NO:1, and the polypeptide has the biological function of SEQ ID NO:1; (c) A derivative polypeptide formed by substituting, deleting or adding one or more (preferably 1-20, more preferably 1-10, even more preferably 1-5) amino acid residues in any of the amino acid sequences shown in SEQ ID NO:1, and retaining the biological function of SEQ ID NO:

1.

2. The Cas protein according to claim 1, wherein The Cas protein relative to the polypeptide having the amino acid sequence shown in SEQ ID NO:1: (1) Contains mutations at one or more sites, such that the cleavage activity or preference is altered; or (2) Mutations occur at one or more sites corresponding to SEQ ID NO:1 in the wild-type Cas protein selected from the following group: V7, T11, S12, Y165, S166, S196, K197, S198, A199, S240, S241, Q242, E243, I244, N276, Y282, D283, A285, I304, L307, Y308, S309, R315, E316, T317, I318, I319, V348, I349, E350, P351, G416, I417, E418, F419, D420, E648, L650, A651, Y652, T681, N682, E683, S684, E748, G749, S751, K865, P866, Y867, N868, I872, M957, F958, Q960, W961; Or (3) Mutations occur at the amino acid sites corresponding to SEQ ID NO:1 in the wild-type Cas protein selected from one or more sites of the following group: (a) The amino acid substitution at D283 is selected from D283N, D283Q, D283S, D283E, preferably selected from D283N, D283Q, D283S; (b) The amino acid substitution at A285 is selected from A285V, A285L, A285I, A285S, A285T, A285G, preferably selected from A285V, A285T, A285P, more preferably selected from A285T, A285P; (c) The amino acid substitution at Y308 is selected from Y308W, Y308F, Y308T, Y308S, Y308A, preferably selected from Y308F, Y308A; (d) The amino acid substitution at R315 is selected from R315K, R315Q, R315N, R315H, R315A, R315S, preferably selected from R315K, R315A, R315S, more preferably selected from R315A, R315S; (e) The amino acid substitution at E316 is selected from E316D, E316R, E316Q, preferably selected from E316R, E316Q; (f) The amino acid substitution at T317 is selected from T317S, T317K, T317R, T317N, T317Q, T317A, preferably selected from T317S, T317K, T317R, more preferably selected from T317K, T317R; (g) The amino acid substitution at I318 is selected from I318L, I318V, I318M, I318A, I318F, I318G, I318K, preferably selected from I318L, I318K; (h) The amino acid substitution at I319 is selected from I319L, I319V, I319M, I319A, I319F, I319Q, I319W, I319G, preferably selected from I319L, I319Q, I319W, more preferably selected from I319Q, I319W; (i) The amino acid substitution at G416 is selected from G416P, G416A, G416V, G416L, G416I, G416S, G416R, preferably selected from G416A, G416S, G416R, more preferably selected from G416S, G416R; (j) The amino acid substitution at I417 is selected from I417L, I417V, I417M, I417A, I417F, I417G, I417H, preferably selected from I417L, I417H, I417V, more preferably selected from I417H, I417V; (k) The amino acid substitution at F419 is selected from F419L, F419V, F419I, F419A, F419Y, F419W, F419T, preferably selected from F419L, F419T, F419V, more preferably selected from F419T, F419V; (l) The amino acid substitution at D420 is selected from D420E, D420T, D420V, preferably selected from D420T, D420V; (m) The amino acid substitution at E648 is selected from E648D, E648L, E648R, E648S, E648T, preferably selected from E648L, E648R, E648S, E648T; (n) The amino acid substitution at A651 is selected from A651V, A651L, A651I, A651G, A651P, A651S, A651T, preferably selected from A651V, A651L, A651G, A651P, A651S, more preferably selected from A651L, A651G, A651P, A651S. More preferably, the variant is a non-natural protein, and the variant has one or more mutations in the wild-type Cas protein corresponding to SEQ ID NO: 1 selected from the group consisting of: V7I, T11V, S12M, Y165W, S166R, S196N, K197F, S198N, A199T, S240L, S241A, Q242S, E243H, I244L, N276S, Y282F, D283N, D283Q, D283S, A285T, A285P, I304A, L307R, Y308A, Y308F, S309A, R315A, R315S, E316R, E316Q, T317K, T317R, I318K, I318L, I319Q, I319W, V348I, I349V, E350L, P351T, G416S, G416R, I417H, I417V, E418Q, F419T, F419V, D420T, D420V, E648L, E648R, E648S, E648T, L650I, A651L, A651G, A651P, A651S, Y652F, T681P, N682P, E683Q, S684C, E748A, G749S, S751G, K865S, P866S, Y867H, N868F, I872L, M957V, F958L, Q960V, W961V; Even more preferably, the mutant is selected from the following group: Y282F + D283Q + A285T + G416S + I417H + E418Q + F419T + D420V; Y165W + S166R + G416R + I417V + E418Q + F419V + D420T; Y282F + D283Q + A285T + V348I + I349V + E350L + P351T; D283S + A285P + L307R + Y308F + S309A; L307R + Y308F + S309A + E648R + L650I + A651P + Y652F; Y282F + D283Q + A285T + E648T + A651S + Y652F; Y282F + D283Q + A285T + E648L + A651G; Y282F + D283Q + A285T + T681P + N682P + E683Q + S684C; Y282F + D283Q + A285T + E748A + G749S + S751G; N276S + D283N + E648S + A651L; Y282F + D283Q + A285T + R315A + E316Q + T317R + I318L + I319Q; Y282F + D283Q + A285T + R315S + E316R + T317K + I318K + I319W + E648S + A651L; V7I + T11V + S12M + Y282F + D283Q + A285T; S240L + S241A + Q242S + E243H + I244L + L307R + Y308F + S309A; S196N + K197F + S198N + A199T + I304A + L307R + Y308A + S309A; Y282F + D283Q + A285T + K865S + P866S + Y867H + N868F + I872L; Y282F + D283Q + A285T + M957V + F958L + Q960V + W961V.

3. A fusion protein, characterized in that, A Cas protein as claimed in claim 1 or 2; and one or more functional domains.

4. An isolated polynucleotide, characterized in that, The polynucleotide encodes the Cas protein as claimed in any one of claims 1 - 2 or the fusion protein as claimed in claim 3; preferably, the polynucleotide has been codon-optimized for expression in eukaryotic cells.

5. A guide RNA (gRNA), characterized in that, The guide RNA includes a direct repeat (DR) sequence capable of binding to the Cas protein described in claim 1 and a spacer sequence capable of targeting a target sequence; preferably, the direct repeat (DR) sequence comprises the nucleotide sequence shown in any one of SEQ ID NO.5 and 71, or a nucleotide sequence having at least about 80% homology with the nucleotide sequence shown in any one of SEQ ID NO.5 and 71; more preferably, the DR sequence comprises a stem-loop structure near the 3' end of the DR sequence, and the stem-loop structure comprises 5'-X1X2X3X4X5NNNNNNNX6X7X8X9X 10 -3'; X1, X2, X3, X4, X5, X6, X7, X8, X9, X 10 is any base comprising A, T, C or G, and N is any base comprising A, T, C or G; wherein X1, X2, X3, X4, X5 and X6, X7, X8, X9, X 10 can hybridize with each other to form a stem and such that NNNNNNN forms a loop; more preferably, wherein the DR sequence comprises a stem-loop structure near the 3' end of the DR sequence as shown below: 5'-CCGTCNNNNNNNGACGG-3' (SEQ ID NO.80); wherein N is any base comprising A, T, C or G.

6. An isolated nucleic acid molecule, characterized in that, Comprising a sequence selected from, or consisting of a sequence selected from: (i) the sequence shown in SEQ ID NO: 5 or 71; (ii) a sequence having one or more base substitutions, deletions or additions (such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared with the sequence shown in SEQ ID NO: 5 or 71; (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity with the sequence shown in SEQ ID NO: 5 or 71; (iv) a sequence that hybridizes with the sequence described in any one of (i) - (iii) under stringent conditions; or (v) the complementary sequence of the sequence described in any one of (i) - (iii); and, the sequence described in any one of (ii) - (v) substantially retains the biological function of the sequence from which it is derived; for example, the isolated nucleic acid molecule is RNA; for example, the isolated nucleic acid molecule comprises direct repeat (DR) sequences in the CRISPR / Cas system.

7. A composite, characterized in that, Comprising: (i) a protein component, selected from the group consisting of: the Cas protein as claimed in any one of claims 1 - 2, the fusion protein as claimed in claim 3, or a combination thereof; and (ii) a nucleic acid component, selected from the group consisting of: the guide RNA as claimed in claim 5, the nucleic acid encoding the guide RNA as claimed in claim 5, the precursor RNA of the guide RNA as claimed in claim 4, the nucleic acid of the precursor RNA encoding the guide RNA as claimed in claim 5, or a combination thereof; wherein, the protein component binds to the nucleic acid component to form a complex.

8. A carrier, characterized in that, Comprising the polynucleotide according to claim 4.

9. A CRISPR-Cas composition, characterized in that, Comprising: (i) A first component selected from the group consisting of: the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, a nucleotide sequence encoding the Cas protein according to any one of claims 1-2 or the fusion protein according to claim 3, and any combination thereof; and (ii) A second component, which is a guide RNA comprising one or more according to claim 5, or a nucleotide sequence encoding the guide RNA comprising one or more according to claim 4; The guide RNA is capable of forming a complex with the protein or protein variant or fusion protein described in (i).

10. A CRISPR-Cas system, characterized in that, Comprising one or more vectors, the one or more vectors comprising: (i) A first nucleic acid, which is a nucleotide sequence encoding the Cas protein according to any one of claims 1-2 or the fusion protein according to claim 3; optionally the first nucleic acid is operably linked to a first regulatory element; and (ii) A second nucleic acid, which encodes a nucleotide sequence comprising the guide RNA according to claim 4; optionally the second nucleic acid is operably linked to a second regulatory element; Wherein: The first nucleic acid and the second nucleic acid are present on the same or different vectors; The guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

11. A kit, characterized in that, Comprising one or more components selected from the following: the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9 or the system according to claim 10.

12. A delivery composition, characterized in that, Comprising a delivery vector and one or more selected from the following: the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9 or the system according to claim 10.

13. A host cell, characterized in that, Comprising the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9, the system according to claim 10 or the delivery composition according to claim 12.

14. An enzyme preparation, characterized in that, The enzyme preparation comprises the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, the complex according to claim 7, the CRISPR-Cas composition according to claim 9 or the system according to claim 10 or the delivery composition according to claim 12.

15. A medicine box, characterized in that, Comprising: A first container, and the complex according to claim 7 or the CRISPR-Cas composition according to claim 9 or the system according to claim 10 as described in the first container, or a drug containing the complex according to claim 7 or the CRISPR-Cas composition according to claim 9 or the system according to claim 10.

16. A medicine box, characterized in that, Comprising: (a1) A first container, and the Cas protein according to any one of claims 1-2 or the fusion protein according to claim 3 or its coding gene or its expression vector located in the first container, or a drug containing the Cas protein variant according to any one of claims 1-2 or the fusion protein according to claim 3 or its coding gene or its expression vector. (b1) Optionally, a second container, and the guide RNA according to claim 5 or its expression vector located in the second container, or a drug containing the guide RNA according to claim 5 or its expression vector.

17. A method for targeting and editing a target gene or cleaving a target gene, characterized in that, Comprising: Contacting the Cas protein according to any one of claims 1-2 or the fusion protein according to claim 3 or the complex according to claim 7 or the composition according to claim 9 or the system according to claim 10 or the delivery composition according to claim 12 or the enzyme preparation according to claim 14 or the kit according to claim 15 or 16 with the target gene, or delivering it into a cell containing the target gene, wherein the target sequence is present in the target gene.

18. A method for inducing a change in cell state, characterized in that, The method comprises contacting the Cas protein according to any one of claims 1-2 or the fusion protein according to claim 3 or the complex according to claim 7 or the composition according to claim 9 or the system according to claim 10 or the delivery composition according to claim 12 or the enzyme preparation according to claim 14 or the kit according to claim 15 or 16 with the target gene in the cell.

19. A method for altering the expression of a gene product, characterized in that, Comprising: Contacting the Cas protein according to any one of claims 1-2 or the fusion protein according to claim 3 or the complex according to claim 7 or the composition according to claim 9 or the system according to claim 10 or the delivery composition according to claim 12 or the enzyme preparation according to claim 14 or the kit according to claim 15 or 16 with the nucleic acid molecule encoding the gene product, or delivering it into a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule. Use of the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9 or the system according to claim 10 or the kit according to claim 11 or the delivery composition according to claim 12 or the enzyme preparation according to claim 14 or the kit according to claim 15 or 16, characterized in that, For preparing a drug or a preparation, the drug or the preparation is used for nucleic acid editing (for example, gene or genome editing).

21. Use of the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, the polynucleotide according to claim 4, the complex according to claim 7, the vector according to claim 8, the CRISPR-Cas composition according to claim 9, or the system according to claim 10, or the kit according to claim 11, or the delivery composition according to claim 12, or the enzyme preparation according to claim 14, or the cartridge according to claim 15 or 16, characterized in that, For preparing a drug or a preparation, the drug or the preparation is used for one or more selected from the following groups: (i) Ex vivo gene or genome editing; (ii) Detection of single-stranded DNA ex vivo; (iii) Editing the target sequence in the target locus to modify a biological or non-human organism; (iv) Treating a disease caused by a defect in the target sequence in the target locus; (v) Treating a disease or disorder of a subject in need.

22. A method for detecting the presence of a target nucleic acid molecule in a sample, characterized in that, The method includes contacting a sample with the Cas protein according to any one of claims 1-2, the fusion protein according to claim 3, or the complex according to claim 7, the CRISPR-Cas composition according to claim 9 or the system according to claim 10, the kit according to claim 11 or the delivery composition according to claim 12, or the enzyme preparation according to claim 14 and a non-target sequence, and detecting a detectable signal generated by cleavage of the non-target sequence, so as to detect a target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA.

Citation Information

Patent Citations

  • Adenosine deaminase, base editor fusion protein, base editor system and application

    CN114634923A

  • Polypeptides comprising multimers of nuclear localization signals or of protein transduction domains and their use for transferring molecules into cells

    WO2001038547A2