A Cas protein, its gene editing system and applications

By performing key amino acid site mutations and fusion functional domains on Cas proteins, the flexibility and efficiency of the existing CRISPR/Cas system in the targeted recognition and cleavage process is solved, and efficient gene editing and disease treatment effects are achieved.

CN116376874BActive Publication Date: 2025-07-22YOLTECH THERAPEUTICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310302690.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-07-22
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

The existing CRISPR/Cas system has unmet application requirements in the targeted recognition and cleavage process, especially Cas9, C2c1 and CasX require two RNAs to be guided, and the PAM sequences are complex and diverse, resulting in insufficient efficiency and flexibility.

Method used

A novel Cas protein and its variants are provided, which reduces dependence on guide RNA by mutation at key amino acid sites and binds to functional domains to form fusion proteins to improve targeted cleavage efficiency and flexibility.

Benefits of technology

It realizes efficient editing and cleavage of target genes, adapts to diverse application needs, improves the flexibility and accuracy of gene editing, and is suitable for gene therapy and disease treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004145885270000151
    Figure BDA0004145885270000151
  • Figure BDA0004145885270000291
    Figure BDA0004145885270000291
  • Figure BDA0004145885270000293
    Figure BDA0004145885270000293
Patent Text Reader

Abstract

The present invention provides a Cas protein, its gene editing system and applications. Specifically, the Cas protein of the present invention has very good gene editing activity, can effectively edit or cleave target genes, and can effectively treat the diseases or disorders of subjects in need.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gene editing, and specifically, to a Cas protein, its gene editing system and applications. Background Art

[0002] The Clustered regularly interspaced short palindromic repeats (CRISPR) system is formed by bacteria and archaea to defend against the DNA of invading phages. The immune interference process of the CRISPR system mainly includes three stages: adaptation, expression, and interference. In the adaptation stage, the CRISPR system integrates short DNA fragments from phages or plasmids between the leader sequence and the first repeat sequence. Each integration is accompanied by the replication of the repeat sequence, thereby forming a new repeat-spacer sequence unit. In the expression stage, the CRISPR locus is transcribed into a CRISPR RNA (crRNA) precursor (pre-crRNA), which is further processed into small crRNAs at the repeat sequence in the presence of Cas proteins and tracrRNA. The mature crRNA forms a Cas / crRNA complex with the Cas protein. In the interference stage, the crRNA guides the Cas / crRNA complex to find the target through its region complementary to the target sequence, and the nuclease activity of the Cas protein causes double-stranded DNA breaks at the target site, thereby rendering the target DNA dysfunctional.

[0003] The CRISPR system is divided into three families: type I, type II, and type III. Among them, the most common type II system is the CRISPR / Cas9 system. The Cas9 protein can process pre-crRNA into mature crRNA that binds to tracrRNA with the assistance of trans-encoded small RNA (tracrRNA). Subsequently, it was found that by artificially constructing a single-stranded chimeric guide RNA (guideRNA) that mimics the crRNA:tracrRNA complex, the Cas9 protein can be effectively mediated to recognize and cleave the target. The three bases immediately adjacent to the 3' end of the target must be in the form of 5'-NGG-3', thereby constituting the protospacer adjacent motif (PAM) structure required for the Cas / crRNA complex to recognize the target. However, the existing different CRISPR / Cas systems each have different advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two RNAs for the guide RNA. The common Cas9, C2c1, CasY, and Cpf1 are usually about 1300 amino acids in size. In addition, the PAM sequences of Cas9, Cpf1, CasX, and CasY are all complex and diverse.

[0004] There is still a need to develop new Cas proteins and CRISPR-Cas systems to meet diverse application requirements. SUMMARY OF THE INVENTION

[0005] The main object of the present invention is to provide a new Cas protein, its gene editing system and applications to meet the above application requirements. Based on this, the present invention also provides a new CRISPR-Cas composition, as well as a gene editing method and a nucleic acid detection method based on this system.

[0006] The first aspect of the present invention provides a protein, which is selected from the following group:

[0007] (a) a polypeptide having the amino acid sequence shown in SEQ ID NO: 1;

[0008] (b) a polypeptide having a homology (or identity) of ≥ 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% with the amino acid sequence shown in SEQ ID NO: 1, and the polypeptide has the biological function of SEQ ID NO: 1;

[0009] (c) a derivative polypeptide formed by substitution, deletion or addition of one or more (preferably 1-20, more preferably 1-10, even more preferably 1-5) amino acid residues in the amino acid sequence shown in any of SEQ ID NO: 1, and retaining the biological function of SEQ ID NO: 1.

[0010] In another preferred embodiment, the protein is an effector protein in the CRISPR / Cas system.

[0011] The second aspect of the present invention provides a protein variant, which is a non-natural protein, and the variant has mutations at one or more core amino acid sites corresponding to SEQ ID NO: 1 in the wild-type protein and related to cleavage activity:

[0012] the aspartic acid (D) site at position 659; and / or

[0013] the aspartic acid (D) site at position 711; and / or

[0014] the glutamic acid (E) site at position 895; and / or

[0015] the aspartic acid (D) site at position 1069.

[0016] In another preferred example, the protein variant is a variant of an effector protein in the CRISPR / Cas system.

[0017] In another preferred example, relative to the cleavage activity of the wild-type protein, the cleavage activity of the protein variant against the target sequence of a target molecule complementary to the guide RNA sequence is reduced (e.g., reduced by 50%, 60%, 70%, 80%, 90%, 95% or more) or substantially lacking.

[0018] In another preferred example, the aspartic acid (D) at position 659 is mutated to any amino acid, preferably to the amino acid types shown in Table A, preferably to one or more amino acids selected from the group consisting of: Ala (A), Val (V), Leu (L), Ile (I), preferably to Ala (A), Val (V), more preferably Ala (A).

[0019] In another preferred example, the aspartic acid (D) at position 711 is mutated to any amino acid, preferably to the amino acid types shown in Table A, preferably to one or more amino acids selected from the group consisting of: Ala (A), Val (V), Leu (L), Ile (I), preferably to Ala (A), Val (V), more preferably Ala (A).

[0020] In another preferred example, the glutamic acid (E) at position 895 is mutated to any amino acid, preferably to the amino acid types shown in Table A, preferably to one or more amino acids selected from the group consisting of: Ala (A), Val (V), Leu (L), Ile (I), preferably to Ala (A), Val (V), more preferably Ala (A).

[0021] In another preferred example, the aspartic acid (D) at position 1069 is mutated to any amino acid, preferably to the amino acid types shown in Table A, preferably to one or more amino acids selected from the group consisting of: Ala (A), Val (V), Leu (L), Ile (I), preferably to Ala (A), Val (V), more preferably Ala (A).

[0022] In another preferred example, the aspartic acid (D) at position 659 is mutated to alanine (A).

[0023] In another preferred example, the aspartic acid (D) at position 711 is mutated to alanine (A).

[0024] In another preferred example, the glutamic acid (E) at position 895 is mutated to alanine (A).

[0025] In another preferred example, the aspartic acid (D) at position 1069 is mutated to alanine (A).

[0026] In another preferred embodiment, the mutation is selected from the group consisting of: D659A, D711A, E895A, D1069A, or a combination thereof.

[0027] In another preferred embodiment, the amino acid sequence of the protein variant is as shown in any one of SEQ ID NOs. 44-47.

[0028] In another preferred embodiment, the protein variant is a polypeptide having the amino acid sequence shown in any one of SEQ ID NOs. 44-47, an active fragment thereof, or a conservative variant polypeptide thereof.

[0029] In another preferred embodiment, except for the mutation (such as at positions 659, 711, 895, and / or 1069), the remaining amino acid sequence of the protein variant is the same as or substantially the same as the sequence of the wild-type protein.

[0030] In another preferred embodiment, the substantially the same means that there are at most 50 (preferably 1-20, more preferably 1-10, even more preferably 1-5) amino acids that are different, where the differences include amino acid substitution, deletion, or addition, and the cleavage activity of the protein variant is reduced.

[0031] In another preferred embodiment, the homology of the variant to the wild-type protein is at least 80%, preferably at least 85% or 90%, more preferably at least 95%, and most preferably at least 98% or 99%.

[0032] In another preferred embodiment, the protein variant is selected from the group consisting of:

[0033] (a) a polypeptide having the amino acid sequence shown in any one of SEQ ID NOs. 44-47;

[0034] (b) a polypeptide derived from (a) formed by substitution, deletion, or addition of one or more (such as 2, 3, 4, or 5) amino acid residues in the amino acid sequence shown in any one of SEQ ID NOs. 44-47, and having reduced cleavage activity.

[0035] In another preferred embodiment, the homology of the derived polypeptide to the sequence shown in any one of SEQ ID NOs.: 44-47 is at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90%, such as 95%, 97%, 99%.

[0036] In another preferred embodiment, the protein variant is formed by mutation of the wild-type protein.

[0037] The third aspect of the present invention provides a fusion protein, comprising the protein described in the first aspect of the present invention or the protein variant described in the second aspect of the present invention; and one or more functional domains.

[0038] In another preferred embodiment, the functional domain is selected from a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage-active polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, and a base excision repair inhibitor (such as uracil-DNA glycosylase inhibitor (UGI)).

[0039] In another preferred embodiment, the functional domain includes one or more of the following enzymatic activities on a target sequence: methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity.

[0040] In another preferred embodiment, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0041] In another preferred embodiment, the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0042] In another preferred embodiment, the adenosine deaminase catalytic domain comprises an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence shown in SEQ ID NO: 28 (005V1 deaminase selected from CN114634923A, the amino acid sequence of which is SEQ ID NO: 2 in this application), and it retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 28.

[0043] In another preferred embodiment, the amino acid sequence of the adenosine deaminase catalytic domain has amino acid additions, insertions, deletions, and substitutions relative to the amino acid sequence shown in SEQ ID NO: 28.

[0044] In another preferred example, the adenosine deaminase catalytic domain includes a mutant of the amino acid sequence shown in SEQ ID NO: 29: Q148G+Q149M+P150R, named deaminase 005V1-10-3.

[0045] In another preferred example, the functional domain is the full length or a functional fragment of TadA8e.

[0046] In another preferred example, the localization signal includes a nuclear localization signal (NLS) and / or a nuclear export signal (NES).

[0047] In another preferred example, the sequence of the nuclear localization signal is as shown in any one of SEQ ID NO: 35-42.

[0048] In another preferred example, the sequence of the nuclear localization signal is located at, near or close to the end (e.g., N-terminus or C-terminus) of the protein as claimed in claim 1.

[0049] In another preferred example, the nuclear export signal includes protein tyrosine kinase 2 (such as human protein tyrosine kinase 2).

[0050] In another preferred example, the reporter protein includes glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, a self-fluorescent protein.

[0051] In another preferred example, the self-fluorescent protein includes green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent protein (e.g., (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire).

[0052] In another preferred example, the DNA binding domain includes a methylation binding protein, LexADBD, Gal4DBD.

[0053] In another preferred example, the epitope tag includes a histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, thioredoxin tag, streptavidin tag.

[0054] In another preferred example, the transcriptional activation domain comprises VP64 and / or VPR.

[0055] In another preferred example, the transcriptional repression domain comprises KRAB and / or SID.

[0056] In another preferred example, the nuclease comprises FokI.

[0057] In another preferred example, the cleavage active polypeptide comprises a polypeptide with single-stranded RNA cleavage activity, a polypeptide with double-stranded RNA cleavage activity, a polypeptide with single-stranded DNA cleavage activity or a polypeptide with double-stranded DNA cleavage activity.

[0058] In another preferred example, the ligase comprises DNA ligase and / or RNA ligase.

[0059] In another preferred example, the functional domain is connected to the N-terminus and / or C-terminus of the protein.

[0060] In another preferred example, the functional domain is inserted between the N-terminus and C-terminus of the protein.

[0061] In another preferred example, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the protein through a linker.

[0062] In another preferred example, the functional domain is inserted between the N-terminus and C-terminus of the protein through a linker.

[0063] In another preferred example, the fusion protein has the following structure from the N-terminus to the C-terminus:

[0064] Z1-Z2(I); or

[0065] Z2-Z1(II); or

[0066] Z3-Z1-Z4(III);

[0067] Wherein, Z1 is cytosine deaminase or adenosine deaminase;

[0068] Z2 is the protein described in claim 1;

[0069] Z3 is the N-terminal fragment of the protein described in claim 1;

[0070] Z4 is the C-terminal fragment of the protein described in claim 1;

[0071] And each "-" is independently a bond or a linker.

[0072] In another preferred example, the fusion protein has the amino acid sequence shown in SEQ ID NO: 43.

[0073] The fourth aspect of the present invention provides an isolated polynucleotide encoding the protein described in the first aspect of the present invention, or the protein variant described in the second aspect of the present invention, or the fusion protein described in the third aspect of the present invention.

[0074] In another preferred embodiment, the polynucleotide sequence comprises the polynucleotide of the sequence shown in SEQ ID NO. 30.

[0075] In another preferred embodiment, the polynucleotide is selected from the group consisting of:

[0076] (a) a polynucleotide having a sequence as shown in SEQ ID NO. 2 or 34;

[0077] (b) a polynucleotide having a nucleotide sequence with a homology of ≥70% (preferably ≥80%, more preferably ≥90%, more preferably ≥95%, most preferably ≥99%) to the sequence shown in SEQ ID NO. 2 or 34 and encoding the polypeptide shown in SEQ ID NO.: 1 or 43;

[0078] (c) a polynucleotide complementary to any one of the polynucleotides described in (a)-(b).

[0079] In another preferred embodiment, the polynucleotide further contains an auxiliary element selected from the group consisting of a signal peptide, a secretion peptide, a tag sequence (such as 6His), or a combination thereof, flanking the ORF of the variant.

[0080] In another preferred embodiment, the polynucleotide is selected from the group consisting of: genomic sequence, cDNA sequence, RNA sequence, or a combination thereof.

[0081] In another preferred embodiment, the polynucleotide further comprises a promoter operably linked to the ORF sequence of the variant.

[0082] In another preferred embodiment, the promoter is selected from the group consisting of: constitutive promoter, tissue-specific promoter, inducible promoter, or strong promoter.

[0083] In another preferred embodiment, the polynucleotide is a polynucleotide optimized for codon usage according to the codon preference of the host cell.

[0084] In another preferred embodiment, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0085] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).

[0086] In another preferred embodiment, the host cell is a prokaryotic cell, such as Escherichia coli.

[0087] In another preferred embodiment, the yeast cell is yeast from one or more sources selected from the group consisting of Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell comprises Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.

[0088] In another preferred embodiment, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.

[0089] The fifth aspect of the present invention provides an isolated nucleic acid molecule comprising a sequence selected from the following, or consisting of a sequence selected from the following:

[0090] (i) The sequence shown in SEQ ID NO: 5;

[0091] (ii) A sequence having one or more base substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base substitutions, deletions, or additions) compared to the sequence shown in SEQ ID NO: 5;

[0092] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to the sequence shown in SEQ ID NO: 5;

[0093] (iv) A sequence that hybridizes to the sequence described in any one of (i)-(iii) under stringent conditions; or

[0094] (v) The complementary sequence of the sequence described in any one of (i)-(iii);

[0095] And, the sequence described in any one of (ii)-(v) substantially retains the biological function of the sequence from which it is derived;

[0096] For example, the isolated nucleic acid molecule is RNA;

[0097] For example, the isolated nucleic acid molecule comprises a direct repeat sequence in the CRISPR / Cas system.

[0098] In another preferred embodiment, the nucleic acid molecule comprises one or more stem-loops or optimized secondary structures;

[0099] For example, the sequence described in any one of (ii)-(v) retains the secondary structure of the sequence from which it is derived.

[0100] In another preferred embodiment, the nucleic acid molecule is a nucleic acid molecule codon-optimized according to the codon preference of the host cell.

[0101] In another preferred embodiment, the nucleic acid molecule comprises a sequence selected from the following, or consists of a sequence selected from the following:

[0102] (a) the nucleotide sequence shown in SEQ ID NO: 5;

[0103] (b) a sequence that hybridizes with the sequence described in (a) under stringent conditions; or

[0104] (c) the complementary sequence of the nucleotide sequence shown in SEQ ID NO: 5.

[0105] The sixth aspect of the present invention provides a guide RNA (gRNA), which includes a direct repeat (DR) sequence capable of binding to the protein described in the first aspect of the present invention and a spacer sequence capable of targeting a target sequence.

[0106] The seventh aspect of the present invention provides a complex, comprising:

[0107] (i) a protein component, selected from the group consisting of: the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, or a combination thereof; and

[0108] (ii) a nucleic acid component, selected from the group consisting of: the guide RNA described in the sixth aspect of the present invention, the nucleic acid encoding the guide RNA described in the sixth aspect of the present invention, the precursor RNA of the guide RNA described in the sixth aspect of the present invention, the nucleic acid of the precursor RNA encoding the guide RNA described in the sixth aspect of the present invention, or a combination thereof;

[0109] wherein the protein component and the nucleic acid component bind to each other to form a complex.

[0110] In another preferred embodiment, the direct repeat (DR) sequence in the guide RNA (gRNA) is linked to the 3'-end or 5'-end of the nucleic acid molecule.

[0111] In another preferred embodiment, the spacer sequence in the guide RNA (gRNA) comprises the complementary sequence of the target sequence.

[0112] The eighth aspect of the present invention provides a vector, comprising the polynucleotide described in the fourth aspect of the present invention or the nucleic acid molecule described in the fifth aspect of the present invention.

[0113] In another preferred embodiment, the vector comprises:

[0114] (1) A first regulatory element, which is operably linked to a nucleotide sequence encoding the protein according to the first aspect of the present invention, or a nucleotide sequence encoding the protein variant according to the second aspect of the present invention, or a nucleotide sequence encoding the fusion protein according to the third aspect of the present invention; and

[0115] (2) A second regulatory element, which is operably linked to a nucleotide sequence encoding a guide RNA, the guide RNA comprising:

[0116] (a) A spacer sequence capable of hybridizing with a target sequence, and

[0117] (b) A direct repeat (DR) sequence, which is linked to the spacer sequence and can direct the protein according to the first aspect of the present invention or the protein variant according to the second aspect of the present invention to bind to the guide RNA to form the complex according to the seventh aspect of the present invention that targets the target sequence.

[0118] In another preferred embodiment, the first regulatory element and the second regulatory element are located on the same or different vectors.

[0119] In another preferred embodiment, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.

[0120] In another preferred embodiment, the vector comprises one or more promoters, which are operably linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, replication origin, selectable marker, nucleic acid restriction site, and / or homologous recombination site.

[0121] In another preferred embodiment, the vector includes a plasmid, a viral vector.

[0122] In another preferred embodiment, the viral vector is selected from the group consisting of: adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or a combination thereof.

[0123] In another preferred embodiment, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, a multifunctional vector.

[0124] The ninth aspect of the present invention provides a host cell, comprising the polynucleotide according to the fourth aspect of the present invention, or the nucleic acid molecule according to the fifth aspect of the present invention, or the vector according to the eighth aspect of the present invention.

[0125] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).

[0126] In another preferred embodiment, the host cell is a prokaryotic cell, such as Escherichia coli.

[0127] In another preferred embodiment, the yeast cell is selected from yeast of one or more sources in the following group: Pichia pastoris, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.

[0128] In another preferred embodiment, the host cell is selected from the following group: Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.

[0129] The tenth aspect of the present invention provides a CRISPR-Cas composition, comprising:

[0130] (i) A first component, selected from the following group: the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, a nucleotide sequence encoding the protein described in the first aspect of the present invention or the protein variant described in the second aspect of the present invention or the fusion protein described in the third aspect of the present invention, and any combination thereof; and

[0131] (ii) A second component, which is a nucleotide sequence comprising one or more guide RNAs described in the sixth aspect of the present invention, or a nucleotide sequence encoding the nucleotide sequence comprising one or more guide RNAs described in the sixth aspect of the present invention;

[0132] The guide RNA can form a complex with the protein or protein variant or fusion protein described in (i).

[0133] In another preferred embodiment, the guide RNA comprises a direct repeat sequence and a spacer sequence from the 5' to 3' direction, and the spacer sequence can hybridize with a target sequence.

[0134] In another preferred embodiment, the direct repeat sequence is the nucleic acid molecule defined in the fifth aspect of the present invention.

[0135] In another preferred embodiment, the composition further comprises a pharmaceutically acceptable carrier.

[0136] In another preferred embodiment, the composition comprises a pharmaceutical composition.

[0137] In another preferred embodiment, the dosage form of the composition is selected from the following group: freeze-dried preparation, liquid preparation, or a combination thereof.

[0138] In another preferred embodiment, the dosage form of the composition is a liquid preparation.

[0139] In another preferred example, the dosage form of the composition is an injection dosage form.

[0140] In another preferred example, the composition is a cell preparation.

[0141] The eleventh aspect of the present invention provides a CRISPR-Cas system, comprising one or more vectors, and the one or more vectors comprise:

[0142] (i) A first nucleic acid, which is a nucleotide sequence encoding the protein described in the first aspect of the present invention, or the protein variant described in the second aspect of the present invention, or the fusion protein described in the third aspect of the present invention; optionally, the first nucleic acid is operably linked to a first regulatory element; and

[0143] (ii) A second nucleic acid, which encodes a nucleotide sequence comprising the guide RNA described in the sixth aspect of the present invention; optionally, the second nucleic acid is operably linked to a second regulatory element;

[0144] Wherein:

[0145] The first nucleic acid and the second nucleic acid are present on the same or different vectors;

[0146] The guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

[0147] In another preferred example, the vector includes a plasmid and a viral vector.

[0148] In another preferred example, the guide RNA includes a spacer sequence capable of hybridizing with a target sequence; and a direct repeat (DR) sequence that is linked to the spacer sequence and is capable of guiding the protein to bind to the guide RNA, thereby forming a CRISPR-Cas composition or complex targeting the target sequence.

[0149] In another preferred example, the guide RNA includes unmodified and modified guide RNAs.

[0150] In another preferred example, the modified guide RNA includes chemical modifications of bases.

[0151] In another preferred example, the chemical modification includes methylation modification, methoxy modification, fluorination modification or thiolation modification.

[0152] In another preferred example, the direct repeat sequence is the nucleic acid molecule defined in claim 5.

[0153] In another preferred example, the first regulatory element and / or the second regulatory element is a promoter, such as an inducible promoter.

[0154] In another preferred embodiment, at least one component in the composition is non-naturally occurring or modified.

[0155] In another preferred embodiment, the spacer sequence is linked to the 3'-end of the direct repeat (DR) sequence.

[0156] In another preferred embodiment, the spacer sequence contains a complementary sequence of the target sequence.

[0157] In another preferred embodiment, when the target sequence is DNA, the target sequence is located at the 3'-end of the protospacer adjacent motif (PAM), and the PAM has a sequence shown as 5'-PAM being TTN, where N is A, T, C or G.

[0158] In another preferred embodiment, the target sequence is DNA from a prokaryotic cell or a eukaryotic cell or a DNA sequence formed by reverse transcription based on RNA; alternatively, the target sequence is non-naturally occurring DNA or a DNA sequence formed by reverse transcription based on RNA.

[0159] In another preferred embodiment, the target sequence includes a cDNA sequence.

[0160] In another preferred embodiment, the target sequence includes single-stranded DNA, double-stranded DNA sequences.

[0161] In another preferred embodiment, the target sequence exists inside the cell.

[0162] In another preferred embodiment, the target sequence exists in the nucleus or in the cytoplasm (e.g., organelles).

[0163] In another preferred embodiment, the cell is a eukaryotic cell.

[0164] In another preferred embodiment, the cell is a prokaryotic cell.

[0165] In another preferred embodiment, the target sequence exists outside the cell.

[0166] In another preferred embodiment, the protein according to the first aspect of the present invention is linked with one or more NLS sequences, or the fusion protein contains one or more NLS sequences.

[0167] In another preferred embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the protein according to the first aspect of the present invention.

[0168] In another preferred embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein according to the first aspect of the present invention.

[0169] The twelfth aspect of the present invention provides a kit, comprising one or more components selected from the following: the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention, or the system described in the eleventh aspect of the present invention.

[0170] In another preferred example, the kit further comprises a label or an instruction manual.

[0171] In another preferred example, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, cleaving a target gene or a non-target gene.

[0172] The thirteenth aspect of the present invention provides a delivery composition, characterized by comprising a delivery vector and one or more selected from the following: the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention, or the system described in the eleventh aspect of the present invention.

[0173] In another preferred example, the delivery vector is a particle.

[0174] In another preferred example, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microbubbles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0175] The fourteenth aspect of the present invention provides an enzyme preparation, which comprises the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the complex described in the seventh aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention, or the system described in the eleventh aspect of the present invention, or the delivery composition described in the thirteenth aspect of the present invention.

[0176] In another preferred example, the enzyme preparation comprises an injection and / or a freeze-dried preparation.

[0177] 15. A medicine box, characterized by comprising:

[0178] A first container, and the complex according to the seventh aspect of the present invention, or the composition according to the tenth aspect of the present invention, or the system according to the eleventh aspect of the present invention, which are located in the first container, or a drug containing the complex according to the seventh aspect of the present invention, or the composition according to the tenth aspect of the present invention, or the system according to the eleventh aspect of the present invention.

[0179] In another preferred example, the drug in the first container is a single-agent preparation containing the complex according to the seventh aspect of the present invention, or the composition according to the tenth aspect of the present invention, or the system according to the tenth aspect of the present invention.

[0180] In another preferred example, the dosage form of the drug is selected from the group consisting of: freeze-dried preparation, liquid preparation, or a combination thereof.

[0181] In another preferred example, the dosage form of the drug is an oral dosage form or an injection dosage form.

[0182] In another preferred example, the medicine box further contains an instruction manual.

[0183] The sixteenth aspect of the present invention provides a medicine box, comprising:

[0184] (a1) A first container, and the protein according to the first aspect of the present invention, or the protein variant according to the second aspect of the present invention, or the fusion protein according to the third aspect of the present invention, or its coding gene or its expression vector, which are located in the first container, or a drug containing the protein according to the first aspect of the present invention, or the protein variant according to the second aspect of the present invention, or the fusion protein according to the third aspect of the present invention, or its coding gene or its expression vector;

[0185] (b1) Optionally, a second container, and the guide RNA according to the sixth aspect of the present invention or its expression vector, which are located in the second container, or a drug containing the guide RNA according to the sixth aspect of the present invention or its expression vector.

[0186] In another preferred example, the first container and the second container are different containers.

[0187] In another preferred example, the drug in the first container is a single-agent preparation containing the protein according to the first aspect of the present invention, or the protein variant according to the second aspect of the present invention, or the fusion protein according to the third aspect of the present invention, or its coding gene or its expression vector.

[0188] In another preferred example, the drug in the second container is a single-agent preparation containing the guide RNA according to the sixth aspect of the present invention or its expression vector.

[0189] In another preferred example, the dosage form of the drug is selected from the group consisting of: freeze-dried preparation, liquid preparation, or a combination thereof.

[0190] In another preferred example, the dosage form of the drug is an oral dosage form or an injection dosage form.

[0191] In another preferred example, the medicine box further contains an instruction manual.

[0192] The seventeenth aspect of the present invention provides a method for targeting and editing a target gene or cleaving a target gene, which is characterized by comprising: contacting the protein described in the first aspect of the present invention, or the protein variant described in the second aspect of the present invention, or the fusion protein described in the third aspect of the present invention, or the complex described in the seventh aspect of the present invention, or the composition described in the tenth aspect of the present invention, or the system described in the eleventh aspect of the present invention, or the delivery composition described in the thirteenth aspect of the present invention, or the enzyme preparation described in the fourteenth aspect of the present invention, or the medicine box described in the fifteenth aspect or the sixteenth aspect of the present invention with the target gene, or delivering it into a cell containing the target gene, and the target sequence is present in the target gene.

[0193] In another preferred example, the target gene is present inside the cell.

[0194] In another preferred example, the cell is a prokaryotic cell.

[0195] In another preferred example, the cell is a eukaryotic cell, such as a mammalian cell (such as a human cell) or a plant cell.

[0196] In another preferred example, the target gene is present in a nucleic acid molecule (for example, a plasmid) in vitro.

[0197] In another preferred example, editing the target gene or cleaving the target gene includes cleavage of the target sequence, such as double-strand break of DNA or single-strand break of RNA, or inserting an exogenous nucleic acid into the cleavage.

[0198] In another preferred example, the target gene includes DNA.

[0199] In another preferred example, the DNA includes single-stranded DNA and double-stranded DNA.

[0200] The eighteenth aspect of the present invention provides a method for inducing a change in cell state, and the method includes contacting the protein described in the first aspect of the present invention, or the protein variant described in the second aspect of the present invention, or the fusion protein described in the third aspect of the present invention, or the complex described in the seventh aspect of the present invention, or the composition described in the tenth aspect of the present invention, or the system described in the eleventh aspect of the present invention, or the delivery composition described in the thirteenth aspect of the present invention, or the enzyme preparation described in the fourteenth aspect of the present invention, or the medicine box described in the fifteenth aspect or the sixteenth aspect of the present invention with the target gene in the cell.

[0201] The nineteenth aspect of the present invention provides a method for altering the expression of a gene product, comprising: contacting the protein described in the first aspect of the present invention, or the protein variant described in the second aspect of the present invention, or the fusion protein described in the third aspect of the present invention, or the complex described in the seventh aspect of the present invention, or the composition described in the tenth aspect of the present invention, or the system described in the eleventh aspect of the present invention, or the delivery composition described in the thirteenth aspect of the present invention, or the enzyme preparation described in the fourteenth aspect of the present invention, or the kit described in the fifteenth or sixteenth aspect of the present invention with a nucleic acid molecule encoding the gene product, or delivering it into a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0202] In another preferred example, the nucleic acid molecule is present in an in vitro nucleic acid molecule (e.g., plasmid).

[0203] In another preferred example, the expression of the gene product is altered (e.g., enhanced or reduced).

[0204] In another preferred example, the gene product is a protein.

[0205] In another preferred example, the protein, fusion protein, polynucleotide, isolated nucleic acid molecule, complex, vector or composition is contained in a delivery vector.

[0206] In another preferred example, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0207] In another preferred example, one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product are altered to modify a cell, cell line or organism.

[0208] The twentieth aspect of the present invention provides a cell or its progeny obtained by the method according to any one of the seventeenth to nineteenth aspects of the present invention, wherein the cell contains a modification that does not exist in its wild type.

[0209] The twenty-first aspect of the present invention provides a cell product of the cell or its progeny described in the twentieth aspect of the present invention.

[0210] The twenty-second aspect of the present invention provides an in vitro, ex vivo or in vivo cell or cell line or their progeny, which cell or cell line or their progeny comprises: the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention, the system described in the eleventh aspect of the present invention or the delivery composition described in the thirteenth aspect of the present invention.

[0211] In another preferred example, the cell is a prokaryotic cell.

[0212] In another preferred example, the cell is a eukaryotic cell, such as a mammalian cell (such as a human cell) or a plant cell.

[0213] In another preferred example, the cell is a stem cell or a stem cell line.

[0214] The twenty-third aspect of the present invention provides the use of the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the nucleic acid molecule described in the fifth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention, the system described in the eleventh aspect of the present invention, the kit described in the twelfth aspect of the present invention, the delivery composition described in the thirteenth aspect of the present invention, the enzyme preparation described in the fourteenth aspect of the present invention or the kit described in the fifteenth or sixteenth aspect of the present invention, for preparing a drug or a preparation for nucleic acid editing (e.g., gene or genome editing).

[0215] In another preferred example, the gene or genome editing includes modifying a gene, knocking out a gene, altering the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide.

[0216] The twenty-fourth aspect of the present invention provides the use of the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the complex described in the seventh aspect of the present invention, the vector described in the eighth aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention, the system described in the eleventh aspect of the present invention, the kit described in the twelfth aspect of the present invention, the delivery composition described in the thirteenth aspect of the present invention, the enzyme preparation described in the fourteenth aspect of the present invention or the kit described in the fifteenth or sixteenth aspect of the present invention, for preparing a drug or a preparation for one or more selected from the group consisting of:

[0217] (i) Ex vivo gene or genome editing;

[0218] (ii) Detection of single-stranded DNA ex vivo;

[0219] (iii) Editing a target sequence at a target locus to modify a biological or non-human organism;

[0220] (iv) Treating a disorder caused by a defect in a target sequence at a target locus;

[0221] (v) Treating a disorder or disease in a subject in need thereof.

[0222] In another preferred example, the disorder or disease includes cancer, infectious diseases, neurological diseases, ophthalmic diseases, and hearing diseases.

[0223] In another preferred example, the disease or disorder includes cystic fibrosis, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy, alpha-1-antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell disease, beta-thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, hypercholesterolemia, transthyretin amyloidosis, retinal diseases, macular degeneration, Wilms tumor, Ewing sarcoma, neuroendocrine tumors, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and urinary bladder cancer.

[0224] In another preferred example, the disorder or disease is caused by a pathogenic point mutation.

[0225] The twenty-fifth aspect of the present invention provides a method for detecting whether a target nucleic acid molecule is present in a sample, the method comprising contacting the sample with the protein described in the first aspect of the present invention, the protein variant described in the second aspect of the present invention, the fusion protein described in the third aspect of the present invention, or the complex described in the seventh aspect of the present invention, the CRISPR-Cas composition described in the tenth aspect of the present invention or the system described in the eleventh aspect of the present invention, the kit described in the twelfth aspect of the present invention or the delivery composition described in the thirteenth aspect of the present invention or the enzyme preparation described in the fourteenth aspect of the present invention and a non-target sequence, and detecting a detectable signal generated by cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA.

[0226] In another preferred embodiment, if the non-target sequence is cleaved by the protein in the complex or CRISPR-Cas composition or system or delivery composition, it indicates that the target nucleic acid molecule is present in the sample; while if the non-target sequence is not cleaved by the protein in the complex or CRISPR-Cas composition or system or delivery composition, it indicates that the target nucleic acid molecule is not present in the sample.

[0227] In another preferred embodiment, the target nucleic acid molecule is target DNA.

[0228] In another preferred embodiment, the target DNA includes DNA formed by reverse transcription of RNA.

[0229] In another preferred embodiment, the target DNA includes cDNA.

[0230] In another preferred embodiment, the target DNA is selected from the group consisting of single-stranded DNA, double-stranded DNA, or a combination thereof.

[0231] It should be understood that within the scope of the present invention, the above technical features of the present invention and the technical features specifically described below (such as in the examples) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be repeated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS

[0232] Figure 1 shows the recombinant expression plasmid of CasY6 ( Figure 1A ) and the recombinant expression map of LbCpf1 ( Figure 1B ).

[0233] Figure 2 shows the prediction of the secondary structure of the direct repeat (DR) sequence corresponding to CasY6.

[0234] Figure 3 shows the Target plasmid map.

[0235] Figure 4Shows the comparison of the editing efficiency of CasY6 and LbCpf1 in Escherichia coli.

[0236] Figure 5 Shows the plasmid map of PHK09T.

[0237] Figure 6 Shows the comparison of the editing efficiency of CasY6 and LbCpf1 in HEK293T cells.

[0238] Figure 7 Shows four mutants of CasY6, which eliminate the catalytic activity (cleavage activity) compared to the wild-type CasY6 protein.

[0239] Figure 8 shows the activity detection of dCasY6 in base editing, where Figure 8A Shows the single-base editing (A>G) efficiency of the base editor composed of dCasY6 at the A2, A4, A13, A15 - A17 sites; Figure 8B Shows the single-base editing (A>G) efficiency of the base editor composed of dCasY6 at A7, A13, A20. Detailed implementation mode

[0240] After extensive and in - depth research, the inventor of the present invention unexpectedly discovered a new Cas protein for the first time. The Cas protein of the present invention has very good gene editing activity, can effectively edit or cleave target genes, and can effectively treat the diseases or disorders of subjects in need (for example, diseases or disorders caused by pathogenic point mutations, including cystic fibrosis, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy, α - 1 - antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell disease, β - thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, hypercholesterolemia, transthyretin amyloidosis, retinal diseases, macular degeneration, Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non - Hodgkin lymphoma, and urinary bladder cancer); compared with the Cas enzymes disclosed in the prior art, the Cas protein of the present invention has an advantage in editing efficiency, providing more choices for base editing tools; the base editor constructed by the Cas enzyme disclosed in the present invention can effectively perform base editing and has potential application prospects. On this basis, the inventor of the present invention completed the present invention.

[0241] Term

[0242] The following examples are only used to describe the present invention and do not limit the present invention. Unless otherwise specified, the experiments and methods described in the examples are basically carried out according to the conventional methods well - known in the art and described in various reference documents.

[0243] In addition, for those conditions not specified in the examples, they are carried out according to the conventional conditions or the conditions recommended by the manufacturer. For reagents or instruments whose manufacturers are not indicated, they are all conventional products that can be obtained through commercial purchase. Those skilled in the art know that the examples describe the present invention by way of example and are not intended to limit the scope claimed by the present invention. All the published cases and other reference materials mentioned herein are incorporated herein by reference in their entirety.

[0244] To more easily understand the present disclosure, certain terms are first defined. As used in this application, unless otherwise clearly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.

[0245] The term "about" can refer to a value or a component within an acceptable error range of a specific value or component determined by a person of ordinary skill in the art, which will depend in part on how the value or component is measured or determined. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0246] As used herein, the terms "comprising" or "including" can be open-ended, semi-closed, and closed. In other words, the terms also include "consisting essentially of", or "consisting of".

[0247] Sequence identity (or homology) is determined by comparing two aligned sequences along a predefined comparison window, which can be 50%, 60%, 70%, 80%, 90%, 95% or 100% of the length of the reference nucleotide sequence or protein, and determining the number of positions at which identical residues occur. Typically, this is expressed as a percentage. Methods for measuring sequence identity of nucleotide sequences are well known to those skilled in the art.

[0248] Cas protein

[0249] In the present invention, Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably, and Cas protein is taken in its broadest sense, including wild-type Cas protein, its derivatives or variants, analogs, and its functional fragments such as oligonucleotide-binding fragments.

[0250] The term "wild-type" has the meaning commonly understood by those skilled in the art, which refers to the typical form of a biological organism, strain, gene, or protein, or the characteristics that distinguish it from mutant or variant forms when it exists in nature, and it can be isolated from natural sources and has not been deliberately modified artificially.

[0251] The terms "variant", "derivative", and "analog" refer to polypeptides that substantially retain the function or activity of the Cas protein of the present invention.

[0252] Generally, derivatization of a protein does not adversely affect the desired activity of the protein (e.g., the activity of binding to a guide RNA, endonuclease activity, the activity of binding to a specific site of a target sequence and cleaving under the guidance of a guide RNA), that is, the derivative of the protein has the same activity as the protein. The modified forms of "derivative" include that one or more amino acids of the protein can be deleted, inserted, modified, and / or substituted. The terms "non-naturally occurring" or "engineered" can be used interchangeably and indicate the involvement of human intervention.

[0253] In one aspect, the present invention provides a Cas protein comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to the amino acid sequence of SEQ ID NO.1 and substantially retaining the biological function of the sequence from which it is derived;

[0254] In one embodiment, the amino acid sequence of the Cas protein has a sequence with one or more amino acid substitutions, deletions or additions as compared to the amino acid sequence of SEQ ID NO.1 and substantially retains the biological function of the sequence from which it is derived;

[0255] In one embodiment, the Cas protein comprises the amino acid sequence shown in SEQ ID NO.1;

[0256] Or a sequence having one or more amino acid substitutions, deletions or additions (such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) as compared to the sequence shown in SEQ ID NO.1; or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence shown in SEQ ID NO.1;

[0257] In one embodiment, the protein has the amino acid sequence shown in SEQ ID NO.1.

[0258] Those skilled in the art will appreciate that the structure of a protein can be altered without adversely affecting its activity and functionality, for example, by introducing one or more conservative amino acid substitutions in the amino acid sequence of the protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule.

[0259] Those skilled in the art are aware of examples and embodiments of conservative amino acid substitutions. Specifically, an amino acid residue can be replaced with another amino acid residue belonging to the same group as the site to be replaced, i.e., a nonpolar amino acid residue is replaced with another nonpolar amino acid residue, a polar uncharged amino acid residue is replaced with another polar uncharged amino acid residue, a basic amino acid residue is replaced with another basic amino acid residue, and an acidic amino acid residue is replaced with another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. As long as the substitution does not result in the inactivation of the biological activity of the protein, conservative substitutions in which one amino acid is replaced with another amino acid belonging to the same group fall within the scope of the present invention. Thus, the proteins of the present invention may contain one or more conservative substitutions in the amino acid sequence, preferably made according to Table A. Additionally, the present invention also encompasses proteins that further contain one or more other non-conservative substitutions, provided that such non-conservative substitutions do not significantly affect the desired function and biological activity of the proteins of the present invention.

[0260] Conservative amino acid substitutions can be made at one or more predicted non-essential amino acid residues. A "non-essential" amino acid residue is an amino acid residue that can be altered (deleted, substituted, or permuted) without changing biological activity, while an "essential" amino acid residue is required for biological activity. A "conservative amino acid substitution" is a substitution in which an amino acid residue is replaced with an amino acid residue having a similar side chain. Amino acid substitutions can be made in non-conserved regions of the Cas enzyme. In general, such substitutions are not made to conserved amino acid residues or to amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conservative or non-conservative changes in conserved regions.

[0261] Table A

[0262]

[0263] Those skilled in the art are aware that changing (substituting, deleting, truncating, or inserting) one or more amino acid residues from the N and / or C terminus of a protein can still retain its functional activity. Thus, proteins in which one or more amino acid residues have been changed from the N and / or C terminus of the Cas proteins of the present invention while retaining their desired functional activity are also within the scope of the present invention. These changes can include those introduced by modern molecular methods such as PCR, which include PCR amplification that modifies or extends the protein coding sequence by including amino acid coding sequences within the oligonucleotides used in the PCR amplification.

[0264] It should be recognized that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known to those skilled in the art.

[0265] For example, amino acid sequence variants of Cas proteins can be prepared by mutagenesis of DNA. It can also be accomplished by other forms of mutagenesis and / or by directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, in combination with relevant screening methods, to effect one or more amino acid substitutions; or one or more amino acid deletions and / or one or more amino acid insertions.

[0266] Those skilled in the art will appreciate that these minor amino acid changes in the Cas proteins of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations present are not close to the catalytic domain, active site, or other functional domains, a lesser impact can be expected.

[0267] Those skilled in the art can identify the essential amino acids of the Cas protein according to methods known in the art, such as site-directed mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domain, active site, or other functional domains of the protein can also be determined by physical analysis of the structure, such as by techniques including nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, in combination with mutagenesis of putative key-site amino acids.

[0268] Orthologue, ortholog

[0269] As used herein, the term "orthologue, ortholog" has the meaning generally understood by those skilled in the art. As further guidance, an "orthologue" of a protein as described herein refers to a protein from a different species that performs the same or similar function as the protein of which it is an orthologue.

[0270] The nucleic acid cleavage of the present invention includes: DNA or RNA breaks in the target nucleic acid generated by the Cas protein (Cis cleavage), and DNA or RNA breaks in the collateral nucleic acid substrate (single-stranded nucleic acid substrate) caused by the collateral activity of the Cas protein (i.e., non-specific or non-targeted cleavage, Trans cleavage or collateral activity). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0271] Trans-cleavage refers to the situation where, in certain environments, the activated Cas12 family proteins remain active after binding to the target sequence and continue to non-specifically cleave non-target oligonucleotides. This collateral cleavage activity can be used to detect the presence of specific target oligonucleotides using the Cas system. For example, the Cas12i system has been engineered to non-specifically cleave ssDNA or transcripts. The collateral cleavage activity is used in a highly sensitive and specific nucleic acid detection platform called SHERLOCK, which can be used for many clinical diagnoses (Gootenberg, J.S. et al., Nucleic acid detection with CRISPR-Cas13a / C2c2. Science 356, 438-442 (2017)).

[0272] Fusion protein

[0273] In one aspect, the present invention provides a fusion protein comprising the Cas protein according to any one of the preceding claims and one or more functional domains.

[0274] In one embodiment, the functional domain comprises one or more of a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methyltransferase, a demethylase, a transcriptional release factor, an HDAC, a cleavage activity polypeptide, a ligase;

[0275] In one example, the "methyltransferase", for example, HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3, ZMET2, CMT1, CMT2, etc.

[0276] "Demethylase" refers to an enzyme that removes a methyl (CH3-) group from nucleic acids, proteins (e.g., histones), and other molecules. Demethylases are important in epigenetic modification mechanisms. Demethylase proteins alter the transcriptional regulation of the genome by controlling the methylation levels that occur on DNA and histones, and thereby regulate the chromatin state at specific loci in an organism, such as TET1 (ten-eleven translocation 1), ten-eleven translocation (TET) dioxygenase 1 (TET1CD), DME, DML1, DML2, ROS1, etc.

[0277] In another preferred example, the transcriptional release factor, for example, eukaryotic release factor 1 (ERF1) activity, eukaryotic release factor 3 (ERF3).

[0278] In one embodiment, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

[0279] In one embodiment, the localization signal includes a nuclear localization signal and / or a nuclear export signal;

[0280] Preferably, the nuclear export signal includes human protein tyrosine kinase 2;

[0281] Preferably, the reporter protein includes one or more of glutathione-S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, or a self-fluorescent protein;

[0282] Preferably, the self-fluorescent protein includes one or more of green fluorescent protein, HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, or blue fluorescent protein;

[0283] Preferably, the DNA binding domain includes one or more of a methylation binding protein, LexA DBD, or Gal4 DBD;

[0284] Preferably, the epitope tag includes one or more of a histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, or thioredoxin tag;

[0285] Preferably, the transcriptional activation domain includes VP64 and / or VPR;

[0286] Preferably, the transcriptional repression domain includes KRAB and / or SID;

[0287] Preferably, the nuclease includes FokI;

[0288] Preferably, the deamination domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TadA;

[0289] Preferably, the cleavage active polypeptide includes a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity;

[0290] Preferably, the ligase includes a DNA ligase and / or an RNA ligase.

[0291] In one embodiment, the functional domain is the full length or a functional fragment of TadA8e.

[0292] Polynucleotide

[0293] In one aspect, the present invention provides a polynucleotide, which is a polynucleotide sequence encoding the Cas protein or a polynucleotide sequence encoding the aforementioned fusion protein.

[0294] In one embodiment, the polynucleotide is a DNA molecule codon-optimized according to the codon preference of the host cell;

[0295] In one embodiment, the host cell includes a prokaryotic cell or a eukaryotic cell;

[0296] In one embodiment, the DNA molecule includes nucleotides having more than 70%, preferably more than 90%, more preferably more than 95%, further preferably 99%, and even more preferably 100% identity with the nucleotide sequence described in any one of SEQ ID NO.2.

[0297] CRISPR system

[0298] The term "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" is used interchangeably and has the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of guiding the activity of the Cas genes.

[0299] CRISPR-Cas composition

[0300] In one aspect, the present invention further provides a CRISPR-Cas composition, which comprises:

[0301] (1) Protein component: the aforementioned Cas protein, or the aforementioned fusion protein; or a nucleic acid molecule encoding the Cas protein or the fusion protein;

[0302] (2) RNA component: guide RNA, or one or more nucleic acids encoding the guide RNA, or precursor RNA of the guide RNA, or a nucleic acid encoding the precursor RNA of the guide RNA;

[0303] The protein component and the nucleic acid component bind to each other to form a complex.

[0304] In one embodiment, the composition is an activated CRISPR complex, and the activated CRISPR complex further comprises: a target sequence of a target nucleic acid bound to the guide RNA.

[0305] In one embodiment, the CRISPR-Cas composition includes one or more vectors, and the one or more vectors comprise:

[0306] (1) A first regulatory element operably linked to a nucleotide sequence encoding the Cas protein or a nucleotide sequence encoding the fusion protein; and

[0307] (2) A second regulatory element operably linked to a nucleotide sequence encoding the guide RNA, and the guide RNA comprises:

[0308] (a) A spacer sequence capable of hybridizing with a target sequence of a target nucleic acid, and

[0309] (b) A direct repeat (DR) sequence linked to the spacer sequence and capable of guiding the Cas protein to bind to the guide RNA to form a CRISPR-Cas complex targeting the target sequence;

[0310] wherein the first regulatory element and the second regulatory element are located on the same or different vectors of the CRISPR-Cas vector system.

[0311] In one embodiment, the first regulatory element or the second regulatory element comprises a promoter, and the promoter comprises one or more of an inducible promoter, a constitutive promoter or a tissue-specific promoter;

[0312] In one embodiment, the promoter comprises one or more of T7, SP6, T3, CMV, EF1a, SV40, PGK1, human β-actin, CAG, U6, H1, T7, T7lac, araBAD, trp, lac or Ptac;

[0313] In one embodiment, the first regulatory element and the second regulatory element are located on the same or different vectors.

[0314] In one embodiment, the vector comprises a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral vector, a herpes simplex vector or a phagemid vector;

[0315] In one embodiment, the vector comprises a plasmid vector.

[0316] In one embodiment, the target nucleic acid comprises DNA derived from a eukaryote or DNA derived from a prokaryote;

[0317] In one embodiment, the eukaryote comprises an animal or a plant;

[0318] In one embodiment, the target nucleic acid includes non-human mammalian DNA, human DNA, insect DNA, avian DNA, reptilian DNA, amphibian DNA, rodent DNA, fish DNA, worm DNA, nematode DNA, or yeast DNA;

[0319] In one embodiment, the non-human mammalian DNA includes non-human primate DNA.

[0320] CRISPR / Cas complex

[0321] The term "CRISPR / Cas complex" refers to a complex formed by the binding of a gRNA (guide RNA) or a mature crRNA (or guide RNA) to a Cas protein, which contains a direct repeat sequence hybridized to a guide sequence of a target sequence and bound to the Cas protein, and the complex is capable of recognizing and cleaving a target nucleotide that can hybridize to the guide RNA or mature crRNA.

[0322] Guide RNA (gRNA, guide RNA)

[0323] The terms "guide RNA (guide RNA, gRNA)", "mature crRNA", "crRNA", "guide sequence", "guide RNA" are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA may include a direct repeat (DR) sequence and a spacer sequence, or consist essentially of or consist of a direct repeat (DR) sequence and a spacer sequence.

[0324] In some cases, the spacer sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize with the target sequence and direct the specific binding of the CRISPR-Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the spacer sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. The guide sequence contains a sequence (such as a direct repeat (DR) sequence) that has sufficient complementarity to a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct the sequence-specific binding of the complex to the target nucleic acid sequence.

[0325] It is known in the art that, based on sufficient complementarity to function, complete complementarity is not required. Thus, if desired, modulation of cleavage efficiency can be achieved by introducing mismatches (e.g., one or more mismatches between a spacer sequence and a target nucleic acid, such as a mismatch of 1 or 2 nucleotides (including the position of the mismatch along the spacer sequence / target sequence)). For example, if a cleavage rate of less than 100% of the target is desired (e.g., in a cell population), 1 or 2 mismatches between the spacer sequence and the target sequence can be introduced into the spacer sequence.

[0326] In one aspect, the present invention provides a guide RNA, which includes a direct repeat (DR) sequence capable of binding to the Cas protein and a spacer sequence capable of targeting a target sequence.

[0327] In one embodiment, the direct repeat (DR) sequence contains the sequence shown in SEQ ID NO.5.

[0328] In one embodiment, the 3' end of the direct repeat sequence contains a stem-loop structure, and further includes a stem formed by hybridization of a first stem nucleotide chain and a second stem nucleotide chain to form the stem-loop structure, and a loop nucleotide chain forms the loop of the stem-loop structure;

[0329] In one embodiment, the direct repeat sequence includes a nucleotide sequence having at least 80% identity with the nucleotide sequence shown in SEQ ID NO.5;

[0330] In one embodiment, the direct repeat sequence includes a nucleotide sequence having at least 85% or more, more preferably 90% or more, and further preferably 95% or more identity with the nucleotide sequence shown in SEQ ID NO.5;

[0331] In one embodiment, the direct repeat sequence includes the nucleotide sequence shown in SEQ ID NO.5.

[0332] In one embodiment, more than 80% of the spacer sequence is complementary to the target nucleic acid;

[0333] In one embodiment, more than 90%, more preferably more than 95%, further preferably more than 99%, and even more preferably 100% of the spacer sequence is complementary to the target nucleic acid;

[0334] In one embodiment, the length of the spacer sequence is 18 - 41 nt;

[0335] In one embodiment, the length of the spacer sequence is 20 nt.

[0336] Target nucleic acid

[0337] In the present invention, the target nucleic acid is used interchangeably with the target sequence, the target nucleic acid sequence, or the target nucleic acid molecule, and refers to a specific nucleic acid that contains a nucleic acid sequence that is fully or partially complementary to the spacer sequence in the guide RNA. The "target sequence" refers to the polynucleotide targeted by the spacer sequence in the guide RNA, such as a sequence that is complementary to the spacer sequence, where hybridization between the target sequence and the spacer sequence will promote the formation of a CRISPR-Cas complex (including the Cas protein and the guide RNA). Perfect complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR-Cas complex. In some embodiments, the target nucleic acid contains non-coding regions (e.g., promoters or terminators). In some embodiments, the target nucleic acid is single-stranded or double-stranded.

[0338] The target sequence can contain any polynucleotide, such as DNA. In certain cases, the target sequence is located intracellularly or extracellularly. In certain cases, the target sequence is located within the nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts) of the cell.

[0339] The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In certain cases, the target sequence should be related to the protospacer adjacent motif (PAM).

[0340] Donor template

[0341] In the present invention, the donor template nucleic acid or the donor template is used interchangeably, and refers to a nucleic acid molecule that one or more cellular proteins can use to alter the structure of the target nucleic acid after the Cas protein described herein has altered the target nucleic acid.

[0342] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid or a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some instances, the donor template nucleic acid is an exogenous nucleic acid molecule. In some instances, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, gene recombination can be achieved using the donor template, and the recombination is homologous recombination.

[0343] Cleavage

[0344] Cleavage refers to a DNA break in the target nucleic acid generated by the Cas protein described herein. In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break.

[0345] In the present invention, the meanings of cleaving a target nucleic acid or modifying a target nucleic acid may overlap. Modifying a target nucleic acid includes not only modification of a single nucleotide, but also insertion or deletion of a nucleic acid fragment.

[0346] Reporter nucleic acid

[0347] A reporter nucleic acid refers to a molecule that can be cleaved by the activated CRISPR system proteins described herein or otherwise inactivated. The reporter nucleic acid contains a nucleic acid element that can be cleaved by a CRISPR protein (e.g., using a single-stranded non-targeting nucleic acid molecule that includes different reporter groups or labeling molecules at both ends). Cleavage of the nucleic acid element generates a detectable signal. Before cleavage, or when the reporter nucleic acid is in an "active" state, the reporter nucleic acid prevents the generation or detection of a positive detectable signal. It will be understood that in certain exemplary embodiments, a minimal background signal may be generated in the presence of an active reporter nucleic acid. The positive detectable signal can be any signal that can be detected using optical, fluorescent, chemiluminescent, electrochemical, or other detection methods known in the art. For example, in certain embodiments, a first signal (i.e., a negative detectable signal) can be detected when the reporter nucleic acid is present, and then it is converted to a second signal (e.g., a positive detectable signal) after detection of the target molecule and cleavage or inactivation by the activated CRISPR protein. The reporter nucleic acid can be a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.

[0348] The detection method described in the present invention can be used for quantitative detection of a target nucleic acid to be detected. The quantitative detection index can be quantified according to the signal strength of the reporter group, such as according to the luminescence intensity of a fluorescent group, or according to the width of a color development band, etc.

[0349] Functional domain

[0350] In the present invention, the functional domain is taken in its broadest sense and includes a protein such as an enzyme or a factor itself or its specific functional fragment / domain. A Cas protein (e.g., a dCas protein) is linked / associated with one or more functional domains selected from a localization signal, a reporter protein, a Cas protein targeting moiety, a DNA binding domain, an epitope tag, a transcriptional activation domain, a transcriptional repression domain, a nuclease, a deamination domain, a methylase, a demethylase, a transcriptional release factor, an HDAC, a cleavage active polypeptide, a ligase, etc. When more than one functional domain is included, the functional domains can be the same or different.

[0351] Deamination domain

[0352] In the present invention, the deamination domain includes a catalytic domain of a deaminase (such as adenosine deaminase or cytidine deaminase). As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts adenine (or the adenine moiety of a molecule) to hypoxanthine (or the hypoxanthine moiety of a molecule), as shown below. In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0353] Adenosine deaminases include, but are not limited to, members of the enzyme family called adenosine deaminases acting on RNA (ADAR), members of the enzyme family called adenosine deaminases acting on tRNA (ADAT), and other family members containing an adenosine deaminase domain (ADAD). According to the present disclosure, adenosine deaminases are capable of targeting adenine in RNA / DNA and RNA duplexes. In certain embodiments, the adenosine deaminase has been modified to increase its ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.

[0354] In some embodiments, the deaminase is a cytidine deaminase. The term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts cytosine (or the cytosine moiety of a molecule) to uracil (or the uracil moiety of a molecule). In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0355] Cytidine deaminases include, but are not limited to, members of the enzyme family called apolipoprotein B mRNA editing complex (APOBEC) family deaminases, activation-induced deaminase (AID), or cytidine deaminase 1 (CDA1). In certain embodiments, APOBEC family deaminases.

[0356] Identity

[0357] "Identity" is used to refer to the sequence match between two polypeptides or between two nucleic acids. "Identity" represents the percentage of the number of identical residues between the polypeptide or nucleic acid sequences out of the total number of residues, and the calculation of the total number of residues is determined based on the type of mutation. Types of mutations include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / replacements of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence.

[0358] Taking a polypeptide sequence as an example, if the mutation type is one or more of the following: substitution / replacement of one or more amino acids / nucleotides, insertion within the sequence, and deletion within the sequence, the total number of residues is calculated based on the larger of the two molecules being compared. If the mutation type also includes insertion (extension) at either or both ends of the sequence or deletion (truncation) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (e.g., the number inserted or deleted at both ends is less than 20) is not included in the total number of residues. When calculating the percentage identity, the sequences being compared are aligned in a way that produces the maximum match between the sequences, and gaps in the alignment (if any) are resolved by a specific algorithm. The calculation of nucleotide identity is the same.

[0359] vector

[0360] A vector is a nucleic acid molecule capable of transporting another nucleic acid molecule linked to it.

[0361] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, no free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and other diverse polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection, enabling the genetic material elements it carries to be expressed in the host cell. A vector can be introduced into a host cell and thereby produce transcripts, proteins, or peptides, including those derived from proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain various elements for controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. A vector can also contain an origin of replication.

[0362] Vectors include plasmids and viral vectors. A plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. In a viral vector, virus-derived DNA or RNA sequences are present in the vector used for packaging the virus. Viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. A viral vector also contains the polynucleotide carried by the virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.

[0363] Other vectors (e.g., non-integrating mammalian vectors) integrate into the genome of the host cell after introduction into the host cell and are thereby replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to as "expression vectors".

[0364] In some embodiments, a vector (e.g., a viral vector or a non-viral vector, such as a lentiviral vector or a plasmid) can be delivered to a target tissue by, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. The above delivery can be carried out via a single dose or multiple doses. Those skilled in the art will understand that the actual dose to be delivered herein can vary to a large extent depending on a variety of factors, including but not limited to vector selection, target cells, organisms, tissues, the general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the mode of administration, and the type of transformation / modification sought.

[0365] Regulatory elements

[0366] "Regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression control elements (e.g., transcriptional termination signals, such as polyadenylation signals, poly-U sequences), the detailed description of which can be found in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or development stage-dependent manner), which may or may not be tissue or cell type specific.

[0367] "Promoter" refers to a non-coding nucleotide sequence located upstream of a gene that can initiate the expression of downstream genes. A constitutive promoter is such a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, will result in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, such as in response to a chemical compound (chemical inducer), or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulatable promoters include, for example, promoters induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounding, or chemicals (such as ethanol, abscisic acid (ABA), jasmonates, salicylic acid, or safeners).

[0368] Host cell

[0369] "Host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism (e.g., a cell line) cultured in the form of a single cell entity, which is used as a recipient of nucleic acid (e.g., an expression vector), and includes the progeny of the original cell that has been genetically modified by nucleic acid.

[0370] It should be understood that the progeny of a single cell may be attributable to natural, accidental, or deliberate mutations and may not necessarily have exactly the same morphology or genome as the original parental cell. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.

[0371] Those skilled in the art will understand that the design of an expression vector can depend on factors such as the choice of host cell to be transformed, the desired level of expression, etc.

[0372] In another aspect, the present invention also provides a host cell or its progeny, wherein the host cell comprises the foregoing Cas protein, or the foregoing fusion protein, or the foregoing polynucleotide, or the foregoing vector system, or the foregoing CRISPR-Cas system, or the foregoing composition.

[0373] In one embodiment, the host cell includes a non-human mammal, a human, an insect, a bird, a reptile, an amphibian, a rodent, a fish, a worm, a nematode, or a yeast cell.

[0374] In one aspect, the present invention also provides a multicellular organism, wherein the multicellular organism comprises the foregoing cell or its progeny.

[0375] In one embodiment, the multicellular organism is an animal model or a plant model for related diseases.

[0376] NLS

[0377] NLS refers to "nuclear localization sequence" or "nuclear localization signal", which is an amino acid sequence that promotes the entry of a protein into the cell nucleus. Nuclear localization sequences are known in the art (for example, as described in International PCT Application PCT / EP2000 / 011690 filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001), and the patent is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the following amino acid sequences: KRTADGSEFESPKKKRKV (SEQ ID NO.35), AVKRPAATKKAGQAKKKKLD (SEQ ID NO.36), KRPAATKKAGQAKKKK (SEQ ID NO.37), KKTELQTTNAENKTKKL (SEQ ID NO.38), KRGINDRNFWRGENGRKTR (SEQ ID NO.39), RKSGKIAAIVVKRPRK (SEQ ID NO.40), PKKKRKV (SEQ ID NO.41), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO.42).

[0378] Operably linked

[0379] "Operably linked" means that a target nucleotide sequence is linked to a regulatory element in a manner that permits expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the types of these vectors can also be selected to target specific types of cells.

[0380] Complementary

[0381] "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. The percentage of complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., if 5, 6, 7, 8, 9, or 10 out of 10 are complementary, the percentage of complementarity is 50%, 60%, 70%, 80%, 90%, and 100%). "Fully complementary" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.

[0382] The term "stringent conditions" in relation to hybridization refers to conditions under which a nucleic acid that is complementary to a target sequence hybridizes predominantly to the target sequence and essentially not to non-target sequences. Stringent conditions are generally sequence-dependent and depend on many factors. In general, the longer the sequence, the higher the temperature at which the sequence hybridizes specifically to its target sequence.

[0383] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding of the bases between these nucleotide residues. The complex can comprise two strands that form a duplex, three or more strands that form a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. The hybridization reaction can constitute a step in a broader process (such as the start of PCR, or the cleavage of a polynucleotide by an enzyme). A sequence that can hybridize to a given sequence is called the "complement" of the given sequence.

[0384] Hybridization of the target sequence with the gRNA means that the nucleic acid sequences of the target sequence and the gRNA can hybridize at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% to form a complex; or it represents that at least 12, 15, 16, 17, 18, 19, 20, or more bases of the nucleic acid sequences of the target sequence and the gRNA can be complementary paired to hybridize and form a complex.

[0385] Expression

[0386] Nucleic acid expression includes one or more of generating an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5′ capping, and / or 3′ end processing), translation of the RNA into a polypeptide or protein, or post-translational modification of the polypeptide or protein.

[0387] Delivery

[0388] "Delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as a combination of DNA / RNA or RNA / RNA or protein-RNA. For example, the Cas protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.

[0389] In one aspect, the present invention also provides a delivery system, which includes the Cas protein or the fusion protein, or the polynucleotide, or the CRISPR-Cas composition described above.

[0390] In one embodiment, the delivery system further includes a delivery vehicle, and the delivery vehicle includes nanoparticles, liposomes, exosomes, microbubbles, gene guns, or electroporation devices.

[0391] In addition, when the delivery target is a plant cell, a delivery method such as using a cell-penetrating peptide (CPP) is also adopted. For example, in a specific embodiment, the Cas protein and / or at least one guide RNA are conjugated with one or more CPPs, so as to effectively transport the CPP conjugated with the Cas protein and / or the guide RNA into the plant cell (e.g., into the protoplast). The CPP is a short peptide with less than 35 amino acids, which is derived from a protein or a chimeric sequence and can transport biomolecules across the cell membrane in a non-receptor-dependent manner. The CPP can be a cationic peptide, a peptide with a hydrophobic sequence, an amphiphilic peptide, a peptide rich in proline and antimicrobial sequences, and a chimeric or bipartite peptide. The CPP can penetrate biological membranes, and thus trigger the movement of different biomolecules across the cell membrane into the cytoplasm, and can improve their intracellular pathways, and thus promote the interaction between the biomolecules and the target.

[0392] Exemplarily, the CPPs include Tat (a nuclear transcriptional activator protein required for viral replication by human immunodeficiency virus type 1), penetratin, Kaposi fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, sweet arrow peptide, etc.

[0393] Linker

[0394] "Linker" refers to a linear polypeptide formed by the connection of multiple amino acid residues through peptide bonds. The linker can be an artificially synthesized amino acid sequence or a naturally occurring polypeptide sequence.

[0395] Detection

[0396] In one aspect, the present invention also provides a method for targeting and editing a target nucleic acid, the method comprising contacting the target nucleic acid with any one of the foregoing CRISPR-Cas systems or compositions.

[0397] In one aspect, the present invention also provides a method for non-specifically degrading single-stranded DNA after identifying a target nucleic acid, the method comprising contacting the target nucleic acid with the foregoing CRISPR-Cas composition.

[0398] In one aspect, the present invention also provides a method for targeting the non-Spacer complementary strand of a double-stranded target nucleic acid and introducing a nick therein after identifying the Spacer complementary strand of the double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with any one of the foregoing CRISPR-Cas systems or compositions.

[0399] In one aspect, the present invention also provides a method for targeting and cleaving a double-stranded target nucleic acid, the method comprising contacting the double-stranded target nucleic acid with any one of the foregoing CRISPR-Cas systems or compositions.

[0400] In one embodiment, before introducing a nick in the Spacer complementary strand of the double-stranded DNA, a nick is introduced in the non-Spacer sequence complementary strand of the double-stranded target nucleic acid.

[0401] In one aspect, the present invention also provides a method for specifically editing a double-stranded nucleic acid, the method comprising contacting the following for a sufficient amount of time under sufficient conditions:

[0402] (1) The foregoing Cas protein, or a fusion protein, or another enzyme having sequence-specific nicking activity, and the guide RNA, which guides the Cas protein or the fusion protein to introduce nicks in the opposing strands relative to the activity of the other sequence-specific nicking enzyme; and (2) the double-stranded nucleic acid; the method results in the formation of a double-strand break.

[0403] In one aspect, the present invention also provides a method for editing a double-stranded nucleic acid, the method comprising contacting the following for a sufficient amount of time under sufficient conditions:

[0404] (1) The foregoing Cas protein, or a fusion protein, and a fusion protein having a protein domain with DNA modification activity, and the RNA guide targeting the double-stranded nucleic acid; and (2) the double-stranded nucleic acid;

[0405] The Cas protein of the fusion protein is modified to introduce a nick in the non-target strand of the double-stranded nucleic acid.

[0406] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at different sites, resulting in staggered cleavage.

[0407] In one embodiment, the two strands of the double-stranded nucleic acid are cleaved at the same site, resulting in a blunt double-strand break.

[0408] In one aspect, the present invention also provides a method for targeting and cleaving a single-stranded target nucleic acid, the method comprising contacting the target nucleic acid with a CRISPR-Cas composition as described in any one of the preceding claims.

[0409] In one aspect, the present invention also provides a method for inducing a change in cell state, the method comprising contacting the aforementioned CRISPR-Cas composition with the target nucleic acid in the cell.

[0410] In one embodiment, the cell state includes apoptosis or dormancy;

[0411] In one embodiment, the cell includes a eukaryotic cell or a prokaryotic cell;

[0412] In one embodiment, the cell includes a mammalian cell or a plant diseased cell;

[0413] In one embodiment, the cell includes a cancer cell;

[0414] In one embodiment, the cell includes an infectious cell or a cell infected with an infectious agent;

[0415] In one embodiment, the cell includes a cell infected with a virus, a cell infected with a prion;

[0416] In one embodiment, the cell includes a fungal cell, a protozoan or a parasite cell.

[0417] In one aspect, the present invention also provides a method for detecting a target nucleic acid in a sample, the method comprising contacting the sample with the aforementioned Cas protein, guide RNA and non-target sequence; detecting a detectable signal generated by the Cas protein cleaving the non-target sequence, thereby detecting the target nucleic acid; the non-target sequence does not hybridize with the guide RNA.

[0418] Kit

[0419] In one aspect, the present invention provides a kit, which includes the aforementioned Cas protein, the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned CRISPR-Cas composition, and the use of the aforementioned host cell in the preparation of the kit, and the components of the kit are in the same or different containers.

[0420] In one aspect, the present invention further provides a container that contains the aforementioned kit.

[0421] In one embodiment, the container includes a sterile container;

[0422] In one embodiment, the container includes a syringe.

[0423] In some embodiments, the kit further includes instructions for using the kit, such as instructions in more than one language. The kit may also contain one or more reagents for use in the process using one or more of the above components. The reagents can be provided in any suitable container. For example, the kit can provide one or more reaction or storage buffers. The above reagents can be provided in a form that requires the addition of one or more other components before use (e.g., in concentrated or lyophilized form); the buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. The buffer can have a suitable pH value (pH), for example, it can be alkaline. In some embodiments, the pH of the buffer is between about 7 and 10.

[0424] Treatment

[0425] "Treatment" means treating or curing a subject's disease, delaying the onset of symptoms of the disease, and / or delaying the severity of the disease. The term "subject" includes but is not limited to various animals, plants, and microorganisms. Animals include mammals, such as bovines, equines, ovines, porcines, canines, felines, lagomorphs (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) has a disease (e.g., a disease caused by a defective disease-related gene). A "plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any mature or developmental stage.

[0426] In one aspect, the present invention further provides the use of the aforementioned Cas protein, the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned CRISPR-Cas composition, and the aforementioned host cell in the preparation of a medicament for treating a disease or disorder of a subject in need thereof.

[0427] In one embodiment, the application includes administering the CRISPR-Cas composition to the subject or to the isolated cells of the subject;

[0428] In one embodiment, the spacer sequence is complementary to at least 15 nucleotides of the target nucleic acid associated with the disorder or disease, and the Cas protein or the fusion protein cleaves the target nucleic acid;

[0429] In one embodiment, the disorder or disease includes cancer or an infectious disease;

[0430] In one embodiment, the cancer includes one or more of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, or urinary bladder cancer;

[0431] In one embodiment, the disorder or disease includes one or more of cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell disease, hypercholesterolemia, transthyretin amyloidosis, or β-thalassemia;

[0432] In one embodiment, the infectious agent of the infectious disease includes one or more of human immunodeficiency virus, herpes simplex virus-1, or herpes simplex virus-2.

[0433] The main advantages of the present invention include:

[0434] (a) The present invention for the first time discovers a new Cas protein. The Cas protein of the present invention has very good gene editing activity, can effectively edit or cleave target genes, and can effectively treat the diseases or disorders of subjects in need (such as one or more of cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell disease, hypercholesterolemia, transthyretin amyloidosis or β-thalassemia).

[0435] (b) The Cas protein of the present invention is a completely new Cas enzyme, which exhibits good nuclease activity in vivo and in vitro and has broad application prospects.

[0436] (c) Compared with the Cas enzymes disclosed in the prior art, the editing efficiency of the Cas protein of the present invention has advantages, providing more choices for base editing tools.

[0437] (d) The base editor constructed by the Cas enzyme disclosed in the present invention can effectively perform base editing and has potential application prospects.

[0438] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are usually carried out under conventional conditions, such as the conditions described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and weight parts.

[0439] Unless otherwise specified, the reagents and materials in the embodiments of the present invention are all commercially available products.

[0440] Example 1 Obtaining of Cas Protein

[0441] The inventors analyzed the metagenome of uncultured organisms, and through analysis such as dereplication and protein clustering, 1 new Cas protein was identified. The Blast analysis results showed that the sequence identity of the Cas protein with the reported Cas proteins was relatively low, and it was named CasY6 in the present invention.

[0442] The amino acid sequences of the above-mentioned CasY6 proteins are respectively shown in SEQ ID NO.1, and the nucleotide sequences after human codon optimization are shown in SEQ ID NO.2.

[0443] Analysis of the direct repeat (DR) sequence of the guide RNA corresponding to the above CasY6 protein showed that:

[0444] The DNA sequence encoding the direct repeat (DR) sequence of the guide RNA corresponding to the CasY6 protein is:

[0445] TATCCATCGTGCCGCCTCTTGGCAC (SEQ ID NO: 5).

[0446] The inventors further analyzed the RNA secondary structure of the DR sequence in pre-crRNA using RNAfold. The analysis results are as Figure 2 shown. It was found that the PAM corresponding to CasY6 is TTN, where N is A / T / C / G. The sgRNA (also known as crRNA) sequence of the CasY6 protein consists of a spacer sequence and a direct repeat (DR) sequence.

[0447] After identification, CasY6 of the present invention belongs to the Cas12 protein family.

[0448] Example 2. Verification of the cleavage activity of Cas protein

[0449] 1. Plasmid construction

[0450] (1) Using the TTR gene as the target, a spacer sequence was designed according to the target sequence of the target gene TTR gene: GCATCTCCCCATTCCATGAG (SEQ ID NO.6).

[0451] According to the DR sequences of the CasY6 protein and the LbCpf1 protein, sgRNA sequences targeting the TTR target gene were designed, as shown in the following table:

[0452]

[0453] According to the requirements of vector expression, a T7 promoter and an rrnB T2 terminator were added to the 5' end and 3' end of the respective sgRNA sequences of CasY6 and LbCpf1 to obtain the CasY6-sgRNA expression fragment sequence:

[0454] Among them, the single-underlined sequence part is the CasY6DR sequence, the double-underlined sequence part is the spacer sequence, the italicized sequence part is the T7 promoter, the wavy-underlined sequence part is the rrnB T2 terminator sequence, the linker sequence is between the spacer sequence and the rrnB T2 terminator sequence, the dotted sequence part is the MfeI restriction site, the bold sequence part is the MluI restriction site, and CACCG is the linker.

[0455] In the same way, the LbCpf1-sgRNA expression fragment sequence was synthesized:

[0456]

[0457] Among them, the single-underlined sequence part is the LbCpf1DR sequence, the double-underlined sequence part is the spacer sequence, the italicized sequence part is the T7 promoter, the wavy-underlined sequence part is the rrnB T2 terminator sequence, the linker sequence is between the spacer sequence and the rrnB T2 terminator sequence, the dotted sequence part is the MfeI restriction site, the bold sequence part is the MluI restriction site, and CACCG is the linker.

[0458] To protect the sequence integrity, when synthesizing the CasY6-sgRNA expression fragment sequence and the LbCpfl-sgRNA expression fragment sequence, AGC was introduced at the 5' end of the sequence and ATA was introduced at the 3' end as protective bases.

[0459] (2) The coding nucleotide sequence (SEQ ID NO.2) fragment of CasY6 protein was synthesized by Suzhou Genewiz Biotechnology Co., Ltd., and the synthesized coding nucleotide sequence fragment of CasY6 protein was constructed into the ABE8e plasmid (Addgene, Plasmid #138489) at positions 466-5160 to obtain the CasY6 recombinant expression plasmid (for the plasmid map, see Figure 1A ).

[0460] The optimized coding nucleotide sequence (SEQ ID NO.13) of LbCpf1 was synthesized by Suzhou Genewiz Biotechnology Co., Ltd., and constructed in the same way to obtain the LbCpf1 recombinant expression plasmid (for the plasmid map, see Figure 1B ).

[0461] (3) Suzhou Hongxun Biotechnology Co., Ltd. synthesized the sgRNA expression sequence fragments as described in step (1) (CasY6-sgRNA expression fragment sequence and LbCpf1-sgRNA expression fragment sequence), and the sgRNA expression fragment sequence was double-digested (MfeI / MluI) and inserted into the CasY6 recombinant expression plasmid vector that was also double-digested (MfeI / MluI) to obtain a recombinant expression plasmid CasY6+sgRNA expression plasmid expressing CasY6 and sgRNA; the LbCpf1+sgRNA expression plasmid was constructed in the same way.

[0462] (4) Construct a Target plasmid with a targeting sequence. The construction process is as follows:

[0463] Suzhou Hongxun Biotechnology Co., Ltd. synthesized the araC-pBAD-CCDB fragment (SEQ ID NO.11) with the TTR target sequence (SEQ ID NO.6), and inserted the araC-pBAD-CCDB fragment into the 1284-1300 site of the pKESK22 (Addgene, Plasmid #64857) plasmid to obtain the Target plasmid. The sequence of the Target plasmid is shown in SEQ ID NO.4, and the plasmid map is shown in Figure 3 .

[0464] 2. Preparation and transformation of E. coli competent cells

[0465] The Target plasmid was transferred into DH5a competent cells, and the cells were separated and inoculated on LB solid medium containing 50 μg / ml kanamycin sulfate using an inoculation loop. The cells were placed in a biochemical incubator and cultured overnight at 37°C. On the second day, a single colony was picked from the plate and inoculated into a LB liquid medium test tube containing 4 ml 50 μg / ml kanamycin sulfate (Sangon Biotech, A100408-0100), and cultured overnight at 37°C and 200 rpm with shaking. The next day, 4 ml of the bacterial solution was inoculated into a 2L shake flask containing 400 ml 50 μg / ml kanamycin sulfate LB liquid medium, and cultured at 37°C and 200 rpm with shaking for 2-3 hours.

[0466] When the OD600nm value of the bacterial solution reaches 0.3 - 0.5, take out the shaking flask and place it on ice for 10 - 15 min. Under sterile conditions, pour the bacterial solution into a pre-cooled 500 ml centrifuge bottle, centrifuge at 4°C and 3000 rpm for 8 min, discard the supernatant, add about 200 ml of pre-cooled CaCl2 solution, pipette and mix well to suspend the bacteria, and place it on ice for 30 min. Then centrifuge the bacterial solution at 4°C and 3000 rpm for 8 min, discard the supernatant, add about 8 ml of pre-cooled CaCl2 solution, re-suspend the bacteria, and aliquot the re-suspended bacteria into 1.5 ml EP tubes, 110 μl per tube, and store them in a -80°C ultra-low temperature freezer for later use.

[0467] 3. Determination of in vivo editing efficiency of Escherichia coli

[0468] Transfer the CasY6 + sgRNA expression plasmid and the LbCpf1 + sgRNA expression plasmid into the competent cells prepared in step 2 respectively. The specific procedure is as follows:

[0469] (1) Take out the competent cells from -80°C, quickly insert them into ice. After about 5 minutes, the bacterial mass melts. Add the CasY6 + sgRNA expression plasmid, and then gently mix by flicking the bottom of the centrifuge tube with your finger. Let it stand in ice for 25 minutes. Heat shock in a 42°C water bath for 45 seconds, quickly put it back into ice and let it stand for 2 minutes. Add 900 μl of sterile LB medium without antibiotics to the centrifuge tube, mix well, and resuscitate at 37°C and 220 rpm for 60 min. Take 100 μl of the bacterial solution respectively and spread them on LB agar plates containing 30 μg / ml carbenicillin resistance (Sangon Biotech, A100358 - 0001) (abbreviated as C-LB medium) and LB agar plates containing both 30 μg / ml carbenicillin resistance and 10 mM L-arabinose (Sangon Biotech, A610071 - 0100) (abbreviated as CL-LB medium). Invert the two LB agar plates and place them in an incubator and culture overnight at 37°C.

[0470] (2) Detection of in vivo editing efficiency of Escherichia coli

[0471] As Figure 3 shown, the Target plasmid carries the PBAD promoter that can be induced by L-arabinose, the CCDB gene regulated by the PBAD promoter, and the CCDB gene can express the CCDB toxic protein. The CCDB toxic protein, as a DNA gyrase inhibitor, can lock the DNA gyrase and the broken double-stranded DNA complex, making the DNA gyrase unable to function, and ultimately leading to cell death.

[0472] Based on this, the inventors designed a method for detecting the in vivo editing efficiency of Escherichia coli:

[0473] In the presence of L - arabinose in the culture medium, if the CasY6 protein or LbCpf1 protein, under the guidance of sgRNA, can specifically target the target sequence (SEQ ID NO.6) of the TTR gene on the Target plasmid and exert a cleavage effect, then the regulatory expression pathway of the PBAD promoter for the CCDB toxic protein is cut off, and the host cell survives because it does not produce the ccdB toxic protein; conversely, if the CasY6 protein or LbCpf1 protein cannot specifically target the TTR target sequence on the Target plasmid, the host cell Escherichia coli will die due to the PBAD promoter induced by L - arabinose regulating the expression of the CCDB gene to produce the CCDB toxic protein.

[0474] Therefore, the editing efficiency of the CasY6 protein in specifically targeting and cleaving the TTR target gene in Escherichia coli can be calculated according to the ratio of the number of bacterial colonies on the CL - LB medium to the number of bacterial colonies on the C - LB medium in step (2).

[0475] The results are as Figure 4 shown. After counting the number of Escherichia coli colonies and calculating the ratio, the editing efficiency of the CasY6 protein is 16.2%, and the editing efficiency of the LbCpf1 protein is 6.4%. The editing efficiency of the CasY6 protein is significantly higher than that of the LbCpf1 protein.

[0476] Example 3 Detection of Editing Efficiency in HEK293T Cells

[0477] 1. Construction of TTR - sgRNA Expression Plasmid

[0478] (1) Design the TTR - sgRNA sequence according to the target sequence (SEQ ID NO.6) of the TTR gene and synthesize oligonucleotides (oligos):

[0479] CasY6 - TTR - sgRNA sequence:

[0480] TATCCATCGTGCCGCCTCTTGGCAC GCATCTCCCCATTCCATGAG (SEQ ID NO.8), the underlined part of the sequence is the DR sequence, and the rest of the sequence is the spacer sequence.

[0481] LbCpf1 - TTR - sgRNA sequence:

[0482] TAATTTCTACTAAGTGTAGAT GCATCTCCCCATTCCATGAG (SEQ ID NO.15), the underlined part of the sequence is the DR sequence, and the rest of the sequence is the spacer sequence.

[0483] (2) Add the CACC sequence to the 5' end of the upstream sequence of TTR-sgRNA and the AAAA sequence to the 5' end of the downstream sequence, and synthesize oligos. The specific sequences are as follows:

[0484]

[0485] After the upstream and downstream primers of the aforementioned TTR-sgRNA are synthesized, annealing is carried out through a preset program (95°C, 5 min; from 95°C to 85°C at -2°C / s; from 85°C to 25°C at -0.1°C / s; maintained at 4°C). Then the annealed product is ligated to the PHK09T vector linearized by BsmBI (NEB, #R0580L). The sequence of the PHK09T vector is shown in SEQ ID NO.3. See the plasmid map in Figure 5 .

[0486] The linearization of the PHK09T vector and its ligation method with the annealed product of TTR-sgRNA are as follows:

[0487] First, linearize the PHK09T vector. Linearization system: 3 μg of PHK09T vector; 6 μL of buffer (NEB: R0539L); 2 μL of BsmBI; supplemented with ddH2O to 60 μL, and digest with enzymes at 50°C overnight.

[0488] Ligation system of the annealed product of TTR-gRNA and the linearized vector:

[0489] 1 μL of T4 ligase buffer (NEB, #M0202L), 20 ng of linearized vector, 5 μL of annealed oligo fragment, 0.5 μL of T4 ligase (NEB, #M0202L), supplemented with ddH2O to 10 μL, and ligate at 16°C overnight to obtain the CasY6-TTR-sgRNA expression plasmid and the LbCpf1-TTR-sgRNA expression plasmid.

[0490] (3) Transfer the CasY6-TTR-sgRNA expression plasmid and the LbCpf1-TTR-sgRNA expression plasmid obtained in step (2) into Escherichia coli DH5a competent cells (Vidy Biotechnology, DL1001). The specific steps are as follows:

[0491] Take the DH5α competent cells out of the -80°C refrigerator and quickly insert them into ice. After 5 minutes, wait for the bacterial mass to thaw, add the ligation product, and gently mix by flicking the bottom of the centrifuge tube with your finger. Let it stand in ice for 25 minutes. Heat shock in a 42°C water bath for 45 seconds, quickly put it back into ice and let it stand for 2 minutes. Add 700 μl of sterile LB medium without antibiotics to the centrifuge tube, mix well, and recover at 37°C and 200 rpm for 60 minutes. Centrifuge at 3000 rpm for 1 minute to collect the bacteria, leave about 100 μl of the supernatant, gently resuspend the bacterial mass by pipetting, and spread it on the LB medium with Amp antibiotic. Invert the plate and incubate it overnight in a 37°C incubator. Pick a single colony, after confirmation by sequencing, shake the positive clone and extract the plasmid (using an endotoxin-free large plasmid extraction kit, TIANGEN: DP120-01), then measure the concentration and store it in a -20°C refrigerator for later use.

[0492] 2. Detection of editing efficiency at the cellular level

[0493] (1) Culture of HEK293T cells

[0494] Inoculate HEK293T cells (purchased from ATCC) into DMEM medium supplemented with 10% FBS (v / v) (Gibco, 11965092), which contains 1% Penicillin Streptomycin (v / v) (Gibco, 15140122), and culture them in a 37°C cell incubator containing 5% CO2. For the cells to be transfected, inoculate them into a 24-well cell culture plate the day before and culture them. Observe the cells the next day, and when the cell density reaches about 80%, perform transfection.

[0495] (2) The CasY6 recombinant expression plasmid (see the plasmid map in Figure 1A ), the CasY6-TTR-sgRNA expression plasmid, the LbCpf1 recombinant expression plasmid (see the plasmid map in Figure 1B ), and the LbCpf1-TTR-sgRNA expression plasmid are transfected into HEK293T cells together with the EGFP-C1 (Addgene, Plasmid, #54759) plasmid respectively.

[0496] The amount of plasmid used for transfection of each well in the 24-well plate is 0.3 μg of the nuclease expression plasmid (CasY6 recombinant expression plasmid or LbCpf1 recombinant expression plasmid), 0.3 μg of the sgRNA expression plasmid (CasY6-TTR-sgRNA expression plasmid or LbCpf1-TTR-sgRNA expression plasmid), and 0.3 μg of the EGFP-C1 plasmid. The specific transfection operation is as follows:

[0497] Mix the CasY6 expression plasmid, CasY6-TTR-sgRNA expression plasmid, and EGFP-C1 plasmid respectively, and dilute them with 25 μl of serum-free transfection medium (Yuanpei Biotech, L530KJ), then add 2 μl of Lipofectamine 3000 (Invitrogen, L3000015) reagent, pipette and mix well to obtain reagent A, and let it stand for 5 minutes. At the same time, dilute 2 μl of Lipofectamine 3000 transfection reagent (Invitrogen, L3000015) with 25 μl of serum-free transfection medium (Yuanpei Biotech, L530KJ) and mix well to obtain reagent B, and let it stand for 5 minutes.

[0498] Mix the above reagent A and reagent B and pipette evenly, then let it stand for 20 minutes. After standing, add the mixed reagent drop by drop to the cells in the 24-well plate to be transfected, and place it back in the incubator at 37 °C and 5% CO2 for culture. After 6 hours of transfection, change the medium to DMEM medium containing 10% FBS.

[0499] In the same way, transfect the LbCpf1 recombinant expression plasmid, LbCpf1-TTR-sgRNA expression plasmid, and EGFP-C1 plasmid into HEK293T cells.

[0500] (3) Detection of editing efficiency

[0501] After 48 hours of transfection, the expression of EGFP fluorescent protein indicates successful cell transfection. Sort the cells with positive EGFP expression for the detection of editing efficiency. Extract the genome of the cells (using the genomic DNA extraction kit, TIANGEN, DP304-03). Design identification primers according to experimental requirements. The sequences of the identification primers used are shown in the following table:

[0502] Primer Name Specific Sequence TTR-F aactgaggaggaatttgtag(SEQ ID NO.9) TTR-R caaaagcaaaaaccaaaacc(SEQ ID NO.10)

[0503] Using the genome as a template, perform PCR amplification of the sequence near the target site with the primers in the above table. The PCR amplification system is as follows:

[0504] 2×Taq Master Mix (Vazyme, P112-03) 25 μL; Primer-F (TTR-F) (10 pmol / μL) 1 μL; Primer-R (TTR-R) (10 pmol / μL) 1 μL; template 1 μL; ddH2O to make up to 50 μL.

[0505] The amplified PCR products were used for identification of editing efficiency by high-throughput deep sequencing (Genewiz Biotechnology Co., Ltd.) or Sanger sequencing (Bio-Pharm Technology (Shanghai) Co., Ltd.).

[0506] By detecting and identifying the editing efficiencies of CasY6 and LbCpf1, the results are as Figure 6 shown. In 293T cells, the editing efficiency of CasY6 was 25%, while that of LbCpf1 was only 18%. The editing efficiency of CasY6 protein was much higher than that of LbCpf1.

[0507] Example 4 Application of CasY6 in Base Editing

[0508] (1) Obtaining of Catalytically Inactive CasY6

[0509] To obtain catalytically inactive (i.e., losing cleavage activity) dCasY6, the inventors constructed CasY6 mutants with single-site mutations of D659A, D711A, E895A, and D1069A: D659A-dCasY6, D711A-dCasY6, E895A-dCasY6, D1069A-dCasY6. The specific construction method is as follows:

[0510] The CasY6+sgRNA expression plasmid obtained in step 1(3) of Example 2 was subjected to site-directed mutagenesis to modify the amino acids at 4 sites of CasY6, namely aspartic acid (Asp, D) at position 659, aspartic acid (Asp, D) at position 711, glutamic acid (Glu, E) at position 895, and aspartic acid (Asp, D) at position 1069 of SEQ ID No.1, and the amino acids at the above sites were mutated to alanine (Ala, A). The codons before and after the mutation of each amino acid are shown in the following table:

[0511] Amino Acid before Mutation Codon Amino Acid after Mutation Codon Aspartic Acid (Asp, D) GAC Alanine (Ala, A) GCA Aspartic Acid (Asp, D) GAC Alanine (Ala, A) GCA Glutamic Acid (Glu, E) GAG Alanine (Ala, A) GCA Aspartic Acid (Asp, D) GAC Alanine (Ala, A) GCA

[0512] Forward and reverse primers were designed and synthesized for the amino acids and their codons in the above table respectively, and then PCR amplification was carried out using the CasY6+sgRNA expression plasmid as a template. After amplification, the amplified products were recovered and purified using a universal DNA purification and recovery kit (Tiangen Biochemical Technology (Beijing) Co., Ltd., DP214). The purified products were transformed into Escherichia coli Dh5a competent cells (Vidy Biotechnology, DL1001), cultured overnight at 37°C, and single colonies were picked and sent for sequencing the next day. After confirmation by sequencing, the positive clones were shaken and the plasmids were extracted (TIANGEN, DP120-01), and then the concentration was measured and stored at -20°C in the refrigerator for later use.

[0513] The recombinant plasmids of the obtained point mutations were respectively named: D659A-dCasY6+sgRNA expression plasmid, D711A-dCasY6+sgRNA expression plasmid, E895A-dCasY6+sgRNA expression plasmid, D1069A-dCasY6+sgRNA expression plasmid.

[0514] Afterwards, the constructed D659A-dCasY6+sgRNA expression plasmid, D711A-dCasY6+sgRNA expression plasmid, E895A-dCasY6+sgRNA expression plasmid, and D1069A-dCasY6+sgRNA expression plasmid were respectively subjected to in vivo editing efficiency detection in Escherichia coli. The detection method and calculation method were the same as those in Steps 2 and 3 of Example 2. The experimental results are as Figure 7 shown. By counting the number of Escherichia coli clones and calculating the ratio, it was considered that D659A-dCasY6, D711A-dCasY6, E895A-dCasY6, and D1069A-dCasY6 lost catalytic activity (cleavage activity). That is to say, the point mutations of D659A, D711A, E895A, and D1069A caused the CasY6 protein to lose cleavage activity.

[0515] 2. Base editing efficiency detection at the cellular level

[0516] (1) Construction of "CasY6-TTR-sgRNA" and "CasY6-TTR-sgRNA" plasmids

[0517] sgRNAs: CasY6-TTR-sgRNA' and CasY6-TTR-sgRNA" sequences were designed according to the target sequence of the TTR gene, and oligonucleotides (oligos) were synthesized:

[0518] CasY6-TTR-sgRNA':

[0519] TATCCATCGTGCCGCCTCTTGGCAC tatatcccttctacaaattc (SEQ ID NO.20);

[0520] CasY6-TTR-sgRNA":

[0521] TATCCATCGTGCCGCCTCTTGGCAC gtgtctatttccactttgta (SEQ ID NO.21), where the underlined part of the sequence is the DR sequence, and the remaining sequence is the spacer sequence.

[0522] (2) Add the CACC sequence to the 5' end of the upstream sequence of each sgRNA and the AAAA sequence to the 5' end of the downstream sequence. The specific form is as follows:

[0523]

[0524] According to the method of step 1 in Example 3, the upstream and downstream sequences of CasY6-TTR-sgRNA' and CasY6-TTR-sgRNA" were annealed and connected to the PHK09T vector to obtain CasY6-TTR-sgRNA' expression plasmid and CasY6-TTR-sgRNA" expression plasmid, which were transferred into Escherichia coli DH5a competent cells for plasmid amplification and culture. After correct sequencing and concentration determination, they were stored for later use.

[0525] (3) Construction of base editor plasmid (using 005V1-10-3 as an example)

[0526] The adenosine deaminase catalytic domain selected by the inventors is selected from the mutant of the amino acid sequence shown in SEQ ID NO: 28 (named 005V1-10-3): Q148G+Q149M+P150R, the amino acid sequence of the mutant is shown in SEQ ID NO.29, and the nucleotide sequence encoding the deaminase 005V1-10-3 is shown in SEQ ID NO.30. A base editor fusion protein consisting of deaminase 005V10-3 and CasY6 protein was constructed by homologous recombination. The specific operation is as follows:

[0527] First, Suzhou Hongxun Biotechnology Co., Ltd. synthesized the 005V1-10-3 nucleotide fragment with homology arm sequence and linker:

[0528] The bold part is the 005V1-10-3 nucleotide sequence, the italic part is the left and right homology arm regions, and the wavy line part is the linker sequence.

[0529] The D1069A-dCasY6+sgRNA expression plasmid obtained in step (1) was subjected to PCR amplification for linearization to obtain a linearized expression vector. The primers used are shown in the following table:

[0530] Primer Name Specific Sequence D1069A-dCasY6-F ATCAAGAACCAAATCATCGG(SEQ ID NO:32) D1069A-dCasY6-R Gactttccgcttcttctttgg(SEQ ID NO:33)

[0531] The nucleotide fragment of 005V1-10-3 with homology arm sequence and linker (SEQ ID NO: 31) and D1069A-dCasY6+sgRNA linearized expression vector were homologously recombined, and Gibson Assembly Master Mix (NEB, E2611S) was used for reaction. After the reaction, the ligation product was transformed into Escherichia coli DH5a competent cells (Weidi Biotechnology, DL1001). The specific process is as follows:

[0532] After the DH5α competent cells were taken out of the -80°C refrigerator, they were quickly inserted into ice. After 5 minutes, when the bacterial mass melted, the ligation product was added and the bottom of the centrifuge tube was gently mixed by hand, and then left standing in ice for 25 minutes. Heat shock at 42°C in a water bath for 45 seconds, quickly put it back into ice and leave it standing for 2 minutes. Add 700 μl of sterile LB medium to the centrifuge tube, mix well, and resuscitate at 37°C and 200 rpm for 60 minutes. Centrifuge at 5000 rpm for 1 minute to collect the bacteria, leave about 100 μl of the supernatant, gently pipette to resuspend the bacterial mass and spread it on the LB medium containing Amp antibiotic. Invert the plate and incubate it overnight in a 37°C incubator. Pick a single colony, after sequencing confirmation, shake the positive clone, extract the base editor plasmid using an endotoxin-free plasmid large extraction kit (TIANGEN: DP120-01), then measure the concentration and store it for later use at -20°C. Among them, the nucleotide sequence encoding the base editor fusion protein 005V1-10-3-D1069A-dCasY6 is shown in SEQ ID NO: 34.

[0533] (4) According to the method in step 2 of Example 3, the base editor plasmid and the EGFP-C1 (Addgene, Plasmid#54759) plasmid were co-transfected into 293T cells with the CasY6-TTR-sgRNA’ and CasY6-TTR-sgRNA” expression plasmids respectively.

[0534] After 48 hours of transfection, the expression of EGFP fluorescent protein indicated successful cell transfection, and the cells with positive EGFP expression were sorted for the detection of editing efficiency. Extract the genome of the 293T cells using a kit (TIANGEN, DP304-03).

[0535] (5) Detect the base editing efficiency according to the method in step (3) of Example 3.

[0536] Design primers according to experimental requirements, and the sequences of the identification primers used are shown in the following table:

[0537] Primer Name Specific Sequence TTR-F’ gggtgtattactttgccatg(SEQ ID NO.26) TTR-R’ aacctttggtcattcatcaccttc(SEQ ID NO.27)

[0538] The results are as Figure 8A 、 8B shown. The base editor composed of D1069A-dCasY6 can achieve effective editing at multiple sites. As can be seen from Figure 8A , there are effective edits at the +2, +4, +13, +15, +16, +17 sites of CasY6-TTR-sgRNA’, and the editing efficiency at the +13, +15, +16, +17 sites reaches nearly 10% to nearly 30%.

[0539] From Figure 8BIt can be seen that effective editing exists at the +7, +13, and +20 sites of "CasY6-TTR-sgRNA". The base editing efficiency at the +13 site reaches more than 10%, and the editing efficiency at the +20 site reaches more than 15%.

[0540] Sequence information

[0541]

[0542]

[0543]

[0544]

[0545]

[0546]

[0547]

[0548]

[0549]

[0550]

[0551]

[0552]

[0553] All documents mentioned in the present invention are cited in this application as references, just as if each document was cited separately as a reference. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

Claims

1. A protein, characterized in that, The amino acid sequence of the said protein is as shown in SEQ ID NO:

1.

2. A protein variant, characterized in that, The said variant is a non-natural protein, and mutations occur at the core amino acid sites related to cleavage activity corresponding to SEQ ID NO: 1 of the wild-type protein, selected from the following groups: Aspartic acid (D) at position 659 is mutated to alanine (A); Aspartic acid (D) at position 711 is mutated to alanine (A); Glutamic acid (E) at position 895 is mutated to alanine (A); or Aspartic acid (D) at position 1069 is mutated to alanine (A).

3. A fusion protein, characterized in that, Comprising the protein according to claim 1 or the protein variant according to claim 2; and one or more functional domains.

4. The fusion protein according to claim 3, wherein The said functional domains are selected from localization signals, reporter proteins, Cas protein targeting moieties, DNA binding domains, epitope tags, transcriptional activation domains, transcriptional repression domains, nucleases, deamination domains, methylases, demethylases, transcriptional release factors, HDACs, cleavage activity polypeptides, ligases, integrases, transposases, recombinases, polymerases, and base excision repair inhibitors.

5. The fusion protein according to claim 3, wherein The said functional domains include one or more of the following enzymatic activities towards the target sequence: methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity, and deglycosylation activity.

6. The fusion protein according to claim 3 or 4, wherein the said functional domain is selected from a nuclear localization signal (NLS) and / or a nuclear export signal (NES).

7. The fusion protein according to claim 3 or 4, wherein The said functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain.

8. The fusion protein according to claim 7, wherein The said adenosine deaminase catalytic domain or cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

9. The fusion protein according to claim 7, characterized in that, The said adenosine deaminase catalytic domain includes a mutant of the amino acid sequence shown in SEQ ID NO: 29: Q148G+Q149M+P150R, named deaminase 005V1-10-3.

10. The fusion protein according to claim 7, wherein The said functional domain is the full length or a functional fragment of TadA8e.

11. The fusion protein according to claim 7, wherein The amino acid sequence of the said fusion protein is as shown in SEQ ID NO:

43.

12. An isolated polynucleotide, characterized in that, The said polynucleotide encodes the protein according to claim 1, the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11.

13. The polynucleotide according to claim 12, wherein, The said polynucleotide is a polynucleotide codon-optimized according to the codon preference of the host cell.

14. The polynucleotide according to claim 13, wherein the said host cell includes a prokaryotic cell or a eukaryotic cell.

15. The polynucleotide according to claim 14, wherein the said host cell is a eukaryotic cell.

16. The polynucleotide according to claim 14 or 15, wherein the said eukaryotic cell is selected from yeast cells, plant cells, or mammalian cells.

17. The polynucleotide according to claim 14, wherein the said host cell is a prokaryotic cell.

18. The polynucleotide according to claim 12, wherein the polynucleotide sequence is as shown in SEQ ID NO: 2 or 34.

19. A guide RNA (gRNA), characterized in that, The guide RNA includes a direct repeat (DR) sequence capable of binding to the protein according to claim 1 or the protein variant according to claim 2, and a spacer sequence capable of targeting a target sequence, wherein the direct repeat (DR) sequence is as shown in SEQ ID NO:

5.

20. A composite, characterized in that, Comprising: (i) a protein component selected from the group consisting of the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, or a combination thereof; and (ii) a nucleic acid component selected from the group consisting of the guide RNA according to claim 19, a nucleic acid encoding the guide RNA according to claim 19, a precursor RNA of the guide RNA according to claim 19, a nucleic acid encoding the precursor RNA of the guide RNA according to claim 19, or a combination thereof; wherein the protein component and the nucleic acid component bind to each other to form a complex.

21. A carrier, characterized in that, Comprising the polynucleotide according to any one of claims 12-18 or a nucleic acid molecule encoding the guide RNA according to claim 19.

22. The vector according to claim 21, comprising: (1) a first regulatory element operably linked to a nucleotide sequence encoding the protein according to claim 1, or a nucleotide sequence encoding the protein variant according to claim 2, or a nucleotide sequence encoding the fusion protein according to any one of claims 3-11; and (2) a second regulatory element operably linked to a nucleotide sequence encoding a guide RNA, the guide RNA comprising: (a) a spacer sequence capable of hybridizing to a target sequence, and (b) a direct repeat (DR) sequence linked to the spacer sequence, capable of guiding the protein according to claim 1, the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11 to bind to the guide RNA to form a complex targeting the target sequence, wherein the direct repeat sequence is as shown in SEQ ID NO:

5.

23. The vector according to claim 22, wherein the first regulatory element and the second regulatory element are located on the same or different vectors.

24. The vector according to claim 22, wherein the first regulatory element and / or the second regulatory element is a promoter.

25. The vector according to claim 24, wherein the vector comprises one or more promoters operably linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, origin of replication, selectable marker, nucleic acid restriction site, and / or homologous recombination site.

26. The vector according to any one of claims 21-25, wherein the vector includes a plasmid, a viral vector.

27. The vector according to claim 26, wherein the viral vector is selected from the group consisting of: adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or a combination thereof.

28. The vector according to any one of claims 21-25, wherein the vector comprises a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, a multifunctional vector.

29. A host cell, characterized in that, Comprising the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the polynucleotide according to any one of claims 12-18, or the nucleic acid molecule encoding the guide RNA according to claim 19, the complex according to claim 20, or the vector according to any one of claims 21-28, wherein the host cell does not include a plant cell.

30. A CRISPR-Cas composition, characterized in that, Comprising: (i) a first component selected from the group consisting of: the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the nucleotide sequence encoding the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, and any combination thereof; and (ii) a second component, which is a nucleotide sequence comprising one or more guide RNAs according to claim 19, or a nucleotide sequence encoding the nucleotide sequence comprising one or more guide RNAs according to claim 19; The guide RNA is capable of forming a complex with the protein or protein variant or fusion protein described in (i).

31. A CRISPR-Cas system, characterized in that, Comprising one or more vectors, the one or more vectors comprising: (i) a first nucleic acid, which is a nucleotide sequence encoding the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11; the first nucleic acid is operably linked to a first regulatory element; and (ii) a second nucleic acid, which encodes a nucleotide sequence comprising the guide RNA according to claim 19; the second nucleic acid is operably linked to a second regulatory element; Wherein: The first nucleic acid and the second nucleic acid are present on the same or different vectors; The guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

32. A kit, characterized in that, Comprising one or more components selected from the following: The protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the polynucleotide according to any one of claims 12-18, the complex according to claim 20, the vector according to any one of claims 21-28, the CRISPR-Cas composition according to claim 30, or the system according to claim 31.

33. A delivery composition, characterized in that, Comprising a delivery vector, and one or more selected from the following: The protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the polynucleotide according to any one of claims 12-18, the complex according to claim 20, the vector according to any one of claims 21-28, the CRISPR-Cas composition according to claim 30, or the system according to claim 31.

34. An enzyme preparation, characterized in that, The enzyme preparation comprises the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the complex according to claim 20, the CRISPR-Cas composition according to claim 30, the system according to claim 31, or the delivery composition according to claim 33.

35. A medicine box, characterized in that, Comprising: A first container, and the complex according to claim 20, the CRISPR-Cas composition according to claim 30, or the system according to claim 31 located in the first container, or a drug containing the complex according to claim 20, the CRISPR-Cas composition according to claim 30, or the system according to claim 31.

36. A medicine box, characterized in that, Comprising: (a1) A first container, and the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, or its coding gene or its expression vector located in the first container, or a drug containing the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, or its coding gene or its expression vector. (b1) Optionally, a second container, and the guide RNA according to claim 19 or its expression vector located in the second container, or a drug containing the guide RNA according to claim 19 or its expression vector.

37. A non-therapeutic and non-diagnostic method for targeting and editing a target gene or cleaving a target gene, characterized in that, Comprising: Contacting the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, or the complex according to claim 20, the CRISPR-Cas composition according to claim 30, the system according to claim 31, the delivery composition according to claim 33, the enzyme preparation according to claim 34, or the kit according to claim 35 or 36 with the target gene, or delivering it into a cell containing the target gene, wherein the target sequence is present in the target gene.

38. A non-therapeutic and non-diagnostic method for inducing a change in cell state, characterized in that, The method comprises contacting the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, or the complex according to claim 20, the CRISPR-Cas composition according to claim 30, the system according to claim 31, the delivery composition according to claim 33, the enzyme preparation according to claim 34, or the kit according to claim 35 or 36 with the target gene in the cell.

39. A non-therapeutic and non-diagnostic method for altering the expression of a gene product, characterized in that, Comprising: Contact the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, or the complex according to claim 20, or the CRISPR-Cas composition according to claim 30, or the system according to claim 31, or the delivery composition according to claim 23, or the enzyme preparation according to claim 34, or the kit according to claim 35 or 36 with a nucleic acid molecule encoding the gene product, or deliver it into a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

40. An in vitro, ex vivo cell or cell line or their progeny, characterized in that, The cell or cell line or their progeny comprise: the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the polynucleotide according to any one of claims 12-18, the complex according to claim 20, the vector according to any one of claims 21-28, the CRISPR-Cas composition according to claim 30, or the system according to claim 31, or the delivery composition according to claim 33, and the cell or cell line or their progeny are not plant cells, animal embryonic stem cells or animal germ cells.

41. Use of the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the polynucleotide according to any one of claims 12-18, the complex according to claim 20, the vector according to any one of claims 21-28, the CRISPR-Cas composition according to claim 30 or the system according to claim 31 or the kit according to claim 32 or the delivery composition according to claim 33 or the enzyme preparation according to claim 34 or the cartridge according to claim 35 or 36, characterized in that, For preparing a drug or preparation for nucleic acid editing.

42. The use according to claim 41, characterized in that, The drug or preparation is for gene or genome editing.

43. Use of the protein according to claim 1, the protein variant according to claim 2, the fusion protein according to any one of claims 3-11, the polynucleotide according to any one of claims 12-18, the complex according to claim 20, the vector according to any one of claims 21-28, the CRISPR-Cas composition according to claim 30 or the system according to claim 31 or the kit according to claim 32 or the delivery composition according to claim 33 or the enzyme preparation according to claim 34 or the cartridge according to claim 35 or 36, characterized in that, For preparing a drug or preparation for one or more selected from the group consisting of: (i) ex vivo gene or genome editing; (ii) detection of single-stranded DNA ex vivo; (iii) editing a target sequence in a target locus to modify a biological or non-human organism; (iv) treating a disorder caused by a defect in a target sequence in a target locus; (v) treating a disorder or disease of a subject in need thereof.

44. The use according to claim 43, characterized in that, The disorder or disease includes cancer, infectious diseases, neurological diseases, ophthalmic diseases, and hearing diseases.

45. The use according to claim 43, characterized in that, The disease or disorder includes cystic fibrosis, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy, α-1-antitrypsin deficiency, Pompe disease (glycogen storage disease type II), myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, hereditary chronic kidney disease, sickle cell disease, β-thalassemia, frontotemporal dementia, Leber congenital amaurosis, hyperlipidemia, transthyretin amyloidosis, retinal diseases, Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, and bladder cancer.

46. The use according to claim 45, wherein the disease or disorder includes hypercholesterolemia, macular degeneration, melanoma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin's lymphoma, and non-Hodgkin's lymphoma.

47. The use according to claim 43, wherein, The disorder or disease is caused by a pathogenic point mutation.

48. A method for detecting the presence of a target nucleic acid molecule in a sample, characterized in that, The method includes contacting a sample with the protein according to claim 1, or the protein variant according to claim 2, or the fusion protein according to any one of claims 3-11, or the complex according to claim 20, the CRISPR-Cas composition according to claim 30 or the system according to claim 31, the kit according to claim 32 or the delivery composition according to claim 33 or the enzyme preparation according to claim 34 and a non-target sequence, and detecting a detectable signal generated by cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA.

Citation Information

Patent Citations

  • Polypeptides comprising multimers of nuclear localization signals or of protein transduction domains and their use for transferring molecules into cells

    WO2001038547A2

  • Adenosine deaminase, base editor fusion protein, base editor system and application

    CN114634923A

  • CRISPR-Cas effector protein as well as gene editing system and application thereof

    CN114921439A