Protein, polynucleotide, vector, vector system, composition¸ kit, cell, target DNA modification method, and production method

Novel Cas-like proteins with modified amino acid sequences address the limitations of CRISPR/Cas9 systems by recognizing and modifying NAAN, NAGN, or NGAN PAM sequences, enhancing genome editing capabilities.

WO2026100694A1PCT designated stage Publication Date: 2026-05-15SETSUROTECH INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SETSUROTECH INC
Filing Date
2025-11-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing CRISPR/Cas9 systems struggle to efficiently modify target DNA sequences with specific PAM sequences such as NAAN, NAGN, or NGAN, limiting their applicability in genome editing.

Method used

Development of novel Cas-like proteins that can recognize and modify DNA sequences with PAM sequences like NAAN, NAGN, or NGAN, including variants with amino acid modifications that retain nuclease activity.

Benefits of technology

Enables efficient modification of target DNA sites previously difficult for conventional Cas proteins, expanding the range of editable sequences in genome editing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025039156_15052026_PF_FP_ABST
    Figure JP2025039156_15052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a novel protein that can be used in modifying target DNA. A protein according to the present disclosure is a protein (a1), a protein (a2), or a protein (a3). (a1) is a protein that comprises the amino acid sequence of any one of SEQ ID NO: 1-19. (a2) is a protein that comprises an amino acid sequence obtained by deleting, inserting, substituting, or adding one or more amino acids in the amino acid sequence of any one of SEQ ID NO: 1-19, and that has nuclease activity. (a3) is a protein that comprises an amino acid sequence having not less than 80% identity with the amino acid sequence of any one of SEQ ID NO: 1-19, and that has nuclease activity.
Need to check novelty before this filing date? Find Prior Art

Description

Protein, polynucleotide, vector, vector system, composition, kit, cell, method for modifying target DNA, and production method

[0001] The present disclosure relates to a protein, a polynucleotide, a vector, a vector system, a composition, a kit, a cell, a method for modifying target DNA, and a production method.

[0002] In recent years, genome editing using the CRISPR / Cas9 (Clustered Regularly Interspaced Palindromic Repeats / CRISPR-associated protein 9) system and the like has been widely used. In addition to Cas9, searches for other Cas proteins such as Cas12 and the creation of variants of Cas proteins such as MAD7 have been carried out (Patent Document 1).

[0003] International Publication No. 2020 / 086475

[0004] Therefore, an object of the present disclosure is to provide a new protein that can be used for modifying target DNA and the like.

[0005] In order to achieve the above object, the protein of the present disclosure includes the protein of the following (a1), (a2), or (a3): (a1) a protein consisting of any one of the amino acid sequences of SEQ ID NOs: 1 to 19; (a2) a protein consisting of an amino acid sequence in which one or several amino acids are deleted, inserted, substituted, or added in any one of the amino acid sequences of SEQ ID NOs: 1 to 19 and having nuclease activity; (a3) a protein consisting of an amino acid sequence having 80% or more identity to any one of the amino acid sequences of SEQ ID NOs: 1 to 19 and having nuclease activity.

[0006] The polynucleotide of the present disclosure encodes the protein of the present disclosure.

[0007] The vector of the present disclosure includes the polynucleotide of the present disclosure.

[0008] The vector system of the present disclosure includes a first vector including the polynucleotide of the present disclosure and a second vector including a polynucleotide encoding a guide RNA for a target site.

[0009] The compositions of the present disclosure include the proteins of the present disclosure, the polynucleotides of the present disclosure, and / or the vectors of the present disclosure, and a guide RNA or polynucleotide encoding a guide RNA to a target site.

[0010] The kits of this disclosure include the proteins of this disclosure, the polynucleotides of this disclosure, and / or the vectors of this disclosure, and a guide RNA or polynucleotide encoding it to a target site.

[0011] The cells of this disclosure include the proteins, polynucleotides, vectors, and / or vector systems of this disclosure.

[0012] The method for modifying target DNA according to the present disclosure includes a modification step of modifying the target DNA by bringing the target DNA into contact with the protein of the present disclosure and a guide RNA for a target site within the target DNA, within a cell.

[0013] A manufacturing method of the present disclosure is a method for manufacturing cells or non-human animals in which target DNA has been modified, comprising an introduction step of introducing the protein of the present disclosure, the polynucleotide of the present disclosure, and / or the vector of the present disclosure and a guide RNA to the target site into cells or non-human animals.

[0014] This disclosure provides novel proteins that can be used to modify target DNA.

[0015] Figure 1 is a graph showing the results of the EGxxFP assay in Example 1. Figure 2 is a graph showing the results of the EGxxFP assay in Example 1. Figure 3 is a graph showing the results of the EGxxFP assay in Example 1. Figure 4 is a graph showing the results of the EGxxFP assay in Example 1. Figure 5 is a schematic diagram showing the sites of the cleavage active sites of ST9 and CR-FD2 in Example 2. Figure 6 is a graph showing the results of the EGxxFP assay in Example 2. Figure 7 is a schematic diagram showing the sites of the cleavage active sites of ST9 and CR-FD3 in Example 3. Figure 8 is a graph showing the results of the EGxxFP assay in Example 3. Figure 9 is a graph showing the results of the EGxxFP assay in Example 4.

[0016] <Definitions> In this specification, “protein” or “peptide” means a polymer composed of unmodified amino acids (natural amino acids), modified amino acids, and / or synthetic amino acids.

[0017] In this specification, "nuclease" means an enzyme capable of cleaving phosphodiester bonds between nucleotides in nucleic acids. Examples of nucleases include exonucleases, which digest nucleic acids from the ends, and endonucleases, which cleave nucleotide chains in the middle. These nucleases can be classified into DNA nucleases, which act on DNA, and RNA nucleases, which act on RNA.

[0018] In this specification, "protospacer-adjacent motif (PAM) sequence" means a sequence adjacent to a sequence targeted by a Cas nuclease in the CRISPR bacterial adaptive immune system. The length and nucleotide sequence of the PAM sequence vary depending on the type of Cas nuclease.

[0019] In this specification, “coding sequence” means a nucleic acid (RNA or DNA molecule) containing a nucleotide sequence that codes for a protein. The coding sequence can be codon-optimized.

[0020] In this specification, “polynucleotide” means a polymer of deoxyribonucleotides (DNA), ribonucleotides (RNA), and / or modified nucleotides. In this specification, when “polynucleotide” is used in combination with a specific protein, “polynucleotide” means a polymer of nucleotides that encode the amino acid sequence of the protein. Examples of such polynucleotides include genomic DNA, cDNA, mRNA, etc. The polynucleotide may be single-stranded or double-stranded, for example. The polynucleotide is interchangeable with “nucleic acid” or “oligonucleotide.”

[0021] In this specification, “target site” means a region or sequence of polynucleotides targeted by the protein-guide RNA complex of this disclosure.

[0022] In this specification, "nucleic acid modification activity" means activity that can modify bases (e.g., A, T, C, G, or U) within nucleic acid sequences such as DNA and RNA. "Nucleic acid modification" means substitution, conversion, deletion, insertion, addition, etc., of bases within a nucleic acid sequence.

[0023] In this specification, “guide RNA” means RNA that can target the proteins of this disclosure to a target DNA sequence. The guide RNA includes a crRNA (CRISPR RNA) that binds to the target site and optionally a tracrRNA (trans activating crRNA) that is involved in the activity of the CRISPR-Cas system, depending on the type of Cas protein. The guide RNA may be, for example, a single-stranded RNA in which the crRNA and tracrRNA are directly or indirectly linked, or it may be two RNAs, the crRNA and the tracrRNA.

[0024] In this specification, "crRNA" means RNA having a base sequence (hereinafter also referred to as "sequence") complementary to the target DNA sequence. The crRNA is an RNA that constitutes the CRISPR / Cas system and is known to perform the function of recognizing the target DNA sequence. The crRNA generally includes a spacer sequence and a repeat sequence. The spacer sequence can also be said to be a sequence (guide sequence) that can form a double helix with the complementary strand of the target DNA sequence. The crRNA can form a double helix with, for example, either the sense strand or the antisense strand of the target DNA.

[0025] In this specification, "tracrRNA" means RNA having a sequence complementary to a portion of the crRNA. The sequence complementary to a portion of the crRNA is also called an anti-repeat region (sequence). The tracrRNA has a sequence that includes an anti-repeat region (sequence) (repeating region) followed by one or more hairpin structures, and is capable of forming a stem-loop. In some Cas proteins, the tracrRNA is known to function as a scaffold for binding the Cas protein to the crRNA. The proteins of this disclosure can, for example, directly bind to crRNA.

[0026] In this specification, “hybridize” means annealing with a complementary polynucleotide resulting from nucleotide complementarity, specifically the complementarity of bases in a nucleotide. That is, it means that two polynucleotides can form a non-covalent pair via hydrogen bonds. The hybridize may be, for example, the joining of two complementary base sequences, or substantially complementary base sequences that have one or more mismatched base pairs.

[0027] In this specification, “complementarity” means that one polynucleotide and another polynucleotide can form a nucleotide pair, i.e., a base pair.

[0028] In this specification, “vector” means a nucleic acid delivered to a host cell in vitro or in vivo, or a recombinant plasmid or virus containing such nucleic acid.

[0029] In this specification, "target DNA sequence" means the base sequence of the target polynucleotide in the target DNA that hybridizes with the guide RNA.

[0030] In this specification, "eukaryote" means an organism that has a nucleus enclosed by a nuclear membrane.

[0031] In this specification, "prokaryotes" refers to organisms that do not have a nucleus enclosed by a nuclear membrane.

[0032] Sequence information for proteins or nucleic acids (e.g., DNA or RNA) encoding them, as described herein, can be obtained from sources such as the Protein Data Bank, UniProt, Ensembl, or GenBank. RNA nucleic acid sequences can also be obtained from corresponding DNA nucleic acid sequences using appropriate sequence conversion software.

[0033] The following is an explanation of the present disclosure with examples, but the present disclosure is not limited to these examples and can be modified as desired. Furthermore, unless otherwise specified, each explanation in the present disclosure is interchangeable with one another. In this specification, when the expression "~" is used, it is used to mean including the numerical or physical value before and after it. Also, in this specification, the expression "A and / or B" includes "A only," "B only," and "both A and B."

[0034] <Proteins> In one embodiment, the present disclosure provides novel proteins that can be used for modifying target DNA. The proteins of the present disclosure are proteins of the following nature (a). The proteins of the following nature (a) may have nuclease activity or nickas activity. (a) Proteins of the following nature (a1), (a2), or (a3): (a1) Proteins consisting of any amino acid sequence of SEQ ID NOs: 1 to 19; (a2) Proteins consisting of amino acid sequences in which one or more amino acids are deleted, inserted, substituted, or added in any amino acid sequence of SEQ ID NOs: 1 to 19; (a3) ​​Proteins consisting of amino acid sequences having 80% or more identity with any amino acid sequence of SEQ ID NOs: 1 to 19.

[0035] The Cas protein recognizes the PAM sequence of target DNA and modifies the target DNA through the nuclease activity of the Cas protein, etc. The PAM sequences that the Cas protein can recognize vary depending on the type and origin of the Cas protein. However, it is known that there is a bias in the PAM sequences that the Cas protein can recognize. For example, NAAN, NAGN, or NGAN (N = A, T, G, or C; the same applies hereinafter) are known to be difficult to recognize by the Cas protein, especially the commonly used Cas9 and Cas12. As a result of diligent searching for various Cas proteins, the inventors have found a protein that can recognize NAAN, NAGN, or NGAN as the PAM sequence, and have completed this disclosure. According to this disclosure, a novel protein that can recognize NAAN, NAGN, or NGAN as the PAM sequence can be provided. Therefore, according to this disclosure, target DNA can be modified even at target sites with base sequences that were difficult to modify with conventional Cas proteins, such as NAAN, NAGN, or NGAN.

[0036] The protein (a1) described above is a protein designed by the present inventors as a Cas-like protein.

[0037]

[0038] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 1, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 1, for example, 1 to 287, 1 to 284, 1 to 215, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2. In this disclosure, the numerical range of the number may be any range in which all positive integers belong.

[0039] In the protein of (a2) above, the substitution is preferably a conservative substitution. A conservative substitution means the substitution of an amino acid residue with an amino acid residue having a similar side chain. Examples of conservative substitutions include substitutions between amino acid residues having basic side chains such as lysine, arginine, and histidine; substitutions between amino acid residues having acidic side chains such as aspartic acid and glutamic acid; substitutions between amino acid residues having non-charged polar side chains such as glycine, asparagine, glutamine, serine, threonine, tyrosine, and cysteine; substitutions between amino acid residues having non-polar side chains such as alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, and tryptophan; substitutions between amino acid residues having β-branched side chains such as threonine, valine, and isoleucine; substitutions between amino acid residues having aromatic side chains such as tyrosine, phenylalanine, tryptophan, and histidine; and so on (the same applies hereinafter).

[0040] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 1, then in the proteins (a1) to (a3), for example, the 1136th amino acid E (glutamic acid) or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G (glycine). In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1209th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A, as indicated by the underline, is substituted with E (glutamic acid), it is even more preferable that the 1207th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1208th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1209th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), and the 1210th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with R (arginine).

[0041]

[0042] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 2, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 2, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0043] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 2, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0044]

[0045] When the protein of the present disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 3, in the protein of (a2), "one or several" may be, for example, within the range in which (a2) functions as a protein having nuclease activity. "One or several" of (a2) is, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2 in the amino acid sequence of SEQ ID NO: 3.

[0046] When the protein of the present disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 3, in the proteins of (a1) to (a3), for example, the 1129th amino acid E shown underlined or the amino acid residue corresponding to the amino acid E may be substituted with G. In this case, since the proteins of (a1) to (a3) can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) shown underlined or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), and it is more preferable that the 1203rd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with R (arginine).

[0047]

[0048] When the protein of the present disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 4, in the protein of (a2), "one or several" may be, for example, within the range in which (a2) functions as a protein having nuclease activity. "One or several" of (a2) in the amino acid sequence of SEQ ID NO: 4 is, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0049] When the protein of the present disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 4, in the proteins of (a1) to (a3), for example, the 1129th amino acid E shown underlined or the amino acid residue corresponding to the amino acid E may be substituted with G. In this case, since the proteins of (a1) to (a3) can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) shown underlined or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), the 1200th amino acid V (valine) or the amino acid residue corresponding to the V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to the P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), and it is more preferable that the 1203rd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with R (arginine).

[0050]

[0051] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 5, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 5, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0052] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 5, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0053]

[0054] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 6, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 6, for example, 1 to 287, 1 to 215, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0055] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 6, then in the proteins (a1) to (a3), for example, the 1136th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1209th amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1207th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1208th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1209th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1210th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0056]

[0057] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 7, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 7, for example, 1 to 285, 1 to 214, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0058] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 7, then in the proteins (a1) to (a3), for example, the 1127th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1198th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1199th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0059]

[0060] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 8, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 8, for example, 1 to 284, 1 to 213, 1 to 142, 1 to 71, 1 to 56, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0061] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 8, then in the proteins (a1) to (a3), for example, the 1120th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1191st amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1192nd amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1194th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0062]

[0063] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 9, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be, for example, 1 to 285, 1 to 214, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0064] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 9, then in the proteins (a1) to (a3), for example, the 1127th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1198th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1199th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0065]

[0066] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 10, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 10, for example, "one or several" may be 1 to 284, 1 to 213, 1 to 142, 1 to 71, 1 to 56, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0067] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 10, then in the proteins (a1) to (a3), for example, the 1120th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1191st amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1192nd amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1194th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0068]

[0069] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 11, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 11, for example, "one or several" may be 1 to 285, 1 to 214, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0070] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 11, then in the proteins (a1) to (a3), for example, the 1128th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1199th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1200th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0071]

[0072] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 12, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 12, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0073]

[0074]

[0075] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 14, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2) or (a3), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 14, such as 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0076] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 14, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0077]

[0078] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 15, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 15, "one or several" may be, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0079] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 15, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0080]

[0081] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 16, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 16, "one or several" may be, for example, 1 to 254, 1 to 190, 1 to 142, 1 to 127, 1 to 63, 1 to 50, 1 to 38, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0082] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 16, then in the proteins (a1) to (a3), for example, the 969th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1040th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1041st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1043rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0083]

[0084] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 17, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 17, for example, 1 to 252, 1 to 189, 1 to 142, 1 to 126, 1 to 63, 1 to 50, 1 to 37, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0085] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 17, then in the proteins (a1) to (a3), for example, the 960th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1033rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1031st amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1032nd amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1033rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1034th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0086]

[0087] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 18, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 18, "one or several" may be, for example, 1 to 254, 1 to 190, 1 to 142, 1 to 127, 1 to 63, 1 to 50, 1 to 38, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0088] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 18, then in the proteins (a1) to (a3), for example, the 969th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1040th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1041st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1043rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0089]

[0090] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 19, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 19, "one or several" may be, for example, 1 to 254, 1 to 190, 1 to 142, 1 to 127, 1 to 63, 1 to 50, 1 to 38, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0091] In the protein (a3) ​​described above, "identity" is defined as, for example, within the range in which (a3) ​​functions as a protein having nuclease activity. The "identity" of (a3) ​​is, for example, 70%, 75%, 80% or more, 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, and 99% or more with respect to the amino acid sequence of (a1). The "identity" can be calculated, for example, by using the default parameters in the homology algorithm BLAST (http: / / www.ncbi.nlm.nih.gov / BLAST / ) of the National Center for Biotechnology Information (NCBI) (hereinafter the same applies).

[0092] The proteins of this disclosure may recognize, for example, protospacer-adjacent motif (PAM) sequences. Examples of PAM sequences include TTTV (SEQ ID NO: 20, V=A, G, or C; the same applies hereinafter), NAAN (SEQ ID NO: 21), NAGN (SEQ ID NO: 22), or NGAN (SEQ ID NO: 23). NAAN (SEQ ID NO: 21), NAGN (SEQ ID NO: 22), or NGAN (SEQ ID NO: 23) are PAM sequences that are generally known to be difficult to recognize with Cas9 or Cas12. Examples of NGAN (SEQ ID NO: 23) include AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), etc. Examples of NAGN (SEQ ID NO: 22) include AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), CAGC (SEQ ID NO: 30), and NAAN (SEQ ID NO: 21) include GAAC (SEQ ID NO: 31), CAAC (SEQ ID NO: 32), and so on.

[0093] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 1, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26), GAGC (SEQ ID NO: 29), and GAAC (SEQ ID NO: 31).

[0094] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 2, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25) and GGAC (SEQ ID NO: 26).

[0095] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 3, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), GAGC (SEQ ID NO: 29), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0096] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 4, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), CAGC (SEQ ID NO: 30), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0097] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 5, then the PAM sequence of the protein in (a) preferably includes, for example, GAGC (SEQ ID NO: 29).

[0098] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 6, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26) and GAGC (SEQ ID NO: 29).

[0099] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 7, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0100] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 8, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0101] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 9, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), AAGC (SEQ ID NO: 28), and CAAC (SEQ ID NO: 32).

[0102] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 10, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26).

[0103] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 11, then the PAM sequence of the protein in (a) preferably includes, for example, GAGC (SEQ ID NO: 29).

[0104] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 12, the PAM sequence of the protein in (a) preferably includes, for example, AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), CAGC (SEQ ID NO: 30), and GAAC (SEQ ID NO: 31).

[0105] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 13, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26) and AAGC (SEQ ID NO: 28).

[0106] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 14, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26).

[0107] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 15, the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), and GAAC (SEQ ID NO: 31).

[0108] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 16, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), CAGC (SEQ ID NO: 30), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0109] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 17, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), and CAAC (SEQ ID NO: 32).

[0110] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 18, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), and GAAC (SEQ ID NO: 31).

[0111] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 19, then the PAM sequence of the protein in (a) preferably includes, for example, AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), CAGC (SEQ ID NO: 30), and GAAC (SEQ ID NO: 31).

[0112] The protein of this disclosure may be, for example, an RNA-dependent DNA nuclease. The protein of this disclosure may form a complex with, for example, a guide RNA as described below and function as a DNA nuclease. The nuclease may be, for example, an endonuclease.

[0113] The proteins of this disclosure may further have, for example, an enzyme activity domain. If the proteins of this disclosure have an enzyme activity domain, they do not have, for example, nuclease activity. In this case, the proteins of this disclosure may or may not have, for example, nickase activity. Examples of such proteins include the proteins (1) to (7) below. The enzyme activity includes, for example, nucleic acid modification activities such as methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity; transcription activation activity, transcription repression activity, transcription release factor activity, DNA integration activity, nucleic acid binding activity, etc. (1) In the amino acid sequence of SEQ ID NO: 1 or 6, it is preferable that the 1136th amino acid E is replaced with something other than G, for example, A, and further, the 1209th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1207th amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1208th amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1209th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1210th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein;(2) In the amino acid sequence of SEQ ID NOs. 2-5, 14, or 15, it is preferable that the 1129th amino acid E is replaced with something other than G, for example, A, and further, the 1202nd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is replaced with E (glutamic acid), the 1200th amino acid V (valine) or the amino acid residue corresponding to the V is replaced with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to the P is replaced with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is replaced with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (3) In the amino acid sequence of SEQ ID NO: 7 or 9, it is preferable that the amino acid sequence E at position 1127 is replaced with something other than G, for example, A, and further, the amino acid A (alanine) at position 1200 or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the amino acid V (valine) at position 1198 or the amino acid residue corresponding to said V is replaced with A (alanine), the amino acid P (proline) at position 1199 or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the amino acid A (alanine) at position 1200 or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the amino acid A (alanine) at position 1201 or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein;(4) In the amino acid sequence of SEQ ID NO: 8 or 10, it is preferable that the amino acid sequence E at position 1120 is replaced with something other than G, for example, A, and further, the amino acid A (alanine) at position 1193 or the amino acid residue corresponding to amino acid A is replaced with E (glutamic acid), the amino acid V (valine) at position 1191 or the amino acid residue corresponding to V is replaced with A (alanine), the amino acid P (proline) at position 1192 or the amino acid residue corresponding to P is replaced with D (aspartic acid), the amino acid A (alanine) at position 1193 or the amino acid residue corresponding to amino acid A is replaced with E (glutamic acid), and the amino acid A (alanine) at position 1194 or the amino acid residue corresponding to amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (5) In the amino acid sequence of Sequence ID No. 11, it is preferable that the 1128th amino acid sequence E is replaced with something other than G, for example, A, and further, the 1201st amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is replaced with E (glutamic acid), the 1199th amino acid V (valine) or the amino acid residue corresponding to the V is replaced with A (alanine), the 1200th amino acid P (proline) or the amino acid residue corresponding to the P is replaced with D (aspartic acid), the 1201st amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is replaced with E (glutamic acid), and the 1202nd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein;(6) In the amino acid sequence of SEQ ID NO: 16 or 18, it is preferable that the 969th amino acid E is replaced with something other than G, for example, A, and further, the 1042nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1040th amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1041st amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1042nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1043rd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (7) Preferably, in the amino acid sequence of Sequence ID No. 17, the 960th amino acid E is substituted with something other than G, for example, A, and further, the 1033rd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), the 1031st amino acid V (valine) or the amino acid residue corresponding to the V is substituted with A (alanine), the 1032nd amino acid P (proline) or the amino acid residue corresponding to the P is substituted with D (aspartic acid), the 1033rd amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), and the 1034th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein.

[0114] The enzyme activity domain is preferably a nucleic acid modification activity domain having nucleic acid modification activity. The nucleic acid modification activity is, for example, an activity that directly or indirectly causes DNA modification by modifying nucleic acids, such as nucleic acid degradation activity, nucleic acid base conversion activity, and DNA hydrolysis activity. The DNA modification is such as a DNA strand cleavage reaction that cleaves the DNA strand catalyzed by the nucleic acid degradation activity, a nucleic acid base conversion reaction that converts substituents on the purine or pyrimidine ring of a nucleic acid base to other groups without cleaving the DNA strand catalyzed by the nucleic acid base conversion activity, and a debase reaction that hydrolyzes the N-glycosidic bond of DNA catalyzed by the DNA glucosidase.

[0115] Nucleic acid base conversion activity includes, for example, deaminase activity. Examples of such deaminase activity include cytidine deaminase activity, adenosine deaminase activity, and guanosine deaminase activity.

[0116] Examples of the aforementioned DNA glycosylase activity include thymine DNA glycosylase activity, oxoguaning glycosylase activity, alkyladenine DNA glycosylase activity, and so on.

[0117] The enzyme-active domain is located, for example, at the 5' end and / or 3' end of the protein of this disclosure.

[0118] The enzyme-active domain is linked, for example, directly or indirectly to the protein of the Disclosure. The direct linkage is, for example, a covalent linkage, where, for example, the protein of the Disclosure and the enzyme-active domain constitute a fusion protein. The indirect linkage is, for example, a linkage via a binding tag-binding partner. The binding tag-binding partner is, for example, a combination of substances with mutually specific binding properties. Examples of the binding tag-binding partner include a combination of biotin and avidin or streptavidin, a combination of an SH3 domain (SH3) and an SH ligand, a combination of nickel and a His tag, an epitope tag such as a flag(trademark)-tag, HA-tag, T7-tag, V5-peptide-tag, and / or Myc-tag, and an antibody against the tag. In this case, the protein of the Disclosure and the enzyme-active domain are linked (bound) via the binding tag-binding partner, for example, after transcription and translation from a polynucleotide.

[0119] The proteins of this disclosure may, for example, have the amino acids of the N-terminal amino acid residues protected by a protecting group (e.g., a formyl group, a C1-6 acyl group such as a C1-6 alkanoyl such as an acetyl group, etc.).

[0120] The proteins of this disclosure may, for example, have a pyroglutamine-oxidized N-terminal glutamine residue that can be cleaved and produced in vivo.

[0121] The proteins of this disclosure may, for example, have substituents on the side chains of amino acids within the molecule (e.g., -OH, -SH, amino group, imidazole group, indole group, guanidino group, etc.) protected by appropriate protecting groups (e.g., formyl group, C1-6 acyl group such as a C1-6 alkanoyl group such as an acetyl group, etc.).

[0122] The protein of this disclosure may be, for example, a complex protein such as a glycoprotein to which sugar chains are attached.

[0123] The proteins of this disclosure may, for example, be in the form of salts with an acid or a base. The salts are not particularly limited and include acidic salts, basic salts, etc. Examples of acidic salts include inorganic acid salts such as hydrochloride, hydrobromide, sulfate, nitrate, and phosphate; organic acid salts such as acetate, propionate, tartrate, fumarate, maleate, malate, citrate, methanesulfonate, and p-toluenesulfonate; and amino acid salts such as aspartate and glutamate. Examples of basic salts include alkali metal salts such as sodium salt and potassium salt; and alkaline earth metal salts such as calcium salt and magnesium salt.

[0124] The proteins disclosed herein can be readily produced according to known genetic engineering techniques. For example, the proteins disclosed herein can be produced using PCR, restriction enzyme digestion, DNA ligation techniques, in vitro transcription and translation techniques, recombinant protein production techniques, etc.

[0125] <Polynucleotides> In another embodiment, the Disclosure provides polynucleotides. The polynucleotides of the Disclosure include coding sequences for proteins of the Disclosure. The polynucleotides of the Disclosure can be described by reference to the description of proteins of the Disclosure.

[0126] The polynucleotides of the disclosed herein can be designed by substituting corresponding codons based on the amino acid sequence of the protein of the disclosed herein. The base sequence of the polynucleotide of the disclosed herein may be, for example, codon-optimized.

[0127] <Vectors> In another embodiment, the Disclosure provides vectors. The vectors of the Disclosure comprise the polynucleotides of the Disclosure. The vectors of the Disclosure can be described by reference to the descriptions of proteins and polynucleotides of the Disclosure.

[0128] The vectors of this disclosure may, for example, include a polynucleotide encoding a guide RNA for a target site in target DNA. In this case, the polynucleotide encoding the guide RNA is hybridizable to the target DNA, i.e., hybridizable to a target sequence in the target DNA. The polynucleotide encoding the guide RNA may further include other components besides the guide RNA.

[0129] The proteins of this disclosure can perform modifications to a target site by forming a complex with a guide RNA having a structure similar to that of a Cas protein guide RNA, for example. The proteins of this disclosure can use, for example, an RNA molecule having a structure or base sequence similar to that of a Cas9 protein guide RNA or a Cas12 protein guide RNA as the guide RNA. The proteins of this disclosure can also use, for example, the guide RNA of a Cas9 protein or the guide RNA of a Cas12 protein as the guide RNA.

[0130] The guide RNA includes, for example, crRNA. Since the crRNA includes, for example, a spacer sequence and a direct repeat sequence, it can also be said that the guide RNA includes, for example, a spacer sequence and a direct repeat sequence. In the guide RNA, the spacer sequence hybridizes, for example, to the target nucleotide sequence of the target DNA. The spacer sequence may be, for example, completely complementary (identical) or substantially complementary (identical) to the target nucleotide sequence of the target DNA. "Substantially complementary" means, for example, that it consists of a base sequence having 50% or more identity with the nucleotide sequence of the target DNA and is capable of hybridizing to the nucleotide sequence of the target DNA. The aforementioned "identity" is, for example, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more with respect to the nucleotide sequence of the target DNA. The substantially complementary sequence is, for example, a sequence in which one or two mismatched base pairs exist between the spacer sequence and the target DNA. In the guide RNA, the length of the spacer sequence is not particularly limited as long as it is long enough to hybridize to the target DNA, and examples include 14 to 30 bases, 16 to 28 bases, 16 to 25 bases, 20 to 25 bases, 20 to 22 bases, or 21 to 23 bases.

[0131] The hybridizes can be detected, for example, by various hybridization assays under stringent conditions. The hybridization assays are not particularly limited, and methods such as those described in Sambrook et al., "Molecular Cloning: A Laboratory Manual 2nd Ed." [Cold Spring Harbor Laboratory Press (1989)] can also be employed.

[0132] The aforementioned "stringent conditions" may be, for example, low-stringent conditions, medium-stringent conditions, or high-stringent conditions. "Low-stringent conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 32°C. "Medium-stringent conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 42°C. "High-stringent conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 50°C. The degree of stringency can be set by a person skilled in the art by appropriately selecting conditions such as temperature, salt concentration, probe concentration and length, ionic strength, and time. The "stringent conditions" can also be those described in Sambrook et al.'s "Molecular Cloning: A Laboratory Manual 2nd Ed." [Cold Spring Harbor Laboratory Press (1989)], for example.

[0133] The spacer sequence preferably further includes, for example, a seed sequence. The seed sequence is, for example, the 5' terminal sequence of the PAM sequence in the spacer sequence and is important for the specificity of the guide RNA to the target DNA. In the guide RNA, the seed sequence is, for example, substantially or completely complementary to the nucleotide sequence of the target DNA, preferably the latter. If the seed sequence is substantially complementary to the nucleotide sequence of the target DNA, the substantially complementary sequence is, for example, a sequence in which one or two mismatched base pairs exist between the seed sequence and the nucleotide sequence of the target DNA.

[0134] In the guide RNA, the length of the direct repeat sequence is not particularly limited, and examples include 19 to 40 nucleotides in length, 21 to 38 nucleotides in length, etc.

[0135] The length of the guide RNA sequence is set so that the protein of this disclosure can target the target site of the target DNA. The length of the guide RNA sequence is not particularly limited and may be, for example, 30-220 nucleotides, 40-220 nucleotides, 50-220 nucleotides, 60-220 nucleotides, 70-220 nucleotides, 80-220 nucleotides, 90-220 nucleotides, 100-220 nucleotides, 110-220 nucleotides, 110-210 nucleotides, 110-200 nucleotides, 110-190 nucleotides, 110-180 nucleotides, 110-170 nucleotides, 110-160 nucleotides, or 110-150 nucleotides. If the guide RNA includes crRNA and tracrRNA as separate RNAs, the length of the guide RNA sequence is, for example, the sum of the lengths of both RNAs. If the guide RNA contains both crRNA and tracrRNA within the same RNA (as in the case of sgRNA described later), the length of the guide RNA sequence is, for example, the length of the sgRNA sequence. If the guide RNA contains crRNA, the length of the guide RNA sequence is, for example, the length of the crRNA sequence.

[0136] The guide RNA may be of one type or of multiple types. If there are multiple types of guide RNA, it is preferable that the guide RNA be configured to hybridize to different target DNA sequences, for example. If there are multiple types of guide RNA, for example, by setting the different target DNA sequences so that the guide RNA targets the 5' and 3' sides of the target site or target region, the protein-guide RNA complex of this disclosure can induce a deletion (large deletion) of the target region between the 5' and 3' sides. If there are multiple types of guide RNA, for example, by setting the guide RNA to target the 5' and 3' sides of the target region and introducing donor DNA having a sequence complementary to the 5' side region (5' arm region) and the 3' side region (3' arm region), the protein-guide RNA complex of this disclosure can recombinantly introduce the donor DNA into the region between the 5' and 3' sides.

[0137] The guide RNA may consist of crRNA alone, or it may consist of crRNA and tracrRNA. If the guide RNA consists of crRNA and tracrRNA, the guide RNA may be an sgRNA (single-stranded guide strand) formed by the fusion of the tracrRNA and the crRNA, or it may not be fused; that is, it may consist of two RNAs, crRNA and tracrRNA. If the tracrRNA and the crRNA are not fused, it is preferable that the tracrRNA and the crRNA have substantially or completely complementary sequences, such as the repeat sequence and the anti-repeat region. The structure of the guide RNA (single-stranded or double-stranded) and the presence or absence of tracrRNA can be determined depending on the protein of this disclosure.

[0138] The tracrRNA comprises a repeating sequence (anti-repeat region) and one or more hairpin structures. The number of hairpin structures is not particularly limited and may be, for example, 1 to 10, 1 to 5, 1 to 3, or 2. The tracrRNA is not particularly limited as long as it can bind to the crRNA and form a complex.

[0139] The tracrRNA sequence may, for example, have nucleotide regions with other functional properties.

[0140] The length of the tracrRNA sequence is not particularly limited and can be set according to the protein of this disclosure, for example. Examples of tracrRNA sequence lengths include 10-200 nucleotides, 20-190 nucleotides, 30-180 nucleotides, 40-180 nucleotides, 50-180 nucleotides, 60-180 nucleotides, 70-180 nucleotides, 70-170 nucleotides, 70-160 nucleotides, 70-150 nucleotides, 70-140 nucleotides, 70-130 nucleotides, 70-120 nucleotides, 70-110 nucleotides, 70-100 nucleotides, 70-90 nucleotides, and so on.

[0141] The length of the crRNA sequence is not particularly limited and can be set according to the protein of this disclosure. Examples of the tracrRNA sequence length include 25 to 80 nucleotides, 30 to 80 nucleotides, 35 to 80 nucleotides, 40 to 80 nucleotides, 49 to 80 nucleotides, 40 to 70 nucleotides, or 40 to 60 nucleotides.

[0142] The crRNA sequence can be configured to target a region adjacent to the PAM sequence of the target DNA.

[0143] An example of the nucleotide sequence of the guide RNA in this disclosure is the following sequence. In the nucleotide sequence of Sequence ID No. 33 below, the underlined N (where N is A, T, G, or U) at positions 36 to 56 is a nucleotide sequence (spacer sequence) that can hybridize with the target DNA sequence in the target DNA. In the nucleotide sequence of Sequence ID No. 33 below, N is 21 nucleotides long, but the length of N is not limited to this, and should be within the range that allows the protein in this disclosure to target the target site of the target DNA. The length of N may be, for example, 14 to 30 nucleotides, 16 to 28 nucleotides, 16 to 25 nucleotides, 20 to 25 nucleotides, 20 to 22 nucleotides, or 21 to 23 nucleotides. Guide RNA (Sequence ID No. 33) 5'-GUCAAAAGACCUUUGGAAUUUCUACUCUUGUAGAUNNNNNNNNNNNNNNNNNNNNN-3'

[0144] As mentioned above, the sequence of the guide RNA is not particularly limited as long as it can hybridize to the target DNA sequence.

[0145] The polynucleotides on the vector of this disclosure and the polynucleotides encoding guide RNA for a target site of target DNA are configured to be transcribed and / or translated in a host such as a cell to form a complex. The protein of this disclosure and the guide RNA are, for example, bound together to form a complex. The complex can, for example, target the protein of this disclosure to a target site of the target DNA in a sequence-dependent manner of the guide RNA, and modify the target DNA in an activity-dependent manner of the protein of this disclosure. The modifications include, for example, double-strand breaks, single-strand breaks, conversion, substitution, deletion, or insertion of one or more nucleotides in the target DNA. The modifications may also be, for example, modifications of the base, sugar, or phosphate portion of the target DNA. The complex of the guide RNA and the protein of this disclosure converts, substitutes, deletes, or inserts one or more nucleotides at the targeted site. The one or more nucleotides may be, for example, 1 to 10, 1 to 5, etc.

[0146] The target DNA may include, for example, chromosomal DNA, genomic DNA, mitochondrial DNA, viral DNA, binding elements, exogenous DNA, plasmid DNA, etc. If the target DNA encodes a gene, it may also be called, for example, a target gene. The target DNA may be single-stranded (ssDNA) or double-stranded (dsDNA).

[0147] The target DNA sequence into which the guide RNA can hybridize is not particularly limited and can be set, for example, according to the PAM (proto-spacer adjacent motif) sequence of the protein of this disclosure. The length of the target DNA sequence is not particularly limited and can be, for example, a length into which the guide RNA can specifically hybridize, such as 14 to 30 nucleotides or 16 to 25 nucleotides.

[0148] The vectors of this disclosure may, for example, include a promoter sequence on the vector. The promoter can be appropriately set, for example, depending on the type of cell expressing the polynucleotide encoding the protein of this disclosure and / or the guide RNA. Examples of the promoters include the T3 promoter, T7 promoter, sp6 promoter, EF1α promoter, SRα promoter, SV40 (Simian virus) promoter, LTR promoter, CMV (Cytomegalovirus) promoter, RSV (Respiratory syncytial virus) promoter, HSV-tk promoter, Cauliflower mosaic virus (CaMV) 35S promoter, actin promoter, heat shock promoter, and REF (Rubber Elongation Examples of promoters include Factor promoter, polyhedrin promoter, p10 promoter, trp promoter, lac promoter, recA promoter, λPL promoter, lpp promoter, tac promoter, GAL1 promoter, GAL10 promoter, PH05 promoter, PGK promoter, GAP promoter, ADH promoter, SpO1 promoter, SpO2 promoter, penP promoter, gyrA-ldh promoter, pgm promoter, rrn4 promoter, P23 promoter, araBAD promoter, cat promoter, cspA promoter, EM7 promoter, J23119 promoter, T5 promoter, tac promoter, U6 promoter, TDH3 promoter, TEF1 promoter, etc.

[0149] The vectors of this disclosure may, for example, include a polynucleotide encoding an enzyme-active protein or its enzyme-active domain. The enzyme activity of the enzyme-active protein may include, for example, nucleic acid modification activities such as methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity; transcription activation activity, transcription repression activity, transcription deactivation factor activity, DNA integration activity, nucleic acid binding activity, etc. The enzyme-active protein may also be, for example, a peptide fragment thereof, as long as it has catalytic activity.

[0150] The vectors of this disclosure may include, for example, a multicloning site, an enhancer, a splicing signal, a poly-A addition signal, a drug resistance gene, a nutritional complement gene, an origin of replication, etc.

[0151] The vectors of this disclosure can be prepared, for example, by inserting a polynucleotide containing the coding sequence of the protein of this disclosure into a skeletal vector (hereinafter also referred to as the "basic vector"). The type of expression vector is not particularly limited and can be appropriately determined, for example, depending on the type of host. Specifically, when synthesizing the vector by genetic engineering, the synthesis of the vector first involves, for example, designing and synthesizing a polynucleotide containing the coding sequence of the protein of this disclosure. This design and synthesis can be carried out by PCR, for example, using a vector containing a polynucleotide encoding a polynucleotide containing the coding sequence of the protein of this disclosure as a template, and using primers designed to synthesize a desired nucleic acid region. Then, a recombinant vector for protein expression (expression vector) can be obtained by ligating the obtained polynucleotide into a suitable vector, and a transformant can be obtained by introducing this recombinant vector into a host so that the target gene can be expressed (Sambrook J. et al., Molecular Cloning, A Laboratory Manual (4th edition) (Cold Spring Harbor Laboratory Press (2012)).

[0152] Examples of such vectors include phage vectors, plasmid vectors, viral vectors, retroviral vectors, chromosomal vectors, episomal vectors, and virus-derived vectors.

[0153] When transforming a host using the heat shock method as the introduction method, the vector may be, for example, a binary vector. Examples of expression vectors include pETDuet-1, pQE-80L, and pUCP26Km. When transforming bacteria such as E. coli, examples of expression vectors include pETDuet-1 vector (Novagen), pQE-80L (QIAGEN), pBR322, pB325, pAT153, and pUC8.

[0154] <Vector Systems> In another embodiment, the Disclosure provides vector systems. The vector systems of the Disclosure comprise a first vector comprising a polynucleotide of the Disclosure and a second vector comprising a polynucleotide encoding a guide RNA for a target site. The vector systems of the Disclosure can be described by reference to the descriptions of proteins, polynucleotides, and vectors of the Disclosure.

[0155] In the vector system of this disclosure, the first vector and the second vector may, for example, be arranged on the same vector or on different vectors.

[0156] <Compositions> In another embodiment, the Disclosure provides compositions. The compositions of the Disclosure comprise a protein, a polynucleotide, and / or a vector of the Disclosure, and a guide RNA or polynucleotide encoding a guide RNA to a target site on target DNA. The compositions of the Disclosure can be described by reference to the descriptions of the proteins, polynucleotides, vectors, and vector systems of the Disclosure.

[0157] <Kits> In another embodiment, the Disclosure provides kits capable of modifying target DNA. The kits of the Disclosure comprise the proteins of the Disclosure, the polynucleotides of the Disclosure, and / or the vectors of the Disclosure, and a guide RNA or polynucleotide encoding it to a target site. The kits of the Disclosure can be described by reference to the descriptions of the proteins, polynucleotides, vectors, vector systems, and compositions of the Disclosure.

[0158] The kits of this disclosure may include, for example, organisms. Examples of such organisms include prokaryotes and eukaryotic cells. Examples of such prokaryotes include lactic acid bacteria, Bacillus subtilis var. natto, Escherichia coli, and cyanobacteria. Examples of such eukaryotes include microorganisms, plants, animals, and insects. Examples of such animals include humans, monkeys, dogs, cats, rabbits, pigs, cows, mice, and rats. There may be one or more types of organisms.

[0159] The kits of this disclosure may include, for example, cells. Examples of such cells include prokaryotic cells and eukaryotic cells. Examples of such prokaryotic cells include lactic acid bacteria, Bacillus subtilis var. natto, Escherichia coli, and cyanobacteria. Examples of such eukaryotic cells include microbial cells, plant cells, and animal cells. Examples of such microbial cells include cells derived from microorganisms such as yeast, fungi, and protozoa. Examples of such plant cells include cells derived from plants such as seed plants, ferns, mosses, and algae. Examples of such animal cells include cells derived from vertebrates such as mammals, birds, reptiles, amphibians, and fish; arthropods such as insects; mollusks; and other animals. Examples of such cells include in vivo cells, in vitro cells, and primary cultured cells. The aforementioned cells may be of one type or multiple types.

[0160] The kit described herein may include, for example, instructions, a manual, etc.

[0161] The kit of this disclosure may include, for example, a culture medium. Examples of such culture media include YM medium, YPD medium, PD medium, DOB medium, SD medium, LB medium, NB medium, SCD medium, MRS medium, BHI medium, etc. The culture medium may also include additives such as: carbon sources such as glucose, dextrin, soluble starch, and sucrose; nitrogen sources such as ammonium salts, nitrates, corn slush liquor, peptone, casein, meat extract, soybean meal, and potato extract; inorganic substances such as calcium chloride, sodium dihydrogen phosphate, and magnesium chloride; vitamins; and growth factors. If the kit of this disclosure includes such additives, these additives may be added to the culture medium beforehand or added during culture. The additions may be continuous or intermittent.

[0162] The kits of the present disclosure may further include, for example, containers for storing the proteins of the present disclosure, the polynucleotides of the present disclosure, the vectors of the present disclosure, the vector systems of the present disclosure, and / or compositions of the present disclosure, and culture media.

[0163] The kits of the present disclosure may, for example, contain each component separately, or some or all of them may be contained in a mixed or unmixed state. In the case where all components of the kits of the present disclosure are contained in a single container in a mixed or unmixed state, the kits of the present disclosure may also be called, for example, culture media.

[0164] The kits of this disclosure can also be suitably used, for example, as test kits or research kits for use in genome editing of organisms and / or cells.

[0165] <Cells> In another embodiment, the present disclosure provides cells. The cells of the present disclosure include the proteins of the present disclosure, the polynucleotides of the present disclosure, the vectors of the present disclosure, and / or the vector systems of the present disclosure. The cells of the present disclosure can be made by reference to the descriptions of the proteins, polynucleotides, vectors, vector systems, and kits of the present disclosure.

[0166] If the cells of the present disclosure include the polynucleotides of the present disclosure, the vectors of the present disclosure, and / or the vector systems of the present disclosure, then the cells of the present disclosure can be said to express the polynucleotides of the present disclosure.

[0167] <Method for Modifying Target DNA> In another embodiment, the Disclosure provides a method for modifying target DNA. The method for modifying target DNA according to the Disclosure includes a modification step of modifying the target DNA by contacting the target DNA with the protein of the Disclosure and a guide RNA for a target site in the target DNA within a cell. The modification method according to the Disclosure can be described by reference to the descriptions of the protein, polynucleotide, vector, vector system, composition, and kit of the Disclosure.

[0168] In the modification step, the complex of the protein of the present disclosure and the guide RNA is targeted to the target DNA, and the target DNA is modified by the complex at the target site.

[0169] The modification method of this disclosure includes, for example, a transfection step of introducing the protein of this disclosure or a polynucleotide encoding it and the guide RNA or a polynucleotide encoding it into the cells. The transfection can be carried out by known methods capable of introducing (transfecting) exogenous proteins or exogenous DNA into host cells, specifically, methods such as transfection using a gene gun such as a particle gun, microinjection, calcium phosphate, polyethylene glycol, lipofection using liposomes, electroporation, ultrasonic nucleic acid transfection, DEAE-dextran, direct injection using microglass tubes, hydrodynamic, cationic liposome, lysozyme, competent, PEG, methods using transfection aids, methods via Agrobacterium, protoplasts, etc. Examples of liposomes include lipofectamine and cationic liposomes, and examples of transfection aids include atelocollagen, nanoparticles, and polymers.

[0170] The introduction step may include a cell culture step after the introduction. The culture can be carried out using known methods or modified methods and conditions, depending on the type of cell. The culture medium used in the culture can be described in the kit.

[0171] In the culture step, the culture pH is not particularly limited as long as it is the optimal pH for the proliferation of the cells.

[0172] In the culture step described above, the culture temperature is not particularly limited as long as it is the optimal temperature for the proliferation of the cells.

[0173] In the culture step, the culture time is not particularly limited as long as it is the time during which the complex is expressed. The culture time is, for example, the time during which the protein and guide RNA of this disclosure are expressed and the target DNA can be modified.

[0174] The culture may be carried out, for example, by static culture or by shaking culture.

[0175] <Manufacturing Method> In another embodiment, the Disclosure provides a method for producing cells or non-human animals in which target DNA has been modified. The method for producing cells or non-human animals in which target DNA has been modified according to the Disclosure includes an introduction step of introducing the proteins, polynucleotides, and / or vectors of the Disclosure and a guide RNA to a target site of target DNA into cells or non-human animals. The manufacturing methods of the Disclosure can be described by reference to the descriptions of the proteins, polynucleotides, vectors, vector systems, compositions, kits, and modification methods of the Disclosure.

[0176] <Cells or Non-Human Animals> In another embodiment, the Disclosure provides cells or non-human animals. The cells or non-human animals of the Disclosure are produced by the production methods of the Disclosure. The cells or non-human animals of the Disclosure can be produced by referring to the descriptions of proteins, polynucleotides, vectors, vector systems, compositions, kits, cells, modification methods, and production methods of the Disclosure.

[0177] The present disclosure will be described in detail below using examples, but the present disclosure is not limited to the embodiments described in the examples.

[0178] [Example 1] (1) Construction of vectors containing polynucleotides encoding each protein Vectors containing polynucleotides encoding each of the proteins of the present disclosure (SEQ ID NOs: 1 to 15) were constructed. Specifically, a coding sequence was inserted downstream of the CBh promoter of an expression vector into which a CAG-TurboRFP expression cassette was inserted. The coding sequence was obtained by adding SV40 NLS (SEQ ID NOs: 34) to the N-terminus of each of the proteins of the present disclosure (SEQ ID NOs: 1 to 15), nucleoplasmin NLS (SEQ ID NOs: 35), a GS linker, and a 3×HA tag (SEQ ID NOs: 36) to the C-terminus of the proteins of the present disclosure. In addition, a polynucleotide encoding a guide RNA (SEQ ID NOs: 37) was inserted downstream of the U6 promoter to obtain the vector of the present disclosure. Furthermore, as other expression vectors, expression vectors were also constructed in which polynucleotides encoding the amino acid sequences of SEQ ID NOs: 16 to 19 (ST9: SEQ ID NOs: 16, ST9_2: SEQ ID NOs: 17, ST9_3LA: SEQ ID NOs: 18, ST9_S: SEQ ID NOs: 19) were inserted instead of the proteins of the present disclosure.

[0179] SV40 NLS (Sequence ID 34) PKKKRKV

[0180] nucleoplasmin NLS (SEQ ID NO: 35) KRPAATKKAGQAKKKK

[0181] 3xHA tag (sequence number 36) YPYDVPDYAYPYDVPDYAYPYDVPDYA

[0182] Guide RNA (SEQ ID NO: 37) 5'-GUCAAAAGACCUUUGGAAUUUCUACUCUUGUAGAUUCUGUCCCCUCCACCCCACAG-3'

[0183]

[0184]

[0185]

[0186]

[0187] (2) Construction of a vector for EGxxFP assay detection A vector was prepared containing the EGFP amino group terminal region, PAM sequences (SEQ ID NOs. 24-32), and inserted downstream of the CAG promoter in the order of the cleaved base sequence and the EGFP carboxyl group terminal region. This EGxxFP assay detection vector is designed so that when cleaved by genome editing, the EGFP sequence is repaired and green fluorescence is observed.

[0188] (3) EGxxFP assay The cleavage activity of the proteins of this disclosure at target sites was examined by the EGxxFP assay. Specifically, HEK293T cells were seeded in 24-well plates and incubated at 37°C for 1.5 to 2 days at CO2 2The cells were cultured under incubator conditions. The culture medium used was DMEM (High Glucose) (Sigma-Aldrich), 10% FBS (Gibco), and a 1x penicillin-streptomycin mixture (Nacalitex). After the culture, 75 μl of Solution A (600 μl Opti-MEM (Gibco), 3 μl 2 ng / nl EGxxFP assay detection vector) was dispensed into each tube, and 0.75 μl of each expression vector constructed in Example 1(1) at 0.4 ng / nl was mixed in. After mixing, 75 μl of Solution B (600 μl Opti-MEM (Gibco), 60 μl PEIMax (Polysciences)) was added to prepare the mixture. After preparation, the mixture was allowed to stand at room temperature (hereinafter referred to as approximately 24°C) for 20 minutes. Subsequently, the culture medium in the 24-well plate was replaced with fresh medium, and 50 μl of the mixture was added to each well. After the addition, the 24-well plate was gently shaken. After shaking, the cells were cultured for 1.5 to 2 days and then harvested by trypsin treatment. After harvesting, a diluted solution of Passive Lysis 5X Buffer (Promega) was added, and the cells were lysed by vigorous mixing. After lysis, the cells were centrifuged at 12,000 rpm for 5 minutes, and the supernatant was collected. 50 μl of the supernatant from each sample was added to two wells of a 96-well plate, and the fluorescence intensities of EGFP and RFP contained in the supernatant were detected using a plate reader (CYTATION® 3 imaging reader (BioTek)). In the detection, the EGFP value was standardized by the RFP value. A sample using an expression vector that did not contain the polynucleotide encoding the guide RNA of this disclosure was used as a negative control, and the negative control value was subtracted from the value of each sample. These results are shown in Figures 1 to 4.

[0189] Figures 1-4 are graphs showing the results of the EGxxFP assay. In Figures 1-4, the vertical axis represents the EGFP value standardized by the RFP value, and the horizontal axis represents the sample type. As shown in Figures 1-4, the proteins disclosed herein were found to have higher nuclease activity compared to the control.

[0190] [Example 2] (1) Construction of a vector containing a polynucleotide encoding a modified protein with different amino acids having a cleavage active site A vector containing a polynucleotide encoding a protein with a modified cleavage active site was constructed. Specifically, a coding sequence was obtained by inserting downstream of the CBh promoter of an expression vector, in which the 969th amino acid of ST9, glutamic acid (E), shown underlined, is replaced with glycine (G), and the 1040-1043rd amino acids, valine, proline, alanine, and alanine, are replaced with alanine, aspartic acid, glutamic acid, and arginine, into the N-terminus of a protein (CR-FD2, SEQ ID NO: 38), and adding SV40 NLS (SEQ ID NO: 34) to the C-terminus of the protein of this disclosure, nucleoplasmin NLS (SEQ ID NO: 35), a GS linker, and a 3×HA tag (SEQ ID NO: 36). The CR-FD2 has mutations in the cleavage active site and adjacent sites (Figure 5). Therefore, CR-FD2 has a different cleavage active site than the non-mutated protein (ST9).

[0191]

[0192] (2) The cleavage activity of modified proteins having mutations at the EGxxFP assay site and adjacent sites was examined using the EGxxFP assay. Specifically, the assay was carried out in the same manner as the EGxxFP assay in Example 1(3), except that the vector constructed in Example 2(1) was used instead of the vector constructed in Example 1(1). The results are shown in Figure 6.

[0193] Figure 6 is a graph showing the results of the EGxxFP assay. In Figure 6, the vertical axis represents the EGFP value standardized by the RFP value, and the horizontal axis represents the sample type. As shown in Figure 6, the protein of this disclosure, with its improved cleavage active site, was found to possess nuclease activity.

[0194] [Example 3] (1) Construction of a vector containing a polynucleotide encoding a modified protein with alanine substitution A vector containing a polynucleotide encoding a modified protein in which the amino acid responsible for the cleavage active site is substituted with alanine was constructed. Specifically, as shown underlined, glutamic acid, the 969th amino acid of ST9, is substituted with alanine, and the amino acids 1040 to 1043, valine, proline, alanine, and alanine, are substituted with alanine, aspartic acid, glutamic acid, and arginine, respectively, of a protein (CR-FD3, SEQ ID NO: 39) had SV40 NLS (SEQ ID NO: 34) added to the N-terminus, and nucleoplasmin NLS (SEQ ID NO: 35), a GS linker, and a 3×HA tag (SEQ ID NO: 36) added to the C-terminus of the protein of this disclosure. This coding sequence was inserted downstream of the CBh promoter of the expression vector to obtain the vector of this disclosure. Note that CR-FD3 has an alanine mutation at the cleavage active site (Figure 7). Therefore, the cleavage active site of CR-FD3 differs from that of the unmutated case (ST9). Furthermore, CR-FD3 can be described as a protein in which mutations have been introduced into the amino acids constituting the active site, as in CR-FD2.

[0195]

[0196] (2) The cleavage activity of modified proteins having an alanine mutation at the EGxxFP assay site was examined using the EGxxFP assay. Specifically, the assay was carried out in the same manner as the EGxxFP assay in Example 1(3), except that the vector constructed in Example 3(1) was used instead of the vector constructed in Example 1(1). The results are shown in Figure 8.

[0197] Figure 8 is a graph showing the results of the EGxxFP assay. In Figure 8, the vertical axis shows the EGFP value standardized by the RFP value, and the horizontal axis shows the sample type. As shown in Figure 8, it was found that proteins in which the 969th amino acid, glutamic acid, was replaced with alanine, and the 1040-1043rd amino acids, valine, proline, alanine, and alanine, were replaced with alanine, aspartic acid, glutamic acid, and arginine, respectively, did not possess nuclease activity.

[0198] [Example 4] The cleavage activity of the protein of this disclosure at the target site when the PAM sequence is TTTC (SEQ ID NO: 40) was investigated by EGxxFP assay. Specifically, the EGxxFP assay was performed in the same manner as in Example 1(3) using each expression vector constructed in Example 1(1) and the EGxxFP assay detection vector of Example 1(2) where the PAM sequence is TTTC. These results are shown in Figure 9.

[0199] Figure 9 is a graph showing the results of the EGxxFP assay. In Figure 9, the vertical axis represents the EGFP value standardized by the RFP value, and the horizontal axis represents the sample type. As shown in Figure 9, the protein of this disclosure was found to recognize the PAM sequence TTTC and to exhibit higher nuclease activity compared to the control.

[0200] Although the present disclosure has been described above with reference to embodiments and examples, the present disclosure is not limited to the above embodiments and examples. Various modifications to the structure and details of the present disclosure are possible, as can be understood by those skilled in the art within the scope of the present disclosure.

[0201] This application claims priority based on Japanese Patent Application No. 2024-195971, filed on 8 November 2024, and incorporates all of its disclosures herein.

[0202] The patents, patent applications, and documents cited herein are incorporated herein by reference in the same manner as their contents are specifically described herein.

[0203] <Note> Some or all of the above embodiments and examples may be described as follows, but are not limited to the following. <Protein> (Note 1) The following proteins (a1), (a2), or (a3): (a1) A protein consisting of any amino acid sequence of SEQ ID NOs. 1 to 19; (a2) A protein having nuclease activity, consisting of an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, or added in any amino acid sequence of SEQ ID NOs. 1 to 19; (a3) ​​A protein having nuclease activity, consisting of an amino acid sequence having 80% or more identity with any amino acid sequence of SEQ ID NOs. 1 to 19. (Note 2) The protein according to Note 1, which recognizes a protospacer adjacent motif (PAM) sequence having a base sequence selected from the group consisting of SEQ ID NOs. 20 to 32. (Note 3) The protein in (a1) above is one of the proteins in (b1) to (b6) or (b7) below, as described in Note 1 or 2: (b1) A protein in which the 1136th amino acid E in the amino acid sequence of SEQ ID NO: 1 or 6 is replaced with G; (b2) A protein in which the 1129th amino acid E in the amino acid sequence of SEQ ID NO: 2 to 5, 14, or 15 is replaced with G; (b3) A protein in which the 1127th amino acid E in the amino acid sequence of SEQ ID NO: 7 or 9 is replaced with G; (b4) A protein in which the 1120th amino acid E in the amino acid sequence of SEQ ID NO: 8 or 10 is replaced with G; (b5) A protein in which the 1128th amino acid E in the amino acid sequence of SEQ ID NO: 11 is replaced with G; (b6) A protein in which the 969th amino acid E in the amino acid sequence of SEQ ID NO: 16 or 18 is replaced with G; (b7) A protein in which the amino acid E at position 960 in the amino acid sequence of SEQ ID NO: 17 is replaced with G.(Note 4) The protein in (a1) above is one of the proteins described in (c1) to (c6) or (c7) below, as described in any of Notes 1 to 3: (c1) A protein in which the amino acid VPAA at positions 1207 to 1210 in the amino acid sequence of SEQ ID NO: 1 or 6 is replaced with ADER; (c2) A protein in which the amino acid VPAA at positions 1200 to 1203 in the amino acid sequence of SEQ ID NO: 2 to 5, 14, or 15 is replaced with ADER; (c3) A protein in which the amino acid sequence VPAA at positions 1198 to 1201 in the amino acid sequence of SEQ ID NO: 7 or 9 is replaced with ADER; (c4) A protein in which the amino acid sequence VPAA at positions 1191 to 1194 in the amino acid sequence of SEQ ID NO: 8 or 10 is replaced with ADER; (c5) A protein in which the amino acid sequence VPAA at positions 1199 to 1202 in the amino acid sequence of SEQ ID NO: 11 is replaced with ADER; (c6) A protein in which the amino acid sequence VPAA at positions 1040 to 1043 in the amino acid sequence of SEQ ID NO: 16 or 18 is replaced with ADER; (c7) A protein in which the amino acid sequence VPAA at positions 1031 to 1034 in the amino acid sequence of SEQ ID NO: 17 is replaced with ADER. (Note 5) A protein according to any one of Notes 1 to 4 that is an RNA-dependent DNA nuclease. (Note 6) A protein according to any one of Notes 1 to 5 that further has an enzyme activity domain. (Note 7) A protein according to any one of Notes 1 to 6, wherein the enzyme activity domain is located at the 5' end and / or 3' end of the protein according to any one of Notes 1 to 6. (Note 8) A protein according to any one of Notes 1 to 7, wherein the enzyme activity is selected from the group consisting of nucleic acid base conversion activity, methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity. <Polynucleotide> (Note 9) A polynucleotide comprising a protein coding sequence described in any of Notes 1 to 8. <Vector> (Note 10) A vector comprising the polynucleotide described in Note 9. (Note 11) The vector described in Note 10, further comprising a polynucleotide encoding a guide RNA for a target site.<Vector System> (Note 12) A vector system comprising a first vector containing the polynucleotide described in Note 9, and a second vector containing a polynucleotide encoding a guide RNA for a target site. <Composition> (Note 13) A composition comprising a protein described in any of Notes 1 to 8, a polynucleotide described in Note 9, and / or a vector described in Note 10 or 11, and a guide RNA for a target site or a polynucleotide encoding it. (Note 14) The composition according to Note 13, wherein the guide RNA is complexable with the protein. (Note 15) The composition according to Note 13 or 14, wherein the guide RNA contains a polynucleotide that can hybridize to a target site in target DNA. (Note 16) The composition according to any of Notes 13 to 15, wherein the guide RNA contains crRNA. <Kit> (Note 17) A kit comprising a protein described in any of Notes 1 to 8, a polynucleotide described in Note 9, and / or a vector described in Note 10 or 11, and a guide RNA for a target site or a polynucleotide encoding it. <Cell> (Note 18) A cell comprising a protein described in any of Notes 1 to 8, a polynucleotide described in Note 9, a vector described in Note 10 or 11, and / or a vector system described in Note 12. <Method for modifying target DNA> (Note 19) A method for modifying target DNA, comprising a modification step of modifying the target DNA by contacting the target DNA with a protein described in any of Notes 1 to 8 and a guide RNA for a target site within the target DNA within the cell. (Note 20) The modification method according to Note 19, comprising an introduction step of introducing the protein described in any of Notes 1 to 8 or a polynucleotide encoding it, and the guide RNA or a polynucleotide encoding it, into the cell. (Note 21) The modification method according to Note 19 or 20, wherein in the modification step, the complex of the protein and the guide RNA is targeted to the target DNA, and the target DNA is modified by the complex at the target site.(Note 22) The modification method according to any one of Notes 19 to 21, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA. (Note 23) The modification method according to any one of Notes 19 to 22, wherein the guide RNA comprises crRNA. (Note 24) The modification method according to any one of Notes 19 to 23, wherein the modification is a double-strand break, a single-strand break, a conversion of one or more nucleotides, a deletion and / or insertion of the target DNA. (Note 25) The modification method according to any one of Notes 19 to 24, wherein the cell is an animal cell or a plant cell. (Note 26) The modification method according to any one of Notes 19 to 25, wherein the cell is a eukaryotic cell or a prokaryotic cell. (Note 27) The modification method according to any one of Notes 19 to 26, wherein the target DNA is chromosomal DNA. <Manufacturing Method> (Note 28) A method for producing cells or non-human animals in which target DNA has been modified, comprising an introduction step of introducing a protein described in any one of Notes 1 to 8, a polynucleotide described in Note 9, and / or a vector described in Note 10 or 11, and a guide RNA for the target site into cells or non-human animals. <Cells or Non-Human Animals> (Note 29) Cells or non-human animals produced by the manufacturing method of Note 28.

[0204] As described above, the proteins of this disclosure provide novel proteins that can be used for modifying target DNA. Therefore, the present invention is extremely useful in fields such as medicine, food, and agriculture.

Claims

1. Proteins of the following types (a1), (a2), or (a3): (a1) Proteins consisting of any of the amino acid sequences of SEQ ID NOs: 1-15, 18, and 19; (a2) Proteins having nuclease activity, consisting of amino acid sequences in which 1 to 142 amino acids are deleted, inserted, substituted, or added to any of the amino acid sequences of SEQ ID NOs: 1-15, 18, and 19; (a3) ​​Proteins having nuclease activity, consisting of amino acid sequences having 90% or more identity with any of the amino acid sequences of SEQ ID NOs: 1-15, 18, and 19.

2. The protein according to claim 1, which recognizes a protospacer adjacent motif (PAM) sequence having a base sequence selected from the group consisting of SEQ ID NOs. 20 to 32.

3. The protein according to claim 1 or 2, wherein the protein in (a1) is one of the proteins in (b1) to (b5) or (b6) below: (b1) A protein in which the 1136th amino acid E in the amino acid sequence of SEQ ID NO: 1 or 6 is replaced with G; (b2) A protein in which the 1129th amino acid E in the amino acid sequence of SEQ ID NO: 2 to 5, 14, or 15 is replaced with G; (b3) A protein in which the 1127th amino acid E in the amino acid sequence of SEQ ID NO: 7 or 9 is replaced with G; (b4) A protein in which the 1120th amino acid E in the amino acid sequence of SEQ ID NO: 8 or 10 is replaced with G; (b5) A protein in which the 1128th amino acid E in the amino acid sequence of SEQ ID NO: 11 is replaced with G; (b6) A protein in which the 969th amino acid E in the amino acid sequence of SEQ ID NO: 18 is replaced with G.

4. The protein according to claim 1 or 2, wherein the protein in (a1) is one of the proteins in (c1) to (c5) or (c6) below: (c1) A protein in which the amino acid VPAA at positions 1207 to 1210 in the amino acid sequence of SEQ ID NO: 1 or 6 is replaced with ADER; (c2) A protein in which the amino acid VPAA at positions 1200 to 1203 in the amino acid sequence of SEQ ID NO: 2 to 5, 14, or 15 is replaced with ADER; (c3) A protein in which the amino acid sequence VPAA at positions 1198 to 1201 in the amino acid sequence of SEQ ID NO: 7 or 9 is replaced with ADER; (c4) A protein in which the amino acid sequence VPAA at positions 1191 to 1194 in the amino acid sequence of SEQ ID NO: 8 or 10 is replaced with ADER; (c5) A protein in which the amino acid sequence VPAA at positions 1199 to 1202 in the amino acid sequence of SEQ ID NO: 11 is replaced with ADER; (c6) A protein in which the amino acid sequence VPAA at positions 1040-1043 in sequence number 18 is replaced with ADER.

5. The protein according to claim 1 or 2, which is an RNA-dependent DNA nuclease.

6. The protein according to claim 1 or 2, further comprising an enzyme-active domain.

7. The protein according to claim 1 or 2, wherein the enzyme-active domain is located at the 5' end and / or 3' end of the protein according to claim 1 or 2.

8. The protein according to claim 1 or 2, wherein the enzyme activity is selected from the group consisting of nucleic acid base conversion activity, methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity.

9. A polynucleotide comprising the protein coding sequence described in claim 1.

10. A vector comprising the polynucleotide described in claim 9.

11. The vector according to claim 10, further comprising a polynucleotide encoding a guide RNA for a target site.

12. A vector system comprising a first vector containing the polynucleotide described in claim 9, and a second vector containing the polynucleotide encoding a guide RNA for a target site.

13. A composition comprising the protein according to claim 1, the polynucleotide according to claim 9, and / or the vector according to claim 10, and a guide RNA for a target site or a polynucleotide encoding the same.

14. The composition according to claim 13, wherein the guide RNA is capable of complexing with the protein.

15. The composition according to claim 13, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA.

16. The composition according to claim 13, wherein the guide RNA comprises crRNA.

17. A kit comprising the protein according to claim 1, the polynucleotide according to claim 9, and / or the vector according to claim 10, and a guide RNA for a target site or a polynucleotide encoding the same.

18. A cell comprising the protein according to claim 1, the polynucleotide according to claim 9, the vector according to claim 10, and / or the vector system according to claim 12.

19. A method for modifying target DNA, comprising a modification step of modifying the target DNA by contacting the target DNA with the protein described in claim 1 or 2 and a guide RNA for a target site within the target DNA within a cell.

20. The modification method according to claim 19, comprising an introduction step of introducing the protein or polynucleotide encoding the same according to claim 1 or 2 and the guide RNA or polynucleotide encoding the same into the cell.

21. The modification method according to claim 19, wherein in the modification step, the complex of the protein and the guide RNA is targeted to the target DNA, and the target DNA is modified by the complex at the target site.

22. The modification method according to claim 19, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA.

23. The modification method according to claim 19, wherein the guide RNA includes crRNA.

24. The modification method according to claim 19, wherein the modification is a double-strand break, a single-strand break, a conversion of one or more nucleotides, a deletion and / or insertion of the target DNA.

25. The modification method according to claim 19, wherein the cells are animal cells or plant cells.

26. The modification method according to claim 19, wherein the cells are eukaryotic cells or prokaryotic cells.

27. The modification method according to claim 19, wherein the target DNA is chromosomal DNA.

28. A method for producing cells or non-human animals in which target DNA has been modified, comprising an introduction step of introducing the protein according to claim 1, the polynucleotide according to claim 9, and / or the vector according to claim 10, and a guide RNA for the target site into cells or non-human animals.