Proteins, polynucleotides, vectors, vector systems, compositions, kits, cells, methods for modifying target DNA, and methods for manufacturing.

Novel Cas-like proteins with specific amino acid modifications enable effective editing of target DNA sequences with PAM sequences like NAAN, NAGN, or NGAN, overcoming limitations of conventional CRISPR/Cas systems.

JP2026083891AActive Publication Date: 2026-05-20SETSUROTECH INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SETSUROTECH INC
Filing Date
2024-11-08
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing genome editing technologies using CRISPR/Cas systems struggle to modify target DNA sequences with specific PAM sequences such as NAAN, NAGN, or NGAN, limiting their effectiveness.

Method used

Development of novel Cas-like proteins with amino acid sequences that can recognize and modify target DNA at sites with PAM sequences like NAAN, NAGN, or NGAN, including proteins with specific amino acid modifications such as conservative substitutions.

Benefits of technology

Enables effective modification of target DNA sequences that were previously difficult to modify with conventional Cas proteins, expanding the range of editable sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083891000001_ABST
    Figure 2026083891000001_ABST
Patent Text Reader

Abstract

This provides a novel protein that can be used to modify target DNA. [Solution] The protein of this disclosure is the protein of (a1), (a2), or (a3) ​​below: (a1) A protein consisting of any one of several specific amino acid sequences; (a2) A protein having nuclease activity, comprising an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, or added in any of the amino acid sequences of the aforementioned multiple specific sequences; (a3) A protein having nuclease activity, comprising an amino acid sequence that has 80% or more identity with any of the amino acid sequences of the aforementioned multiple specific sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to proteins, polynucleotides, vectors, vector systems, compositions, kits, cells, methods for modifying target DNA, and production methods.

Background Art

[0002] In recent years, genome editing using systems such as CRISPR / Cas9 (Clustered Regularly Interspaced Palindromic Repeats / CRISPR-associated protein 9) has been widely used. In addition to Cas9, other Cas proteins such as Cas12 have been searched for, and variants of Cas proteins such as MAD7 have been created (Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Therefore, an object of the present disclosure is to provide a new protein that can be used for modifying target DNA.

Means for Solving the Problems

[0005] To achieve the above object, the protein of the present disclosure includes the protein of the following (a1), (a2), or (a3): (a1) A protein consisting of an amino acid sequence of any one of SEQ ID NOs: 1 to 19; (a2) A protein consisting of an amino acid sequence in which one or several amino acids are deleted, inserted, substituted, or added in the amino acid sequence of any one of SEQ ID NOs: 1 to 19 and having nuclease activity; (a3) A protein having nuclease activity, consisting of an amino acid sequence that has 80% or more identity with any of the amino acid sequences of Sequence ID No. 1 to 19.

[0006] The polynucleotides of this disclosure encode proteins of this disclosure.

[0007] The vectors of this disclosure include the polynucleotides of this disclosure.

[0008] The vector system of this disclosure comprises a first vector containing the polynucleotide of this disclosure, It includes a second vector containing a polynucleotide encoding a guide RNA for the target site.

[0009] The compositions of the Disclosure include the proteins of the Disclosure, the polynucleotides of the Disclosure, and / or the vectors of the Disclosure. It includes a guide RNA for the target site or a polynucleotide encoding it.

[0010] The kit of this disclosure comprises the protein of this disclosure, the polynucleotide of this disclosure, and / or the vector of this disclosure, It includes a guide RNA for the target site or a polynucleotide encoding it.

[0011] The cells of this disclosure include the proteins, polynucleotides, vectors, and / or vector systems of this disclosure.

[0012] The method for modifying target DNA according to the present disclosure includes a modification step of modifying the target DNA by bringing the target DNA into contact with the protein of the present disclosure and a guide RNA for a target site within the target DNA, within a cell.

[0013] The present disclosure is a method for producing cells or non-human animals in which target DNA has been modified, An introduction step of introducing the protein of the present disclosure, the polynucleotide of the present disclosure, and / or the vector of the present disclosure, and a guide RNA for the target site into a cell or a non-human animal is included.

Advantages of the Invention

[0014] According to the present disclosure, it is possible to provide a new protein that can be used for modifying target DNA, etc.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0017] As used herein, "nuclease" means an enzyme capable of cleaving the phosphodiester bond between nucleotides of a nucleic acid. Examples of the nuclease include an exonuclease that digests a nucleic acid from the end and an endonuclease that cleaves in the middle of a nucleotide chain. The nuclease can be classified into a DNA nuclease that acts on DNA and an RNA nuclease that acts on RNA.

[0018] As used herein, "protospacer adjacent motif (PAM) sequence" means a sequence adjacent to a sequence targeted by a Cas nuclease in the CRISPR bacterial adaptive immune system. The length and nucleotide sequence of the PAM sequence vary depending on the type of the Cas nuclease.

[0019] As used herein, "coding sequence" means a nucleic acid (RNA or DNA molecule) containing a nucleotide sequence encoding a protein. The coding sequence can be codon-optimized.

[0020] As used herein, "polynucleotide" means a polymer of deoxyribonucleotides (DNA), ribonucleotides (RNA), and / or modified nucleotides. As used herein, when a "polynucleotide" is used in combination with a specific protein, the "polynucleotide" means a polymer of nucleotides encoding the amino acid sequence of the protein. Examples of the polynucleotide include genomic DNA, cDNA, mRNA, etc. The polynucleotide may be, for example, single-stranded or double-stranded. The polynucleotide can be read interchangeably with "nucleic acid" or "oligonucleotide".

[0021] In this specification, “target site” means a region or sequence of polynucleotides targeted by the protein-guide RNA complex of this disclosure.

[0022] In this specification, "nucleic acid modification activity" means activity that can modify bases (e.g., A, T, C, G, or U) within nucleic acid sequences such as DNA and RNA. "Nucleic acid modification" means substitution, conversion, deletion, insertion, addition, etc., of bases within a nucleic acid sequence.

[0023] In this specification, “guide RNA” means RNA that can target the proteins of this disclosure to a target DNA sequence. The guide RNA comprises a crRNA (CRISPR RNA) that binds to the target site and optionally a tracrRNA (trans-activating crRNA) that is involved in the activity of the CRISPR-Cas system, depending on the type of Cas protein. The guide RNA may be, for example, a single-stranded RNA in which the crRNA and tracrRNA are directly or indirectly linked, or it may be two RNAs, the crRNA and the tracrRNA.

[0024] In this specification, "crRNA" means RNA having a base sequence (hereinafter also referred to as "sequence") complementary to the target DNA sequence. The crRNA is an RNA that constitutes the CRISPR / Cas system and is known to perform the function of recognizing the target DNA sequence. The crRNA generally includes a spacer sequence and a repeat sequence. The spacer sequence can also be described as a sequence (guide sequence) that can form a double helix with the complementary strand of the target DNA sequence. The crRNA can form a double helix with, for example, either the sense strand or the antisense strand of the target DNA.

[0025] In this specification, "tracrRNA" means RNA having a sequence complementary to a portion of the crRNA. The sequence complementary to a portion of the crRNA is also called an anti-repeat region (sequence). The tracrRNA has a sequence that includes an anti-repeat region (sequence) (repeating region) followed by one or more hairpin structures, and is capable of forming a stem-loop. In some Cas proteins, the tracrRNA is known to function as a scaffold for binding the Cas protein to the crRNA. The proteins of this disclosure can, for example, directly bind to crRNA.

[0026] In this specification, “hybridize” means annealing with a complementary polynucleotide resulting from nucleotide complementarity, specifically the complementarity of bases in a nucleotide. That is, it means that two polynucleotides can form a non-covalent pair via hydrogen bonds. The hybridize may be, for example, the joining of two complementary base sequences, or substantially complementary base sequences that have one or more mismatched base pairs.

[0027] In this specification, “complementary” means that one polynucleotide can form a nucleotide pair, i.e., a base pair, with another polynucleotide.

[0028] In this specification, "vector" means in vitro or in vivo In this context, it means nucleic acids delivered to host cells, or recombinant plasmids or viruses containing such nucleic acids.

[0029] In this specification, "target DNA sequence" means the base sequence of the target polynucleotide in the target DNA that hybridizes with the guide RNA.

[0030] In this specification, "eukaryote" means an organism that has a nucleus enclosed by a nuclear membrane.

[0031] In this specification, "prokaryotes" refers to organisms that do not have a nucleus enclosed by a nuclear membrane.

[0032] Sequence information for proteins or nucleic acids (e.g., DNA or RNA) encoding them, as described herein, can be obtained from sources such as the Protein Data Bank, UniProt, Ensembl, or GenBank. RNA nucleic acid sequences can also be obtained from corresponding DNA nucleic acid sequences using appropriate sequence conversion software.

[0033] The following examples illustrate the provisions of this disclosure, but the disclosure is not limited to these examples and can be modified as appropriate. Furthermore, unless otherwise specified, the explanations in this disclosure are interchangeable with each other. In this specification, the expression "~" includes the numerical or physical values ​​before and after it. Also, in this specification, the expression "A and / or B" includes "A only," "B only," and "both A and B."

[0034] <Protein> In one embodiment, the present disclosure provides a novel protein that can be used for modifying target DNA. The protein of the present disclosure is the protein described below (a). The protein described below (a) may have nuclease activity or niccas activity. (a) Proteins of the following (a1), (a2), or (a3): (a1) A protein consisting of any of the amino acid sequences of SEQ ID NOs: 1 to 19; (a2) Proteins consisting of amino acid sequences in which one or more amino acids are deleted, inserted, substituted, or added in any of the amino acid sequences of SEQ ID NOs: 1 to 19; (a3) A protein consisting of an amino acid sequence that has 80% or more identity with any of the amino acid sequences of sequence numbers 1 to 19.

[0035] The Cas protein recognizes the PAM sequence of target DNA and modifies the target DNA through the nuclease activity of the Cas protein. The PAM sequences that the Cas protein can recognize vary depending on the type and origin of the Cas protein. However, it is known that there is a bias in the PAM sequences that the Cas protein can recognize. For example, NAAN, NAGN, or NGAN (N=A, T, G, or C; the same applies hereinafter) are known to be difficult to recognize by the Cas protein, especially the commonly used Cas9 and Cas12. As a result of diligent searching for various Cas proteins, the inventors have found a protein that can recognize NAAN, NAGN, or NGAN as the PAM sequence, and have completed this disclosure. According to this disclosure, a novel protein that can recognize NAAN, NAGN, or NGAN as the PAM sequence can be provided. Therefore, according to this disclosure, target DNA can be modified even at target sites having base sequences that were difficult to modify with conventional Cas proteins, such as NAAN, NAGN, or NGAN.

[0036] The protein (a1) described above is a protein designed by the present inventors as a Cas-like protein.

[0037] ST9.5_13 (Sequence ID 1) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0038] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 1, then in the protein of (a2), "one or several" may mean, for example, any number within the range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may mean, for example, 1 to 287, 1 to 284, 1 to 215, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2. In this disclosure, the numerical range of the number may mean, for example, all positive integers belonging to that range.

[0039] In the protein described in (a2) above, substitutions are preferably conservative substitutions. A conservative substitution means the substitution of an amino acid residue with an amino acid residue having a similar side chain. Examples of such conservative substitutions include substitutions between amino acid residues having basic side chains such as lysine, arginine, and histidine; substitutions between amino acid residues having acidic side chains such as aspartic acid and glutamic acid; substitutions between amino acid residues having non-charged polar side chains such as glycine, asparagine, glutamine, serine, threonine, tyrosine, and cysteine; substitutions between amino acid residues having non-polar side chains such as alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, and tryptophan; substitutions between amino acid residues having β-branched side chains such as threonine, valine, and isoleucine; substitutions between amino acid residues having aromatic side chains such as tyrosine, phenylalanine, tryptophan, and histidine; and so on (the same applies hereinafter).

[0040] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 1, then in the proteins (a1) to (a3), for example, the 1136th amino acid E (glutamic acid) or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G (glycine). In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1209th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A, as indicated by the underline, is substituted with E (glutamic acid), it is even more preferable that the 1207th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1208th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1209th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with E (glutamic acid), and the 1210th amino acid A (alanine) or the amino acid residue corresponding to the amino acid A is substituted with R (arginine).

[0041] ST9.5_14 (Sequence ID 2) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0042] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 2, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 2, for example, "one or several" may be 1-286, 1-214, 1-143, 1-142, 1-71, 1-57, 1-43, 1-28, 1-14, 1-10, 1-7, 1-6, 1-5, 1-4, 1-3, 1 or 2.

[0043] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 2, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0044] ST9.5_38 (Sequence No. 3) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0045] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 3, then in the protein of (a2), "one or several" may mean, for example, any range in which (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 3, "one or several" may mean, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0046] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 3, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0047] ST9.5_42 (Sequence ID 4) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0048] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 4, then in the protein of (a2), "one or several" may mean, for example, that (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 4, "one or several" may mean, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0049] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 4, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0050] ST9.5_44 (Sequence ID 5) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0051] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 5, then in the protein of (a2), "one or several" may mean, for example, that (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 5, "one or several" may mean, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0052] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 5, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0053] ST9.5_58 (Sequence ID 6) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0054] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 6, then in the protein of (a2), "one or several" may mean, for example, a range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may mean, for example, 1 to 287, 1 to 215, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0055] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 6, then in the proteins (a1) to (a3), for example, the 1136th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1209th amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1207th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1208th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1209th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1210th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0056] ST9.5_2R (Sequence ID 7) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0057] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 7, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 7, for example, 1 to 285, 1 to 214, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0058] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 7, then in the proteins (a1) to (a3), for example, the 1127th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1198th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1199th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0059] ST9.5_214 (Sequence ID 8) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0060] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 8, then in the protein of (a2), "one or several" may mean, for example, any range in which (a2) is a protein that functions as a protein having nuclease activity. In (a2), "one or several" may mean, for example, 1 to 284, 1 to 213, 1 to 142, 1 to 71, 1 to 56, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2 amino acids in the amino acid sequence of SEQ ID NO: 8.

[0061] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 8, then in the proteins (a1) to (a3), for example, the 1120th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1191st amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1192nd amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1194th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0062] ST9.5_2W (Sequence ID 9) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0063] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 9, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 9, for example, 1 to 285, 1 to 214, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0064] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 9, then in the proteins (a1) to (a3), for example, the 1127th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1198th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1199th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1200th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0065] ST9.5_243 (Sequence ID 10) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0066] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 10, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 10, "one or several" may be, for example, 1 to 284, 1 to 213, 1 to 142, 1 to 71, 1 to 56, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0067] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 10, then in the proteins (a1) to (a3), for example, the 1120th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1191st amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1192nd amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1193rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1194th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0068] ST9.5_2P (Sequence ID 11) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0069] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 11, then in the protein of (a2), "one or several" may mean, for example, that (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 11, "one or several" may mean, for example, 1 to 285, 1 to 214, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0070] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 11, then in the proteins (a1) to (a3), for example, the 1128th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1199th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1200th amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1201st amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0071] ST9.5_S3 (Sequence ID 12)

[0072] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 12, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 12, for example, "one or several" may be 1-286, 1-214, 1-143, 1-142, 1-71, 1-57, 1-42, 1-28, 1-14, 1-10, 1-7, 1-6, 1-5, 1-4, 1-3, or 1 or 2.

[0073] ST9.5_S4 (Sequence ID 13) If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 13, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 13, "one or several" may be, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 42, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0074] ST9.5_3LA1 (Sequence ID 14) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVKQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0075] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 14, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In (a2) or (a3), "one or several" may be any range in the amino acid sequence of SEQ ID NO: 14, such as 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0076] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 14, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0077] ST9.5_3LA4 (Sequence ID 15) E DLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVKQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0078] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 15, then in the protein of (a2), "one or several" may mean, for example, any range in which (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 15, "one or several" may mean, for example, 1 to 286, 1 to 214, 1 to 143, 1 to 142, 1 to 71, 1 to 57, 1 to 43, 1 to 28, 1 to 14, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 or 2.

[0079] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 15, then in the proteins (a1) to (a3), for example, the 1129th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1200th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0080] ST9 (Sequence ID 16) MGEKTNYFNNFIGISPLTKTLRNALIPTTEITQKHIIEYGIIKDDELRKENRQILKTIMDDYYRSFLTEKSAIHDIDWKPLFLEMENELRHGNNKVNLEKEQKAKRKAINKYFSEDERYKKMFSAKLLSDILPEFVIHNDEYSAEEKEEKQVIKLFSRFATSLKEYFRNRSNTFSADNISTSACHRIVNDNASIFLENVMAYRKIISSLPADELDKLESIQDKLKIKSISEVYTYDNYGKYITQEGIDLYNDICGKVNSFMNLYCQQNKENKNVYKMRKLHKQILCIAKTSYEVGEGYTSDEEVLEVFRNTLNKSEIFSSIKKLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEYDDIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLKEIIIQKVDEIYKVYGSSEKLFDADFVLEKSLKNDAVVAIMKDLLDSVKSFENYIKAFFGEGK ETNRDESFYGDFVLAYDILLKVDHIYDAIRNYVTQKPYSTEKFKLNFNSPTLARGWSKSKEYSNNAIILSKDGILYLGIFNVKNKPDKKIIEGHIRENDGDYKKMVYNLLPGANKMLPKVLISSKSGVETYKPSDYILEGYSNKHLKSSDTFDINYCHDLIDYYKDCIDIHPEWKNFDFNFSDTADFEDISGFYREAVEAQGYKIDWTYISGENIEELQEKGQLFLFKIYNKDFSTKSTGTD NLHTMYLKNLFSEENLKDVVLKLNGEAEIFYRKSSIKNPIIHKKGSMLVNTYIEEEEEHDGKIVEVRKIVPEEIYKELYNHFNNNGKAELSEEAKRLESLVCHEAAKDIVKDCRYTYDKFFIHLPMTINFKASGSSSLNDMMLQYISGQNDMHIIGIDRGERNLIYVSVIDAHGNIVKQKSFNIVAGYDYQEKLKMQEGARQNARKEWKIKEIKEGYLSLVIHEIAQMVIEYNAIIAM EDLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0081] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 16, then in the protein of (a2), "one or several" may mean, for example, any range in which (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 16, "one or several" may mean, for example, 1 to 254, 1 to 190, 1 to 142, 1 to 127, 1 to 63, 1 to 50, 1 to 38, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0082] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 16, then in the proteins (a1) to (a3), for example, the 969th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1040th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1041st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1043rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0083] ST9_2 (Sequence ID 17) MGEKTNYFNNFIGISPLTKTLRNALIPTTEITQKHIIEYGIIKDDELRKENRQILKTIMDDYYRSFLTEKSAIHDIDWKPLFLEMENELRHGNNKVNLEKEQKAKRKAINKYFSEDERYKKMFSAKLLSDILPEFVIHNDEYSAEEKEEKTQVIKLFSRFATSLKEYFRNRSNTFSADNISTSACHRIVNDNASIFLENVMAYRKIISSLPADELDKLEASIQDKLKIKSISEVYTYDNYGKYITQEGIDLYNDICGKVNSFMNLYCQQNKENKNVYKMRKLHKQILCIAKTSYEVPYRFESDEELYSSVNEFLSDVEKKGIIERLQKIGENYDKYELDKVYIAGKYYETFSQKVYKEWNTINSALEIYYTNTLHGSGKSKAEKVRKAVKNDSQKSIAQINELVAEYNLCSDESVKAQNYISEITAIIGQSNMHPLEYNPEVNLVEDKGKASELKKSLDVIMNIMHWCSAFITEEIVSK DAEFYSEIEDIYDELTPVVKLYNMVRNYVTQKPYSTEKFKLNFNSPTLARGWSKSKEYSNNAIILSKDGILIYLGIFNVKNKPDKKIIEGHIRENDGDYKKMVYNLLPGANKMLPKVLISSKSGVETYKPSDYILEGYKSNKHLKSSDTFDINYCHDLIDYYKDCIDIHPEWKNFDFNFSDTADFEDISGFYREVEAQGYKIDWTYISGENIEELQEKGQLFLFKIYNKDFSTKSTGTDNLHTMYLKNLFSEENLKDVVLKLNGEAEIFYRKSSIKNPIIHKKGSMLVNKTYIEEEHDGKIVEVRKIVPEEIYKELYNHFNNNGKAELSEEAKRLESLVCHEAAKDIVKDCRYTYDKFFIHLPMTINFKASGSSSLNDMMLQYISGQNDMHIIGIDRGERNLIYVSVIDAHGNIVKQKSFNIVAGYDYQEKLKMQEGARQNARKEWKEIGKIKEIKEGYLSLVIHEIAQMVIEYNAIIAM EDLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0084] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 17, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 17, "one or several" may be, for example, 1 to 252, 1 to 189, 1 to 142, 1 to 126, 1 to 63, 1 to 50, 1 to 37, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0085] If the protein of this disclosure is a protein consisting of the amino acid sequence of Sequence ID No. 17, then in the proteins (a1) to (a3), for example, the 960th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1033rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1031st amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1032nd amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1033rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1034th amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0086] ST9_3LA (Sequence ID 18) MGKNTNYFNRFIGISSSQKTLRNALIPTEITQKHIIEYGIIKDDELRKENRQILKTIMDDYYRSFLTEKSAIHDIDWKPLFLEMENELRHGNNKVNLEKEQKAKRKAINKYFSEDERYKKMFSAKLLSDILPEFVIHNDEYSAEEKEEKQVIKLFSRFATSLKEYFRNRSNTFSADNISTSACHRIVNDNASIFLENVMAYRKIISSLPADELDKLESIQDKLKIKSISEVYTYDNYGKYITQEGIDLYNDICGKVNSFMNLYCQQNKENKNVYKMRKLHKQILCIAKTSYEVGEGYTSDEEVLEVFRNTLNKSEIFSSIKKLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEYDDIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLKEIIIQKVDEIYKVYGSSEKLFDADFVLEKSLKNDAVVAIMKDLLDSVKSFENYIKAFFGEGK ETNRDESFYGDFVLAYDILLKVDHIYDAIRNYVTQKPYSTEKFKLNFNSPTLARGWSKSKEYSNNAIILSKDGILILYLGIFNVKNKPDKKIIEGHIRENDGDYKKMVYNLLSGSAKMLPKVVFAGSNEKIFGHLISKRILEIREKKLYTAAAGDRKAVAEWIDFMKSAIAIHPEWNEYFKFKFKNTAEYDNANKFYEDIDKQTYKIDWTYISGENIEELQEKGQLFLFKIYNKDFSTKSTGTDNLHTMYLKNLFSEENLKDVVLKLNGEAEIFYRKSSIKNPIIHKKGSMLVNKTYIEEEHDGKIVEVRKIVPEEIYKELYNHFNNGKAELSEEAKRLESLVCHEAAKDIVKDYRYTYDKFFIHLPMVINFKAGDSNSLNDMILQYISGQNDMHIIGIDRGERNLIYVSVIDVHGNIVKQKSFNIVGGYDYQEKLKMQEGARQNARKEWKEIGKIKEIKEGYLSLVIHEIAQMVIEYNAIIAM EDLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY VPAA YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVKQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0087] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 18, then in the protein of (a2), "one or several" may mean, for example, that (a2) is a protein that functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 18, "one or several" may mean, for example, 1 to 254, 1 to 190, 1 to 142, 1 to 127, 1 to 63, 1 to 50, 1 to 38, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0088] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 18, then in the proteins (a1) to (a3), for example, the 969th amino acid E or the amino acid residue corresponding to amino acid E, indicated by the underline, may be substituted with G. In this case, since the proteins (a1) to (a3) ​​can retain nuclease activity, it is preferable that the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A, indicated by the underline, is substituted with E (glutamic acid), and it is even more preferable that the 1040th amino acid V (valine) or the amino acid residue corresponding to V is substituted with A (alanine), the 1041st amino acid P (proline) or the amino acid residue corresponding to P is substituted with D (aspartic acid), the 1042nd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with E (glutamic acid), and the 1043rd amino acid A (alanine) or the amino acid residue corresponding to amino acid A is substituted with R (arginine).

[0089] ST9_S (Sequence ID 19)

[0090] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 19, then in the protein of (a2), "one or several" may be any range in which (a2) functions as a protein having nuclease activity. In the amino acid sequence of SEQ ID NO: 19, "one or several" may be, for example, 1 to 254, 1 to 190, 1 to 142, 1 to 127, 1 to 63, 1 to 50, 1 to 38, 1 to 25, 1 to 12, 1 to 10, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 or 2.

[0091] In the protein (a3) ​​described above, "identity" is defined as, for example, within the range in which (a3) ​​functions as a protein having nuclease activity. The "identity" of (a3) ​​is, for example, 70%, 75%, 80% or more, 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, and 99% or more relative to the amino acid sequence of (a1). The "identity" can be calculated, for example, by using the default parameters in the homology algorithm BLAST (http: / / www.ncbi.nlm.nih.gov / BLAST / ) of the National Center for Biotechnology Information (NCBI) (the same applies hereinafter).

[0092] The proteins of this disclosure may recognize, for example, protospacer-adjacent motif (PAM) sequences. Examples of PAM sequences include TTTV (SEQ ID NO: 20, V=A, G, or C; the same applies hereinafter), NAAN (SEQ ID NO: 21), NAGN (SEQ ID NO: 22), or NGAN (SEQ ID NO: 23). NAAN (SEQ ID NO: 21), NAGN (SEQ ID NO: 22), or NGAN (SEQ ID NO: 23) are generally known to be difficult to recognize with Cas9 or Cas12. Examples of NGAN (SEQ ID NO: 23) include AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), etc. Examples of NAGN (SEQ ID NO: 22) include AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), CAGC (SEQ ID NO: 30), and examples of NAAN (SEQ ID NO: 21) include GAAC (SEQ ID NO: 31), CAAC (SEQ ID NO: 32), etc.

[0093] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 1, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26), GAGC (SEQ ID NO: 29), and GAAC (SEQ ID NO: 31).

[0094] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 2, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25) and GGAC (SEQ ID NO: 26).

[0095] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 3, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), GAGC (SEQ ID NO: 29), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0096] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 4, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), CAGC (SEQ ID NO: 30), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0097] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 5, then the PAM sequence of the protein in (a) preferably includes, for example, GAGC (SEQ ID NO: 29).

[0098] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 6, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26) and GAGC (SEQ ID NO: 29).

[0099] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 7, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0100] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 8, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0101] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 9, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), AAGC (SEQ ID NO: 28), and CAAC (SEQ ID NO: 32).

[0102] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 10, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26).

[0103] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 11, then the PAM sequence of the protein in (a) preferably includes, for example, GAGC (SEQ ID NO: 29).

[0104] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 12, then the PAM sequence of the protein in (a) preferably includes, for example, AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), CAGC (SEQ ID NO: 30), and GAAC (SEQ ID NO: 31).

[0105] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 13, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26) and AAGC (SEQ ID NO: 28).

[0106] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 14, then the PAM sequence of the protein in (a) preferably includes, for example, GGAC (SEQ ID NO: 26).

[0107] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 15, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), and GAAC (SEQ ID NO: 31).

[0108] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 16, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), CAGC (SEQ ID NO: 30), GAAC (SEQ ID NO: 31), and CAAC (SEQ ID NO: 32).

[0109] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 17, then the PAM sequence of the protein in (a) preferably includes, for example, TTTV (SEQ ID NO: 20), AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), GAGC (SEQ ID NO: 29), and CAAC (SEQ ID NO: 32).

[0110] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 18, then the PAM sequence of the protein in (a) preferably includes, for example, TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), and GAAC (SEQ ID NO: 31).

[0111] If the protein of this disclosure is a protein consisting of the amino acid sequence of SEQ ID NO: 19, then the PAM sequence of the protein in (a) preferably includes, for example, AGAC (SEQ ID NO: 24), TGAC (SEQ ID NO: 25), GGAC (SEQ ID NO: 26), CGAC (SEQ ID NO: 27), AAGC (SEQ ID NO: 28), CAGC (SEQ ID NO: 30), and GAAC (SEQ ID NO: 31).

[0112] The protein of this disclosure may be, for example, an RNA-dependent DNA nuclease. The protein of this disclosure may form a complex with, for example, a guide RNA as described below, and function as a DNA nuclease. The nuclease may be, for example, an endonuclease.

[0113] The proteins of this disclosure may further have, for example, an enzyme activity domain. If the proteins of this disclosure have an enzyme activity domain, they do not have, for example, nuclease activity. In this case, the proteins of this disclosure may or may not have, for example, nickase activity. Examples of such proteins include the proteins (1) to (7) below. The enzyme activity includes, for example, nucleic acid modification activities such as methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity; transcription activation activity, transcription repression activity, transcription release factor activity, DNA integration activity, nucleic acid binding activity, etc. (1) In the amino acid sequence of SEQ ID NO: 1 or 6, it is preferable that the 1136th amino acid E is replaced with something other than G, for example, A, and further, the 1209th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1207th amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1208th amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1209th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1210th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (2) In the amino acid sequence of Sequence IDs 2-5, 14, or 15, it is preferable that the 1129th amino acid E is replaced with something other than G, for example, A, and further, the 1202nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1200th amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1201st amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1202nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1203rd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (3) In the amino acid sequence of SEQ ID NO: 7 or 9, it is preferable that the amino acid sequence E at position 1127 is replaced with something other than G, for example, A, and further, the amino acid A (alanine) at position 1200 or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the amino acid V (valine) at position 1198 or the amino acid residue corresponding to said V is replaced with A (alanine), the amino acid P (proline) at position 1199 or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the amino acid A (alanine) at position 1200 or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the amino acid A (alanine) at position 1201 or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (4) In the amino acid sequence of SEQ ID NO: 8 or 10, it is preferable that the amino acid sequence E at position 1120 is replaced with something other than G, for example, A, and further, the amino acid A (alanine) at position 1193 or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the amino acid V (valine) at position 1191 or the amino acid residue corresponding to said V is replaced with A (alanine), the amino acid P (proline) at position 1192 or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the amino acid A (alanine) at position 1193 or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the amino acid A (alanine) at position 1194 or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (5) In the amino acid sequence of Sequence ID No. 11, it is preferable that the 1128th amino acid sequence E is replaced with something other than G, for example, A, and further, the 1201st amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1199th amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1200th amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1201st amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1202nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (6) In the amino acid sequence of SEQ ID NO: 16 or 18, it is preferable that the 969th amino acid E is replaced with something other than G, for example, A, and further, the 1042nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1040th amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1041st amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1042nd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1043rd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein; (7) In the amino acid sequence of Sequence ID No. 17, it is preferable that the 960th amino acid E is replaced with something other than G, for example, A, and further, the 1033rd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), the 1031st amino acid V (valine) or the amino acid residue corresponding to said V is replaced with A (alanine), the 1032nd amino acid P (proline) or the amino acid residue corresponding to said P is replaced with D (aspartic acid), the 1033rd amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with E (glutamic acid), and the 1034th amino acid A (alanine) or the amino acid residue corresponding to said amino acid A is replaced with R (arginine), and the protein has reduced nuclease activity compared to the unsubstituted protein.

[0114] The enzyme activity domain is preferably a nucleic acid modification activity domain having nucleic acid modification activity. The nucleic acid modification activity is, for example, an activity that directly or indirectly causes DNA modification by modifying nucleic acids, such as nucleic acid degradation activity, nucleic acid base conversion activity, and DNA hydrolysis activity. The DNA modification is such as a DNA strand cleavage reaction that cleaves the DNA strand catalyzed by the nucleic acid degradation activity, a nucleic acid base conversion reaction that converts substituents on the purine or pyrimidine ring of a nucleic acid base to other groups without cleaving the DNA strand catalyzed by the nucleic acid base conversion activity, and a debase reaction that hydrolyzes the N-glycosidic bond of DNA catalyzed by the DNA glucosidase.

[0115] Nucleic acid base conversion activity includes, for example, deaminase activity. Examples of such deaminase activity include cytidine deaminase activity, adenosine deaminase activity, and guanosine deaminase activity.

[0116] Examples of the DNA glycosylase activity include thymine DNA glycosylase activity, oxoguaning glycosylase activity, alkyladenine DNA glycosylase activity, and so on.

[0117] The enzyme-active domain is located, for example, at the 5' and / or 3' ends of the protein of this disclosure.

[0118] The enzyme-active domain is linked, for example, directly or indirectly to the protein of the Disclosure. The direct linkage is, for example, a covalent linkage, where, for example, the protein of the Disclosure and the enzyme-active domain constitute a fusion protein. The indirect linkage is, for example, a linkage via a binding tag-binding partner. The binding tag-binding partner is, for example, a combination of substances with mutually specific binding properties. Examples of the binding tag-binding partner include a combination of biotin and avidin or streptavidin, a combination of an SH3 domain (SH3) and an SH ligand, a combination of nickel and a His tag, an epitope tag such as a flag(trademark)-tag, HA-tag, T7-tag, V5-peptide-tag, and / or a Myc-tag, and an antibody against the tag. In this case, the protein of the Disclosure and the enzyme-active domain are linked (bound) via the binding tag-binding partner, for example, after transcription and translation from a polynucleotide.

[0119] The proteins of this disclosure may, for example, have the amino acids of the N-terminal amino acid residues protected by a protecting group (e.g., a formyl group, a C1-6 acyl group such as a C1-6 alkanoyl such as an acetyl group, etc.).

[0120] The proteins of this disclosure may, for example, have a pyroglutamine-oxidized N-terminal glutamine residue that can be cleaved and produced in vivo.

[0121] The proteins of this disclosure may, for example, have substituents on the side chains of amino acids within the molecule (e.g., -OH, -SH, amino group, imidazole group, indole group, guanidino group, etc.) protected by appropriate protecting groups (e.g., formyl group, C1-6 acyl group such as a C1-6 alkanoyl group such as an acetyl group, etc.).

[0122] The protein of this disclosure may be, for example, a complex protein such as a glycoprotein to which sugar chains are attached.

[0123] The proteins of this disclosure may, for example, be in the form of salts with an acid or a base. The salts are not particularly limited and include acidic salts, basic salts, etc. Examples of acidic salts include inorganic acid salts such as hydrochloride, hydrobromide, sulfate, nitrate, and phosphate; organic acid salts such as acetate, propionate, tartrate, fumarate, maleate, malate, citrate, methanesulfonate, and p-toluenesulfonate; and amino acid salts such as aspartate and glutamate. Examples of basic salts include alkali metal salts such as sodium salt and potassium salt; and alkaline earth metal salts such as calcium salt and magnesium salt.

[0124] The proteins disclosed herein can be readily produced according to known genetic engineering techniques. For example, the proteins disclosed herein can be produced using PCR, restriction enzyme digestion, DNA ligation techniques, in vitro transcription and translation techniques, recombinant protein production techniques, etc.

[0125] <Polynucleotides> In another embodiment, the Disclosure provides polynucleotides, which include coding sequences for proteins of the Disclosure. The polynucleotides of the Disclosure can be described by reference to the description of proteins of the Disclosure.

[0126] The polynucleotides of the disclosed herein can be designed by substituting corresponding codons based on the amino acid sequence of the protein of the disclosed herein. The base sequence of the polynucleotide of the disclosed herein may be, for example, codon-optimized.

[0127] <Vector> In another embodiment, the Disclosure provides vectors. The vectors of the Disclosure comprise the polynucleotides of the Disclosure. The vectors of the Disclosure can be described by reference to the descriptions of proteins and polynucleotides of the Disclosure.

[0128] The vectors of this disclosure may, for example, include a polynucleotide encoding a guide RNA for a target site in target DNA. In this case, the polynucleotide encoding the guide RNA is hybridizable to the target DNA, i.e., hybridizable to a target sequence in the target DNA. The polynucleotide encoding the guide RNA may further include other components besides the guide RNA.

[0129] The proteins of this disclosure can perform modifications to a target site by forming a complex with a guide RNA having a structure similar to that of a Cas protein guide RNA, for example. The proteins of this disclosure can use, for example, an RNA molecule having a structure or base sequence similar to that of a Cas9 protein guide RNA or a Cas12 protein guide RNA as the guide RNA. The proteins of this disclosure can also use, for example, the guide RNA of a Cas9 protein or the guide RNA of a Cas12 protein as the guide RNA.

[0130] The guide RNA includes, for example, crRNA. Since the crRNA includes, for example, a spacer sequence and a direct repeat sequence, it can also be said that the guide RNA includes, for example, a spacer sequence and a direct repeat sequence. In the guide RNA, the spacer sequence hybridizes to, for example, the target nucleotide sequence of the target DNA. The spacer sequence may be, for example, completely complementary (identical) or substantially complementary (identical) to the target nucleotide sequence of the target DNA. "Substantially complementary" means, for example, that it consists of a base sequence having 50% or more identity with the nucleotide sequence of the target DNA and is capable of hybridizing to the nucleotide sequence of the target DNA. The aforementioned "identity" is, for example, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more with respect to the nucleotide sequence of the target DNA. The substantially complementary sequence is, for example, a sequence in which there are one or two mismatched base pairs between the spacer sequence and the target DNA. In the guide RNA, the length of the spacer sequence is not particularly limited as long as it is long enough to hybridize to the target DNA, and examples include 14-30 nucleotides, 16-28 nucleotides, 16-25 nucleotides, 20-25 nucleotides, 20-22 nucleotides, or 21-23 nucleotides.

[0131] The hybridizes can be detected, for example, by various hybridization assays under stringent conditions. The hybridization assays are not particularly limited, and methods such as those described in Sambrook et al., "Molecular Cloning: A Laboratory Manual 2nd Ed." [Cold Spring Harbor Laboratory Press (1989)] can also be employed.

[0132] The aforementioned "stringent conditions" may be, for example, low-stringent conditions, medium-stringent conditions, or high-stringent conditions. "Low-stringent conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 32°C. "Medium-stringent conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 42°C. "High-stringent conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 50°C. The degree of stringency can be set by those skilled in the art by appropriately selecting conditions such as temperature, salt concentration, probe concentration and length, ionic strength, and time. The "stringent conditions" can also be those described in Sambrook et al.'s "Molecular Cloning: A Laboratory Manual 2nd Ed." [Cold Spring Harbor Laboratory Press (1989)], for example.

[0133] The spacer sequence preferably further includes, for example, a seed sequence. The seed sequence is, for example, the 5' terminal sequence of the PAM sequence in the spacer sequence and is important for the specificity of the guide RNA to the target DNA. In the guide RNA, the seed sequence is, for example, substantially or completely complementary to the nucleotide sequence of the target DNA, preferably the latter. If the seed sequence is substantially complementary to the nucleotide sequence of the target DNA, the substantially complementary sequence is, for example, a sequence in which there is one or two mismatched base pairs between the seed sequence and the nucleotide sequence of the target DNA.

[0134] In the guide RNA, the length of the direct repeat sequence is not particularly limited, and examples include 19 to 40 nucleotides in length, 21 to 38 nucleotides in length, etc.

[0135] The length of the guide RNA sequence is set so that the protein of this disclosure can target the target site of the target DNA. The length of the guide RNA sequence is not particularly limited and may be, for example, 30-220 nucleotides, 40-220 nucleotides, 50-220 nucleotides, 60-220 nucleotides, 70-220 nucleotides, 80-220 nucleotides, 90-220 nucleotides, 100-220 nucleotides, 110-220 nucleotides, 110-210 nucleotides, 110-200 nucleotides, 110-190 nucleotides, 110-180 nucleotides, 110-170 nucleotides, 110-160 nucleotides, or 110-150 nucleotides. If the guide RNA includes crRNA and tracrRNA as separate RNAs, the length of the guide RNA sequence is, for example, the sum of the lengths of both RNAs. If the guide RNA contains both crRNA and tracrRNA within the same RNA (as in the case of sgRNA described later), the length of the guide RNA sequence is, for example, the length of the sgRNA sequence. If the guide RNA contains crRNA, the length of the guide RNA sequence is, for example, the length of the crRNA sequence.

[0136] The guide RNA may be of one type or of multiple types. If there are multiple types of guide RNA, it is preferable that the guide RNA be configured to hybridize to different target DNA sequences, for example. If there are multiple types of guide RNA, for example, by setting the different target DNA sequences so that the guide RNA targets the 5' and 3' sides of the target site or target region, the protein-guide RNA complex of this disclosure can induce a deletion (large deletion) of the target region between the 5' and 3' sides. If there are multiple types of guide RNA, for example, by setting the guide RNA to target the 5' and 3' sides of the target region and introducing donor DNA having a sequence complementary to the 5' side region (5' arm region) and the 3' side region (3' arm region), the protein-guide RNA complex of this disclosure can recombinantly introduce the donor DNA into the region between the 5' and 3' sides.

[0137] The guide RNA may consist of crRNA alone, or it may consist of crRNA and tracrRNA. If the guide RNA consists of crRNA and tracrRNA, the guide RNA may be an sgRNA (single-stranded guide strand) formed by the fusion of the tracrRNA and the crRNA, or it may not be fused; that is, it may consist of two RNAs, crRNA and tracrRNA. If the tracrRNA and the crRNA are not fused, it is preferable that the tracrRNA and the crRNA have substantially or completely complementary sequences, such as the repeat sequence and the anti-repeat region. The structure of the guide RNA (single-stranded or double-stranded) and the presence or absence of tracrRNA can be determined depending on the protein of this disclosure.

[0138] The tracrRNA comprises a repeating sequence (anti-repeat region) and one or more hairpin structures. The number of hairpin structures is not particularly limited and may be, for example, 1 to 10, 1 to 5, 1 to 3, or 2. The tracrRNA is not particularly limited as long as it can bind to the crRNA and form a complex.

[0139] The tracrRNA sequence may, for example, have nucleotide regions with other functionalities.

[0140] The length of the tracrRNA sequence is not particularly limited and can be set according to the protein of this disclosure, for example. Examples of tracrRNA sequence lengths include 10-200 nucleotides, 20-190 nucleotides, 30-180 nucleotides, 40-180 nucleotides, 50-180 nucleotides, 60-180 nucleotides, 70-180 nucleotides, 70-170 nucleotides, 70-160 nucleotides, 70-150 nucleotides, 70-140 nucleotides, 70-130 nucleotides, 70-120 nucleotides, 70-110 nucleotides, 70-100 nucleotides, 70-90 nucleotides, and so on.

[0141] The length of the crRNA sequence is not particularly limited and can be set according to the protein of this disclosure. Examples of the tracrRNA sequence length include 25-80 nucleotides, 30-80 nucleotides, 35-80 nucleotides, 40-80 nucleotides, 49-80 nucleotides, 40-70 nucleotides, or 40-60 nucleotides.

[0142] The crRNA sequence can be configured to target a region adjacent to the PAM sequence of the target DNA.

[0143] An example of the guide RNA sequence of this disclosure is the following sequence. In the sequence of Sequence ID No. 33 below, the underlined N (where N is A, T, G, or U) at positions 36-56 is a hybridizable sequence (spacer sequence) with the target DNA sequence in the target DNA. In the sequence of Sequence ID No. 33 below, N is 21 nucleotides long, but the length of N is not limited to this, and should be within the range that allows the protein of this disclosure to target the target site of the target DNA. The length of N may be, for example, 14-30 nucleotides, 16-28 nucleotides, 16-25 nucleotides, 20-25 nucleotides, 20-22 nucleotides, or 21-23 nucleotides. Guide RNA (SEQ ID NO: 33) 5'-GUCAAAAGACCUUUGGAAUUUCUACUCUUGUAGAU NNNNNNNNNNNNNNNNNNNNN -3'

[0144] As mentioned above, the sequence of the guide RNA is not particularly limited as long as it can hybridize to the target DNA sequence.

[0145] The polynucleotides on the vector of the disclosed herein and the polynucleotides encoding guide RNA for a target site on target DNA are configured to be transcribed and / or translated in a host such as a cell to form a complex. The protein of the disclosed herein and the guide RNA are, for example, bound together to form a complex. The complex can, for example, target the protein of the disclosed herein to a target site on target DNA in a sequence-dependent manner of the guide RNA, and modify the target DNA in an activity-dependent manner of the protein of the disclosed herein. The modifications include, for example, double-strand breaks, single-strand breaks, conversion, substitution, deletion, or insertion of one or more nucleotides in the target DNA. The modifications may also be, for example, modifications of the base, sugar, or phosphate portion of the target DNA. The complex of the guide RNA and the protein of the disclosed herein causes, for example, one or more nucleotides to be converted, substituted, deleted, or inserted at the targeted site. The one or more nucleotides may be, for example, 1 to 10, 1 to 5, etc.

[0146] The target DNA may include, for example, chromosomal DNA, genomic DNA, mitochondrial DNA, viral DNA, bindable elements, exogenous DNA, plasmid DNA, etc. If the target DNA codes for a gene, it may also be called, for example, a target gene. The target DNA may be single-stranded (ssDNA) or double-stranded (dsDNA).

[0147] The target DNA sequence into which the guide RNA can hybridize is not particularly limited and can be set, for example, according to the PAM (proto-spacer adjacent motif) sequence of the protein of this disclosure. The length of the target DNA sequence is not particularly limited and can be, for example, a length into which the guide RNA can specifically hybridize, such as 14-30 nucleotides or 16-25 nucleotides.

[0148] The vectors of this disclosure may, for example, include a promoter sequence on the vector. The promoter can be appropriately configured, for example, depending on the type of cell expressing the polynucleotide encoding the protein of this disclosure and / or the guide RNA. Examples of promoters include the T3 promoter, T7 promoter, sp6 promoter, EF1α promoter, SRα promoter, SV40 (Simian virus) promoter, LTR promoter, CMV (Cytomegalovirus) promoter, RSV (Respiratory syncytial virus) promoter, HSV-tk promoter, Cauliflower mosaic virus (CaMV) 35S promoter, actin promoter, heat shock promoter, and REF (Rubber Elongation Examples of promoters include Factor promoter, polyhedrin promoter, p10 promoter, trp promoter, lac promoter, recA promoter, λPL promoter, lpp promoter, tac promoter, GAL1 promoter, GAL10 promoter, PH05 promoter, PGK promoter, GAP promoter, ADH promoter, SpO1 promoter, SpO2 promoter, penP promoter, gyrA-ldh promoter, pgm promoter, rrn4 promoter, P23 promoter, araBAD promoter, cat promoter, cspA promoter, EM7 promoter, J23119 promoter, T5 promoter, tac promoter, U6 promoter, TDH3 promoter, TEF1 promoter, etc.

[0149] The vectors of this disclosure may, for example, include a polynucleotide encoding an enzyme-active protein or its enzyme-active domain. The enzyme activity of the enzyme-active protein may include, for example, nucleic acid modification activities such as methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity; transcription activation activity, transcription repression activity, transcription deactivation factor activity, DNA integration activity, nucleic acid binding activity, etc. The enzyme-active protein may also be, for example, a peptide fragment thereof, as long as it has catalytic activity.

[0150] The vectors of this disclosure may include, for example, a multicloning site, an enhancer, a splicing signal, a poly-A addition signal, a drug resistance gene, a nutritional complement gene, an origin of replication, etc.

[0151] The vectors of this disclosure can be prepared, for example, by inserting a polynucleotide containing the coding sequence of the protein of this disclosure into a skeletal vector (hereinafter also referred to as the "basic vector"). The type of expression vector is not particularly limited and can be appropriately determined, for example, depending on the type of host. Specifically, when synthesizing the vector by genetic engineering, the synthesis of the vector first involves, for example, designing and synthesizing a polynucleotide containing the coding sequence of the protein of this disclosure. This design and synthesis can be carried out by PCR, for example, using a vector containing a polynucleotide encoding a polynucleotide containing the coding sequence of the protein of this disclosure as a template, and using primers designed to synthesize a desired nucleic acid region. Then, a recombinant vector for protein expression (expression vector) can be obtained by ligating the obtained polynucleotide into a suitable vector, and a transformant can be obtained by introducing this recombinant vector into a host so that the target gene can be expressed (Sambrook J. et al., Molecular Cloning, A Laboratory Manual (4th edition) (Cold Spring Harbor Laboratory Press (2012)).

[0152] Examples of such vectors include phage vectors, plasmid vectors, viral vectors, retroviral vectors, chromosomal vectors, episomal vectors, and virus-derived vectors.

[0153] When transforming a host using the heat shock method as the introduction method, the vector may be, for example, a binary vector. Examples of expression vectors include pETDuet-1, pQE-80L, and pUCP26Km. When transforming bacteria such as E. coli, examples of expression vectors include pETDuet-1 vector (Novagen), pQE-80L (QIAGEN), pBR322, pB325, pAT153, and pUC8.

[0154] <Vector System> In another embodiment, the Disclosure provides a vector system. The vector system of the Disclosure comprises a first vector comprising a polynucleotide of the Disclosure and a second vector comprising a polynucleotide encoding a guide RNA to a target site. The vector system of the Disclosure can be described by reference to the descriptions of proteins, polynucleotides, and vectors of the Disclosure.

[0155] In the vector system of this disclosure, the first vector and the second vector may, for example, be arranged on the same vector or on different vectors.

[0156] <Composition> In another embodiment, the Disclosure provides compositions. A composition of the Disclosure comprises a protein, a polynucleotide, and / or a vector of the Disclosure, and a guide RNA or polynucleotide encoding a guide RNA to a target site on target DNA. The descriptions of the proteins, polynucleotides, vectors, and vector systems of the Disclosure may be incorporated herein.

[0157] <Kit> In another embodiment, the Disclosure provides a kit capable of modifying target DNA. The kit of the Disclosure comprises a protein of the Disclosure, a polynucleotide of the Disclosure, and / or a vector of the Disclosure, and a guide RNA or polynucleotide encoding it to a target site. The kit of the Disclosure can be described by reference to the descriptions of the protein, polynucleotide, vector, vector system, and composition of the Disclosure.

[0158] The kits of this disclosure may include, for example, organisms. Examples of such organisms include prokaryotes and eukaryotic cells. Examples of such prokaryotes include lactic acid bacteria and natto bacteria. Bacillus subtilis var. natto ), E. coli ( Escherichia coli ), cyanobacteria ( Cyanobacteria Examples include the following: The eukaryotes include, for example, microorganisms, plants, animals, insects, etc. The animals include, for example, humans, monkeys, dogs, cats, rabbits, pigs, cows, mice, rats, etc. The organisms may be one type or multiple types.

[0159] The kit of this disclosure may include, for example, cells. Examples of such cells include prokaryotic cells and eukaryotic cells. Examples of such prokaryotic cells include lactic acid bacteria and natto bacteria. Bacillus subtilis var. natto ), E. coli ( Escherichia coli ), cyanobacteria ( Cyanobacteria Examples of cells include ) and others. Examples of eukaryotic cells include microbial cells, plant cells, animal cells, etc. Examples of microbial cells include cells derived from microorganisms such as yeast, fungi, and protozoa. Examples of plant cells include cells derived from plants such as seed plants, ferns, mosses, and algae. Examples of animal cells include cells derived from animals such as vertebrates such as mammals, birds, reptiles, amphibians, and fish; arthropods such as insects; mollusks; etc. The cells are, in vivo Cells (in vivo cells), in vitroExamples include cells (cultured cells), primary cultured cells, etc. The cells may be of one type or multiple types.

[0160] The kit described herein may include, for example, instructions, a manual, etc.

[0161] The kit of this disclosure may include, for example, a culture medium. Examples of such culture media include YM medium, YPD medium, PD medium, DOB medium, SD medium, LB medium, NB medium, SCD medium, MRS medium, BHI medium, etc. The culture medium may include, for example, a carbon source such as glucose, dextrin, soluble starch, sucrose; a nitrogen source such as ammonium salts, nitrates, corn slush liquor, peptone, casein, meat extract, soybean meal, potato extract; inorganic substances such as calcium chloride, sodium dihydrogen phosphate, magnesium chloride; vitamins; growth factors; etc. If the kit of this disclosure includes such additives, the additives may be additives to be added to the culture medium before culturing, or additives to be added during culturing. The additions may be continuous or intermittent.

[0162] The kits of the present disclosure may further include, for example, containers for storing the proteins of the present disclosure, the polynucleotides of the present disclosure, the vectors of the present disclosure, the vector systems of the present disclosure, and / or compositions of the present disclosure, and culture media.

[0163] The kits of the present disclosure may, for example, contain each component separately, or some or all of them may be contained mixed or unmixed. If, in the kits of the present disclosure, all components are contained in a single container, either mixed or unmixed, the kits of the present disclosure may also be called, for example, culture media.

[0164] The kits of this disclosure can also be suitably used, for example, as test kits or research kits for use in genome editing of organisms and / or cells.

[0165] <cell> In another embodiment, the Disclosure provides cells. The cells of the Disclosure include the proteins of the Disclosure, the polynucleotides of the Disclosure, the vectors of the Disclosure, and / or the vector systems of the Disclosure. The cells of the Disclosure can be made by reference to the descriptions of the proteins, polynucleotides, vectors, vector systems, and kits of the Disclosure.

[0166] If the cells of the present disclosure include the polynucleotides of the present disclosure, the vectors of the present disclosure, and / or the vector systems of the present disclosure, then the cells of the present disclosure can be said to express the polynucleotides of the present disclosure.

[0167] <Methods for modifying target DNA> In other embodiments, the Disclosure provides a method for modifying target DNA. The method for modifying target DNA according to the Disclosure comprises a modification step of modifying the target DNA by contacting the target DNA with the protein of the Disclosure and a guide RNA for a target site in the target DNA within a cell. The modification method according to the Disclosure can be described by reference to the descriptions of the protein, polynucleotide, vector, vector system, composition, and kit of the Disclosure.

[0168] In the modification step, the complex of the protein of the present disclosure and the guide RNA is targeted to the target DNA, and the target DNA is modified by the complex at the target site.

[0169] The modification method of this disclosure includes, for example, a transfection step of introducing the protein of this disclosure or a polynucleotide encoding it and the guide RNA or a polynucleotide encoding it into the cells. The transfection can be carried out by known methods capable of introducing (transfecting) exogenous proteins or exogenous DNA into host cells, specifically, methods such as transfection using a gene gun such as a particle gun, microinjection, calcium phosphate, polyethylene glycol, lipofection using liposomes, electroporation, ultrasonic nucleic acid transfection, DEAE-dextran, direct injection using microglass tubes, hydrodynamic, cationic liposome, lysozyme, competent, PEG, methods using transfection aids, methods via Agrobacterium, protoplasts, etc. Examples of liposomes include lipofectamine and cationic liposomes, and examples of transfection aids include atelocollagen, nanoparticles, and polymers.

[0170] The introduction step may include a cell culture step after the introduction. The culture can be carried out using known methods or modified methods and conditions, depending on the type of cell. The culture medium used in the culture can be described in the kit.

[0171] In the culture step, the culture pH is not particularly limited as long as it is the optimal pH for the proliferation of the cells.

[0172] In the culture step described above, the culture temperature is not particularly limited as long as it is the optimal temperature for the proliferation of the cells.

[0173] In the culture step, the culture time is not particularly limited as long as it is the time during which the complex is expressed. The culture time is, for example, the time during which the protein and guide RNA of this disclosure are expressed and the target DNA can be modified.

[0174] The culture may be carried out, for example, by static culture or by shaking culture.

[0175] <Manufacturing method> In another embodiment, the Disclosure provides a method for producing cells or non-human animals in which target DNA has been modified. The method for producing cells or non-human animals in which target DNA has been modified according to the Disclosure includes an introduction step of introducing the proteins, polynucleotides, and / or vectors of the Disclosure and a guide RNA to a target site of target DNA into cells or non-human animals. The methods for producing the methods of the Disclosure may be described by reference to the descriptions of the proteins, polynucleotides, vectors, vector systems, compositions, kits, and modification methods of the Disclosure.

[0176] <Cells or non-human animals> In another embodiment, the Disclosure provides cells or non-human animals. The cells or non-human animals of the Disclosure are produced by the production methods of the Disclosure. The cells or non-human animals of the Disclosure can be produced by referring to the descriptions of proteins, polynucleotides, vectors, vector systems, compositions, kits, cells, modification methods, and production methods of the Disclosure. [Examples]

[0177] The present disclosure will be described in detail below using examples, but the present disclosure is not limited to the embodiments described in the examples.

[0178] [Example 1] (1) Construction of a vector containing polynucleotides encoding each protein Vectors containing polynucleotides encoding the proteins of the disclosed herein (SEQ ID NOs: 1-15) were constructed. Specifically, a coding sequence was inserted downstream of the CBh promoter of an expression vector into which a CAG-TurboRFP expression cassette was inserted. This sequence included SV40 NLS (SEQ ID NOs: 34) at the N-terminus of each of the disclosed proteins (SEQ ID NOs: 1-15), nucleoplasmin NLS (SEQ ID NOs: 35), a GS linker, and a 3×HA tag (SEQ ID NOs: 36) at the C-terminus of each of the disclosed proteins. Additionally, a polynucleotide encoding a guide RNA (SEQ ID NOs: 37) was inserted downstream of the U6 promoter to obtain the vectors of the disclosed herein. Furthermore, as other expression vectors, expression vectors were also constructed in which polynucleotides encoding the amino acid sequences of SEQ ID NOs: 16-19 (ST9: SEQ ID NOs: 16, ST9_2: SEQ ID NOs: 17, ST9_3LA: SEQ ID NOs: 18, ST9_S: SEQ ID NOs: 19) were inserted instead of the proteins of the disclosed herein.

[0179] SV40 NLS (Sequence ID 34) PKKKRKV

[0180] nucleoplasmin NLS (SEQ ID NO: 35) KRPAATKKAGQAKKKK

[0181] 3 x HA tags (Sequence number 36) YPYDVPDYAYPYDVPDYAYPYDVPDYA

[0182] Guide RNA (SEQ ID NO: 37) 5'-GUCAAAAGACCUUUGGAAUUUCUACUCUUGUAGAUUCUGUCCCCUCCACCCCACAG-3'

[0183] ST9 (Sequence ID 16)

[0184] ST9_2 (Sequence ID 17)

[0185] ST9_3LA (Sequence ID 18)

[0186] ST9_S (Sequence ID 19)

[0187] (2) Construction of vectors for EGxxFP assay detection A vector was prepared containing the EGFP amino-terminal region, PAM sequences (SEQ ID NOs. 24-32), the nucleotide sequence to be cleaved, and the EGFP carboxyl-terminal region, inserted downstream of the CAG promoter in that order. This EGxxFP assay detection vector is designed so that when cleaved by genome editing, the EGFP sequence is repaired and green fluorescence is observed.

[0188] (3) EGxxFP assay The cleavage activity of the proteins disclosed herein against target sites was investigated by the EGxxFP assay. Specifically, HEK293T cells were seeded in 24-well plates and cultured for 1.5 to 2 days at 37°C in a CO2 incubator. The culture media used were DMEM (High Glucose) (Sigma-Aldrich), 10% FBS (Gibco), and a 1× penicillin-streptomycin mixture (Nacalitex). After the culture, 75 μl of Solution A (600 μl Opti-MEM (Gibco), 3 μl of 2 ng / nl EGxxFP assay detection vector) was dispensed into each tube, and 0.75 μl of each expression vector constructed in Example 1(1) at 0.4 ng / nl was mixed in. After the above mixing, 75 μl of solution B (600 μl of Opti-MEM (Gibco) and 60 μl of PEIMax (Polysciences)) was added to prepare the mixture. After preparation, the mixture was allowed to stand at room temperature (hereinafter referred to as approximately 24°C) for 20 minutes. Then, the culture medium in the 24-well plate was replaced with fresh medium, and 50 μl of the mixture was added to each well. After the addition, the 24-well plate was gently shaken. After shaking, the cells were cultured for 1.5 to 2 days and harvested by trypsin treatment. After harvesting, a diluted solution of Passive Lysis 5X Buffer (Promega) was added, and the cells were vigorously mixed to lyse them. After lysis, the cells were centrifuged at 12000 rpm for 5 minutes, and the supernatant was collected. 50 μl of the supernatant from each sample was added to two wells of a 96-well plate, and the fluorescence intensities of EGFP and RFP contained in the supernatant were detected using a plate reader (CYTATION® 3 imaging reader (BioTek)). In the detection, the EGFP value was standardized by the RFP value. A sample using an expression vector that did not contain the polynucleotide encoding the guide RNA of this disclosure was used as a negative control, and the negative control value was subtracted from the value of each sample. These results are shown in Figures 1-4.

[0189] Figures 1-4 are graphs showing the results of the EGxxFP assay. In Figures 1-4, the vertical axis represents the EGFP value standardized by the RFP value, and the horizontal axis represents the sample type. As shown in Figures 1-4, the proteins disclosed herein were found to have higher nuclease activity compared to the control.

[0190] [Example 2] (1) Construction of a vector containing polynucleotides encoding modified proteins with different amino acids having cleavage active sites. A vector containing a polynucleotide encoding a protein with a modified cleavage active site was constructed. Specifically, a protein (CR-FD2, SEQ ID NO: 38) in which the 969th amino acid of ST9, glutamic acid (E), is replaced with glycine (G), and the amino acids 1040-1043, valine, proline, alanine, and alanine, are replaced with alanine, aspartic acid, glutamic acid, and arginine, respectively, was obtained by inserting a coding sequence with SV40 NLS (SEQ ID NO: 34) attached to the N-terminus of the protein (CR-FD2, SEQ ID NO: 38) and nucleoplasmin NLS (SEQ ID NO: 35), a GS linker, and a 3×HA tag (SEQ ID NO: 36) attached to the C-terminus of the protein of this disclosure downstream of the CBh promoter of the expression vector. Note that CR-FD2 has mutations in the cleavage active site and adjacent sites (Figure 5). Therefore, CR-FD2 has a different cleavage active site from the protein without mutations (ST9).

[0191] CR-FD2 (Sequence ID 38) MGEKTNYFNNFIGISPLTKTLRNALIPTTEITQKHIIEYGIIKDDELRKENRQILKTIMDDYYRSFLTEKSAIHDIDWKPLFLEMENELRHGNNKVNLEKEQKAKRKAINKYFSEDERYKKMFSAKLLSDILPEFVIHNDEYSAEEKEEKQVIKLFSRFATSLKEYFRNRSNTFSADNISTSACHRIVNDNASIFLENVMAYRKIISSLPADELDKLESIQDKLKIKSISEVYTYDNYGKYITQEGIDLYNDICGKVNSFMNLYCQQNKENKNVYKMRKLHKQILCIAKTSYEVGEGYTSDEEVLEVFRNTLNKSEIFSSIKKLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEYDDIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLKEIIIQKVDEIYKVYGSSEKLFDADFVLEKSLKNDAVVAIMKDLLDSVKSFENYIKAFFGEGK ETNRDESFYGDFVLAYDILLKVDHIYDAIRNYVTQKPYSTEKFKLNFNSPTLARGWSKSKEYSNNAIILSKDGILYLGIFNVKNKPDKKIIEGHIRENDGDYKKMVYNLLPGANKMLPKVLISSKSGVETYKPSDYILEGYSNKHLKSSDTFDINYCHDLIDYYKDCIDIHPEWKNFDFNFSDTADFEDISGFYREAVEAQGYKIDWTYISGENIEELQEKGQLFLFKIYNKDFSTKSTGTD NLHTMYLKNLFSEENLKDVVLKLNGEAEIFYRKSSIKNPIIHKKGSMLVNTYIEEEEEHDGKIVEVRKIVPEEIYKELYNHFNNNGKAELSEEAKRLESLVCHEAAKDIVKDCRYTYDKFFIHLPMTINFKASGSSSLNDMMLQYISGQNDMHIIGIDRGERNLIYVSVIDAHGNIVKQKSFNIVAGYDYQEKLKMQEGARQNARKEWKIKEIKEGYLSLVIHEIAQMVIEYNAIIAM GDLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY ADER YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0192] (2) EGxxFP assay The cleavage activity of modified proteins with mutations at the cleavage site and adjacent sites was investigated using the EGxxFP assay. Specifically, the assay was performed in the same manner as the EGxxFP assay in Example 1(3), except that the vector constructed in Example 2(1) was used instead of the vector constructed in Example 1(1). These results are shown in Figure 6.

[0193] Figure 6 is a graph showing the results of the EGxxFP assay. In Figure 6, the vertical axis represents the EGFP value standardized by the RFP value, and the horizontal axis represents the sample type. As shown in Figure 6, the protein of this disclosure, with its improved cleavage active site, was found to possess nuclease activity.

[0194] [Example 3] (1) Construction of a vector containing a polynucleotide encoding a modified protein with alanine substitution. A vector containing a polynucleotide encoding a modified protein in which the amino acid responsible for cleavage activity is substituted with alanine was constructed. Specifically, a protein (CR-FD3, SEQ ID NO: 39) in which the 969th amino acid of ST9, glutamic acid, is substituted with alanine, and the 1040-1043rd amino acids, valine, proline, alanine, and alanine, are substituted with alanine, aspartic acid, glutamic acid, and arginine, respectively, was obtained by inserting a coding sequence downstream of the CBh promoter of the expression vector. The CR-FD3 has an alanine mutation at its cleavage activity site (Figure 7). Therefore, the cleavage activity site of CR-FD3 is different from that of the unmutated ST9. Furthermore, CR-FD3 can be described as a protein in which mutations have been introduced into the amino acids constituting the active site in CR-FD2.

[0195] CR-FD3 (Sequence ID 39) MGEKTNYFNNFIGISPLTKTLRNALIPTTEITQKHIIEYGIIKDDELRKENRQILKTIMDDYYRSFLTEKSAIHDIDWKPLFLEMENELRHGNNKVNLEKEQKAKRKAINKYFSEDERYKKMFSAKLLSDILPEFVIHNDEYSAEEKEEKQVIKLFSRFATSLKEYFRNRSNTFSADNISTSACHRIVNDNASIFLENVMAYRKIISSLPADELDKLESIQDKLKIKSISEVYTYDNYGKYITQEGIDLYNDICGKVNSFMNLYCQQNKENKNVYKMRKLHKQILCIAKTSYEVGEGYTSDEEVLEVFRNTLNKSEIFSSIKKLEKLFKNFDEYSSAGIFVKNGPAISTISKDIFGEWNVIRDKWNAEYDDIHLKKKAVVTEKYEDDRRKSFKKIGSFSLEQLQEYADADLSVVEKLKEIIIQKVDEIYKVYGSSEKLFDADFVLEKSLKNDAVVAIMKDLLDSVKSFENYIKAFFGEGK ETNRDESFYGDFVLAYDILLKVDHIYDAIRNYVTQKPYSTEKFKLNFNSPTLARGWSKSKEYSNNAIILSKDGILYLGIFNVKNKPDKKIIEGHIRENDGDYKKMVYNLLPGANKMLPKVLISSKSGVETYKPSDYILEGYSNKHLKSSDTFDINYCHDLIDYYKDCIDIHPEWKNFDFNFSDTADFEDISGFYREAVEAQGYKIDWTYISGENIEELQEKGQLFLFKIYNKDFSTKSTGTD NLHTMYLKNLFSEENLKDVVLKLNGEAEIFYRKSSIKNPIIHKKGSMLVNTYIEEEEEHDGKIVEVRKIVPEEIYKELYNHFNNNGKAELSEEAKRLESLVCHEAAKDIVKDCRYTYDKFFIHLPMTINFKASGSSSLNDMMLQYISGQNDMHIIGIDRGERNLIYVSVIDAHGNIVKQKSFNIVAGYDYQEKLKMQEGARQNARKEWKIKEIKEGYLSLVIHEIAQMVIEYNAIIAM ADLNYGFKKGRFKVERQVYQKFETMLINKLNYLVFKNRGVTEDGGLLRGYQLTYIPESLKNVGRQCGCIFY ADER YTSKIDPTTGFADIFRFKNLTVEEKRDFIRKFDYIKYDLEKEMFVFAFDFKNFVTQNVEMSKNDWCVYTNGIRVKRRYANGRFTNETDEIDINVLMKKTFERTDIDWNDGHNL IDEIIDYDLESQIVDIFKMAVQMRNSRSEAEDRDYDRLVSPVLNESGDFFDSSKVGENLPKDADANGAYCIALKGLYKVRQIQENWNEEEKFSRAKLRISNKDWFDFVQNKRYL

[0196] (2) EGxxFP assay The cleavage activity of modified proteins containing alanine mutations at the cleavage site was investigated using the EGxxFP assay. Specifically, the assay was performed in the same manner as the EGxxFP assay in Example 1(3), except that the vector constructed in Example 3(1) was used instead of the vector constructed in Example 1(1). These results are shown in Figure 8.

[0197] Figure 8 is a graph showing the results of the EGxxFP assay. In Figure 8, the vertical axis shows the EGFP value standardized by the RFP value, and the horizontal axis shows the sample type. As shown in Figure 8, proteins in which the 969th amino acid, glutamic acid, is substituted with alanine, and the 1040-1043rd amino acids, valine, proline, alanine, and alanine, are substituted with alanine, aspartic acid, glutamic acid, and arginine, respectively, were found to lack nuclease activity.

[0198] [Example 4] The cleavage activity of the protein of this disclosure at the target site when the PAM sequence is TTTC (SEQ ID NO: 40) was investigated by the EGxxFP assay. Specifically, the EGxxFP assay was performed in the same manner as in Example 1(3) using each expression vector constructed in Example 1(1) and the EGxxFP assay detection vector of Example 1(2) with the PAM sequence being TTTC. These results are shown in Figure 9.

[0199] Figure 9 is a graph showing the results of the EGxxFP assay. In Figure 9, the vertical axis represents the EGFP value standardized by the RFP value, and the horizontal axis represents the sample type. As shown in Figure 9, the protein disclosed herein was found to recognize the PAM sequence TTTC and exhibit higher nuclease activity compared to the control.

[0200] Although the present disclosure has been described above with reference to embodiments and examples, the present disclosure is not limited to the above embodiments and examples. Various modifications to the structure and details of the present disclosure are possible, as can be understood by those skilled in the art within the scope of the present disclosure.

[0201] The patents, patent applications, and documents cited herein are incorporated herein by reference in the same manner as their contents are specifically described herein.

[0202] <Note> Some or all of the above embodiments and examples may be described as follows, but are not limited to the following. <Protein> (Note 1) The following proteins: (a1), (a2), or (a3): (a1) A protein consisting of any of the amino acid sequences of SEQ ID NOs: 1 to 19; (a2) A protein having nuclease activity, consisting of an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, or added in any of the amino acid sequences of Sequence ID No. 1 to 19; (a3) A protein having nuclease activity, consisting of an amino acid sequence that has 80% or more identity with any of the amino acid sequences of Sequence ID No. 1 to 19. (Note 2) The protein described in Appendix 1 recognizes a protospacer adjacent motif (PAM) sequence having a base sequence selected from the group consisting of SEQ ID NOs. 20 to 32. (Note 3) The protein in (a1) is one of the proteins listed in (b1) to (b6) or (b7) below, as described in Appendix 1 or 2: (b1) Proteins in which amino acid E at position 1136 in the amino acid sequence of SEQ ID NO: 1 or 6 is substituted with G; (b2) Proteins in which amino acid E at position 1129 is substituted with G in the amino acid sequence of sequence numbers 2-5, 14, or 15; (b3) Proteins in which the amino acid sequence of sequence number 7 or 9 has a substitution of amino acid E at position 1127 with G; (b4) Proteins in which the amino acid sequence of sequence number 8 or 10 has the amino acid sequence at position 1120, E, replaced by G; (b5) A protein in which the amino acid sequence of sequence number 11 of sequence number 11 has the amino acid sequence E at position 1128 replaced with G; (b6) Proteins in which the amino acid E at position 969 of sequence number 16 or 18 is substituted with G; (b7) A protein in which the amino acid E at position 960 of sequence number 17 is replaced with G. (Note 4) The protein in (a1) is one of the proteins listed in (c1) to (c6) or (c7) below, as described in any of the notes in Appendix 1 to 3: (c1) Proteins in which amino acids VPAA at positions 1207-1210 in the amino acid sequence of SEQ ID NO: 1 or 6 are substituted with ADER; (c2) Proteins in which the amino acid VPAA at positions 1200-1203 is substituted with ADER in the amino acid sequences of SEQ ID NOs. 2-5, 14, or 15; (c3) Proteins in which the amino acid sequence VPAA at positions 1198-1201 in sequence number 7 or 9 is replaced with ADER; (c4) Proteins in which the amino acid sequence VPAA at positions 1191-1194 in sequence number 8 or 10 is replaced with ADER; (c5) A protein in which the amino acid sequence VPAA at positions 1199-1202 in sequence number 11 of sequence number 11 is replaced with ADER; (c6) Proteins in which the amino acid sequence VPAA at positions 1040-1043 in sequence number 16 or 18 is replaced with ADER; (c7) A protein in which the amino acid sequence VPAA at positions 1031-1034 of sequence number 17 is replaced with ADER. (Note 5) An RNA-dependent DNA nuclease, the protein described in any of the appendices 1 to 4. (Note 6) Furthermore, a protein having an enzyme activity domain, as described in any of the appendices 1 to 5. (Note 7) The enzyme-active domain is located at the 5' end and / or 3' end of the protein described in any of Appendix 1 to 6, the protein described in any of Appendix 1 to 6. (Note 8) The enzyme activity is selected from the group consisting of nucleic acid base conversion activity, methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity, as described in any of the appendices 1 to 7. <Polynucleotides> (Note 9) A polynucleotide containing the coding sequence for a protein described in any of the appendices 1 to 8. <Vector> (Note 10) A vector containing the polynucleotides described in Appendix 9. (Note 11) Furthermore, the vector described in Appendix 10 includes a polynucleotide encoding a guide RNA for the target site. <Vector System> (Note 12) The first vector containing the polynucleotide described in Appendix 9, A vector system comprising a second vector containing a polynucleotide encoding a guide RNA for a target site. <Composition> (Note 13) A protein as described in any of Appendix 1 to 8, a polynucleotide as described in Appendix 9, and / or a vector as described in Appendix 10 or 11, A composition comprising a guide RNA for a target site or a polynucleotide encoding it. (Note 14) The guide RNA is the composition described in Appendix 13, which is capable of complexing with the protein. (Note 15) The composition according to Appendix 13 or 14, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA. (Note 16) The guide RNA is a composition according to any one of appendices 13 to 15, comprising crRNA. <Kit> (Note 17) A protein as described in any of Appendix 1 to 8, a polynucleotide as described in Appendix 9, and / or a vector as described in Appendix 10 or 11, A kit comprising a guide RNA for a target site or a polynucleotide encoding it. <cell> (Note 18) A cell comprising a protein as described in any of the appendices 1 to 8, a polynucleotide as described in appendice 9, a vector as described in appendice 10 or 11, and / or a vector system as described in appendice 12. <Methods for modifying target DNA> (Note 19) A method for modifying target DNA, comprising a modification step of modifying the target DNA by bringing the target DNA into contact with a protein described in any of the appendices 1 to 8 and a guide RNA for a target site within the target DNA within a cell. (Note 20) The modification method according to Appendix 19, comprising an introduction step of introducing the protein or polynucleotide encoding the same described in any of Appendix 1 to 8 and the guide RNA or polynucleotide encoding the same into the cell. (Note 21) The modification method according to Appendix 19 or 20, wherein in the modification step, the complex of the protein and the guide RNA is targeted to the target DNA, and the target DNA is modified by the complex at the target site. (Note 22) The modification method according to any one of appendices 19 to 21, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA. (Note 23) The aforementioned guide RNA includes crRNA, and the modification method is as described in any of appendices 19 to 22. (Note 24) The modification method according to any one of the appendices 19 to 23, wherein the modification is a double-strand break, a single-strand break, a conversion of one or more nucleotides, a deletion and / or insertion of the target DNA. (Note 25) The modification method according to any one of the appendices 19 to 24, wherein the cells are animal cells or plant cells. (Note 26) The modification method according to any one of the appendices 19 to 25, wherein the cells are eukaryotic cells or prokaryotic cells. (Note 27) The modification method described in any of appendices 19 to 26, wherein the target DNA is chromosomal DNA. <Manufacturing method> (Note 28) A method for producing cells or non-human animals in which target DNA has been modified, A method for manufacturing a cell or a non-human animal comprising an introduction step of introducing a protein described in any of appendices 1 to 8, a polynucleotide described in appendice 9, and / or a vector described in appendice 10 or 11, and a guide RNA for the target site. <Cells or non-human animals> (Note 29) Cells or non-human animals produced by the manufacturing method described in Appendix 28. [Industrial applicability]

[0203] As described above, the proteins disclosed herein can provide novel proteins that can be used for modifying target DNA. Therefore, the present invention is extremely useful in fields such as medicine, food, and agriculture.

Claims

1. The following proteins (a1), (a2), or (a3): (a1) A protein consisting of any of the amino acid sequences of SEQ ID NOs: 1-15, 18, and 19; (a2) A protein having nuclease activity, consisting of an amino acid sequence in which 1 to 142 amino acids are deleted, inserted, substituted, or added in any of the amino acid sequences of SEQ ID NOs: 1 to 15, 18, and 19; (a3) A protein having nuclease activity, consisting of an amino acid sequence that has 90% or more identity with any of the amino acid sequences of SEQ ID NOs: 1-15, 18, and 19.

2. The protein according to claim 1, which recognizes a protospacer adjacent motif (PAM) sequence having a base sequence selected from the group consisting of SEQ ID NOs. 20 to 32.

3. The protein in (a1) is the protein in (b1) to (b5) or (b6) described below, according to claim 1 or 2: (b1) Proteins in which the amino acid E at position 1136 is replaced with G in the amino acid sequence of SEQ ID NO: 1 or 6; (b2) Proteins in which amino acid E at position 1129 is substituted with G in the amino acid sequence of SEQ ID NOs: 2-5, 14, or 15; (b3) Proteins in which the amino acid sequence of SEQ ID NO: 7 or 9 has the amino acid sequence E at position 1127 replaced with G; (b4) Proteins in which the amino acid sequence of SEQ ID NO: 8 or 10 has the 1120th amino acid sequence E replaced by G; (b5) A protein in which the amino acid sequence of Sequence ID No. 11 has the amino acid sequence E at position 1128 replaced with G; (b6) A protein in which the amino acid E at position 969 of sequence number 18 is replaced with G.

4. The protein in (a1) is the protein in (c1) to (c5) or (c6) described below, according to claim 1 or 2: (c1) Proteins in which amino acids VPAA at positions 1207-1210 in the amino acid sequence of SEQ ID NO: 1 or 6 are replaced with ADER; (c2) Proteins in which the amino acid VPAA at positions 1200-1203 is substituted with ADER in the amino acid sequences of SEQ ID NOs. 2-5, 14, or 15; (c3) Proteins in which the amino acid sequence VPAA at positions 1198-1201 in sequence number 7 or 9 is replaced with ADER; (c4) Proteins in which the amino acid sequence VPAA at positions 1191-1194 in sequence number 8 or 10 is replaced with ADER; (c5) A protein in which the amino acid sequence VPAA at positions 1199-1202 in SEQ ID NO: 11 is replaced with ADER; (c6) A protein in which the amino acid sequence VPAA at positions 1040-1043 in sequence number 18 is replaced with ADER.

5. The protein according to claim 1 or 2, which is an RNA-dependent DNA nuclease.

6. Furthermore, the protein according to claim 1 or 2, having an enzyme-active domain.

7. The protein according to claim 1 or 2, wherein the enzyme-active domain is located at the 5' end and / or 3' end of the protein according to claim 1 or 2.

8. The protein according to claim 1 or 2, wherein the enzyme activity is selected from the group consisting of nucleic acid base conversion activity, methylase activity, demethylase activity, base modification activity, histone modification activity, RNA cleavage activity, and DNA cleavage activity.

9. A polynucleotide comprising the coding sequence for the protein described in claim 1.

10. A vector comprising the polynucleotide described in claim 9.

11. Furthermore, the vector according to claim 10 comprises a polynucleotide encoding a guide RNA for a target site.

12. A first vector comprising the polynucleotide described in claim 9, A vector system comprising a second vector containing a polynucleotide encoding a guide RNA for a target site.

13. A protein according to claim 1, a polynucleotide according to claim 9, and / or a vector according to claim 10, A composition comprising a guide RNA for a target site or a polynucleotide encoding it.

14. The composition according to claim 13, wherein the guide RNA is capable of complexing with the protein.

15. The composition according to claim 13, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA.

16. The composition according to claim 13, wherein the guide RNA comprises crRNA.

17. A protein according to claim 1, a polynucleotide according to claim 9, and / or a vector according to claim 10, A kit comprising a guide RNA for a target site or a polynucleotide encoding it.

18. A cell comprising the protein according to claim 1, the polynucleotide according to claim 9, the vector according to claim 10, and / or the vector system according to claim 12.

19. A method for modifying target DNA, comprising a modification step of modifying the target DNA by bringing the target DNA into contact with the protein described in claim 1 or 2 and a guide RNA for a target site within the target DNA within a cell.

20. The modification method according to claim 19, comprising an introduction step of introducing the protein or polynucleotide encoding the same according to claim 1 or 2 and the guide RNA or polynucleotide encoding the same into the cell.

21. The modification method according to claim 19, wherein in the modification step, the complex of the protein and the guide RNA is targeted to the target DNA, and the target DNA is modified by the complex at the target site.

22. The modification method according to claim 19, wherein the guide RNA comprises a polynucleotide capable of hybridizing to a target site in the target DNA.

23. The modification method according to claim 19, wherein the guide RNA includes crRNA.

24. The modification method according to claim 19, wherein the modification is a double-strand break, a single-strand break, a conversion of one or more nucleotides, a deletion and / or insertion of the target DNA.

25. The modification method according to claim 19, wherein the cells are animal cells or plant cells.

26. The modification method according to claim 19, wherein the cells are eukaryotic cells or prokaryotic cells.

27. The modification method according to claim 19, wherein the target DNA is chromosomal DNA.

28. A method for producing cells or non-human animals in which target DNA has been modified, A method for producing a protein according to claim 1, a polynucleotide according to claim 9, and / or a vector according to claim 10, and a guide RNA for the target site, comprising the steps of introducing these into a cell or a non-human animal.