Method for knocking-in desired nucleotide sequence, method for producing knock-in cell, and kit

A novel nuclease domain-based method with a site-specific nicking system using single-stranded DNA donor DNA addresses the inefficiencies of CRISPR-Cas and TALEN, enabling precise and efficient nucleotide sequence insertion.

WO2025206046A1PCT designated stage Publication Date: 2025-10-02SUMITOMO CHEM CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012229
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-03-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Conventional genome editing methods using the CRISPR-Cas system require guide RNA and can cause unintended insertion or deletion of nucleotide sequences, and existing technologies using proteins like TALEN are not efficient for knock-in applications.

Method used

A method utilizing a novel nuclease domain (ND1) and its deficient variant (dND1) to form a nicking enzyme, which uses a site-specific nicking system with single-stranded DNA as donor DNA, avoiding double-strand cleavage and enhancing knock-in efficiency.

Benefits of technology

The method allows for precise and efficient insertion of desired nucleotide sequences into target regions using only proteins, reducing unintended mutations and improving knock-in efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012229_02102025_PF_FP_ABST
    Figure JP2025012229_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a method for inserting a donor sequence into a target region on target DNA, the method including a step in which donor DNA and a site-specific nicking system are brought into contact with the target DNA, wherein: the donor DNA is single-stranded DNA that comprises, from the 5' end, a 5'-end homology arm sequence, the donor sequence, and a 3'-end homology arm sequence in this order; and the site-specific nicking system includes a first fusion protein that includes a first DNA binding domain and an ND1 domain and a second fusion protein that includes a second DNA binding domain and a dND1 domain, wherein the ND1 domain and the dND1 domain together form a dimer to introduce a nick in the target region or in the vicinity of the target region.
Need to check novelty before this filing date? Find Prior Art

Description

Method for knocking in a desired nucleotide sequence, method for producing knock-in cells, and kit

[0001] The present invention relates to a method for knocking in a desired nucleotide sequence and a method for producing knock-in cells, and more specifically to a method for inserting (knock-in) a desired nucleotide sequence into a target region on target DNA, a method for producing cells in which a desired nucleotide sequence has been inserted into a target region on target DNA (knock-in cells), and kits for use in these methods.

[0002] A technique for inserting or substituting a desired nucleotide sequence into a specific region on target DNA such as genomic DNA is called "knock-in." For example, a method is known in which double-stranded DNA is cleaved using the CRISPR-Cas system and a repair mechanism called homology-directed repair (HDR) is used at the cleavage site to insert or replace donor DNA containing the nucleotide sequence to be inserted at or near the cleavage site. For example, Y Wu et al., Cell Stem Cell 13, December 5, 2013, pp. 659-662 (Non-Patent Document 1) reports that double-stranded DNA (DNA double-strand break; DSB) was cleaved in mouse fertilized eggs using the CRISPR-Cas9 system and a mutation was introduced near the cleavage site by HDR.

[0003] In such knock-in technology using the CRISPR-Cas system, it is known that a single-strand DNA (ssDNA) having sequences (homology arms) at the 5' and 3' ends of the nucleotide sequence to be inserted that are homologous to the peripheral sequences of the target region to be inserted is used as the donor DNA, or, when the sequence to be inserted is long, a donor plasmid containing the nucleotide sequence of such a single-strand DNA is used. For example, see Li H. et al., bioRxiv, 21 Aug 2017, doi: https: / / doi.org / 10.1038 / ssDNA.2017.03.0026 org / 10.1101 / 178905 (Non-Patent Document 2) reports that by using lssDNA (long single strand DNA: lssDNA) as donor DNA with homology arms of about 500 bases or more at both ends, efficient knock-in can be achieved using the CRISPR-Cas9 system even in cultured cells. Furthermore, lssDNA is difficult to produce and its use has been limited, but for example, Inoue et al., Cells, 2021, 10, 1076. https: / / doi.org / 10.3390 / cells10051076 (Non-Patent Document 3) reports a simpler and cheaper method for preparing lssDNA.

[0004] However, the knock-in technique using double-stranded DNA cleavage as described above has been problematic in that it causes the insertion or deletion of unintended nucleotide sequences. As a technique that does not use double-stranded DNA cleavage, for example, a technique in which a donor plasmid is combined with a nickase-type Cas9 (nCas9, Cas9D10A) in which the double-stranded DNA cleavage activity of Cas9 is partially deleted, and the nucleotide sequence of the donor plasmid is inserted into the site where single-stranded DNA cleavage (nicking) is caused by nCas9 (combination of single nicks in the target gene and donor plasmid: SNGD), is described by Nakajima et al., Genome Research, 28, 2018, pp. 223-230 (Non-Patent Document 4).

[0005] However, all of these conventional technologies rely on the CRISPR-Cas system, and in order to use them to edit the genome, guide RNA must also be introduced into the cell. Therefore, there is a need for a safer technology that enables genome editing using only proteins.

[0006] A known technology that enables genome editing using only proteins is TALEN (Transcription activator-like effector nuclease), an artificial restriction enzyme developed as a second-generation genome editing technology in 2010. TALEN is a fusion protein in which the nuclease domain of the type IIS restriction enzyme FokI (FokI nuclease domain) serves as the nuclease domain and a TALE serves as the DNA-binding domain; the TALEs of a pair of TALENs bind to opposite strands of the target DNA, and the FokI nuclease domains form a dimer with each other, thereby exerting site-specific double-stranded DNA cleavage activity (nuclease activity). Regarding the FokI nuclease domain, there have been reports of linking two FokI nuclease domains and binding them to a zinc finger array (ZF) or TALE to induce double-stranded DNA cleavage (Minczuk et al., Nucleic Acids Research, 36(12), 2008, pp. 3926-3938 (Non-Patent Document 5)), and of inducing single-stranded DNA cleavage (nicks) using a construct in which one side of the FokI nuclease domain has been mutated (Yan Luo et al., Scientific Reports, 6: 20657, 2016, doi: 10.1038 / srep20657 (Non-Patent Document 6)).

[0007] In addition, the present inventors have developed a novel nuclease domain 1 (ND1 domain) that is different from the conventional FokI nuclease domain, and have succeeded in editing target sites on target DNA using an artificial nucleic acid cleaving enzyme containing this ND1 domain and a DNA binding domain such as a ZF or TALE (WO 2020 / 045281 (Patent Document 1)).

[0008] International Publication No. 2020 / 045281

[0009] Y Wu et al. , Cell Stem Cell 13, Dec 5, 2013, p. 659-662Li H. et al. , bioRxiv, 21 Aug 2017, doi: https: / / doi. org / 10.1101 / 178905 Inoue et al. , Cells, 2021, 10, 1076. https: / / doi. org / 10.3390 / cells10051076Nakajima et al. , Genome Research, 28, 2018, p. 223-230 Minczuk et al. , Nucleic Acids Research, 36(12), 2008, p. 3926-3938Yan Luo et al. , Scientific Reports, 6:20657, 2016, doi: 10.1038 / srep20657

[0010] The present invention has been made in consideration of the problems associated with the above-mentioned conventional technologies, and aims to provide a method for specifically and highly efficiently inserting (knock-in) a desired nucleotide sequence into a target region by cleaving single-stranded DNA using only a protein, a method for producing cells in which a desired nucleotide sequence has been inserted into a target region (knock-in cells), and kits for use in these methods.

[0011] As a result of intensive research to achieve the above object, the present inventors have discovered that by using the nuclease domain 1 (ND1 domain) described in Patent Document 1, which is an alternative factor to the FokI nuclease domain, as a dND1 domain into which a nuclease activity-deficient mutation has been introduced, and by combining ND1 and dND1 to form a dimer, a nicking enzyme (nickase active form) can be produced using only the protein, which exhibits nickase activity that can replace the nickase activity of nickase-type Cas9 (nCas9). Furthermore, in order to guide this nicking enzyme to a target region on the target DNA, a DNA-binding domain is bound to each domain to form a fusion protein, and a site-specific nicking system has been successfully developed.

[0012] Based on the description in Non-Patent Document 6 in which the FokI nuclease domain was modified, it is expected that a nick would be introduced into the strand that a domain that has not been introduced with a nuclease activity-deficient mutation recognizes and binds to. However, in contrast to this, the inventors surprisingly found that with the above-mentioned nicking enzyme in which the ND1 domain has been modified, a nick is introduced into the strand that is recognized by the dND1 domain in which a nuclease activity-deficient mutation has been introduced.

[0013] Furthermore, the present inventors have found that in this new nicking system, the use of single-stranded DNA (ssDNA) as the donor DNA, rather than a donor plasmid, significantly increases the knock-in efficiency. This is a phenomenon specific to this nicking system that is not observed with nickase-type Cas9, and led to the completion of the present invention.

[0014] The present invention is provided based on these findings in the following aspects. [1] A method for inserting a donor sequence into a target region on a target DNA, the method comprising the step of contacting the target DNA with: donor DNA, which is a single-stranded DNA comprising a nucleotide sequence arranged in the following order from the 5' side: a 5' homology arm sequence, a donor sequence, and a 3' homology arm sequence; and a site-specific nicking system comprising: a first fusion protein comprising a first DNA-binding domain and an ND1 domain; and a second fusion protein comprising a second DNA-binding domain and a dND1 domain, wherein the ND1 domain and the dND1 domain form a dimer to introduce a nick within or near the target region, wherein the ND1 domain is one of the following (a) to (c): (a) a polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 40; or (b) a polypeptide comprising an amino acid sequence in which one or more amino acids in the amino acid sequence set forth in SEQ ID NO: 40 have been substituted, deleted, inserted, and / or added, and in which the amino acids corresponding to the 66th and 83rd amino acids in the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid. (c) an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and comprising an amino acid sequence in which the amino acids corresponding to the 66th and 83rd amino acids of the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid, and the dND1 domain is at least one polypeptide selected from the group consisting of (da) to (dc) below: (da) a polypeptide comprising an amino acid sequence in which the aspartic acid at the 66th position and / or the aspartic acid at the 83rd position of the amino acid sequence set forth in SEQ ID NO: 40 are substituted with any other amino acid; (db) a polypeptide comprising an amino acid sequence in which one or more amino acids are substituted, deleted, inserted and / or added in the amino acid sequence set forth in SEQ ID NO: 40, and wherein the amino acid corresponding to the aspartic acid at the 66th position and / or the aspartic acid at the 83rd position of the amino acid sequence set forth in SEQ ID NO: 40 are substituted with any amino acid other than aspartic acid.(dc) an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and comprising an amino acid sequence in which the amino acid corresponding to the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any amino acid other than aspartic acid. [2] The method according to [1], wherein the donor DNA is a long single-stranded DNA, the lengths of which are independently 100 bases or more for the 5' homology arm sequence and the 3' homology arm sequence. [3] A method for producing a cell in which a donor sequence has been inserted into a target region on a target DNA, the method comprising the steps of introducing or expressing into a cell and contacting with the target DNA: donor DNA, which is a single-stranded DNA comprising a nucleotide sequence arranged in the following order from the 5' side: a 5' side homology arm sequence, a donor sequence, and a 3' side homology arm sequence; and a site-specific nicking system comprising: a first fusion protein comprising a first DNA-binding domain and an ND1 domain; and a second fusion protein comprising a second DNA-binding domain and a dND1 domain, wherein the ND1 domain and the dND1 domain form a dimer to introduce a nick within or near the target region; and the donor DNA is a single-stranded DNA comprising a nucleotide sequence arranged in the following order from the 5' side: a 5' side homology arm sequence, a donor sequence, and a 3' side homology arm sequence. (b) a polypeptide comprising an amino acid sequence in which one or more amino acids have been substituted, deleted, inserted and / or added in the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acids corresponding to the 66th and 83rd amino acids in the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid; and (c) a polypeptide having an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acids corresponding to the 66th and 83rd amino acids in the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid, and the dND1 domain is a domain comprising at least one polypeptide selected from the group consisting of the following (da) to (dc):(da) a polypeptide comprising an amino acid sequence in which the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any other amino acid; (db) a polypeptide comprising an amino acid sequence in which one or more amino acids are substituted, deleted, inserted and / or added in the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acid corresponding to the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any other amino acid; and (dc) a polypeptide having an amino acid sequence having 90% or more homology to the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acid corresponding to the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any other amino acid. [4] The method according to [3], wherein the donor DNA is a long single-stranded DNA, the lengths of which are independently 100 bases or more for the 5' homology arm sequence and the 3' homology arm sequence. [5] A kit for use in the method according to any one of [1] to [4], comprising: at least one selected from the group consisting of a first fusion protein, a first fusion protein expression vector that expresses the first fusion protein, and a polynucleotide encoding the first fusion protein; and at least one selected from the group consisting of a second fusion protein, a second fusion protein expression vector that expresses the second fusion protein, and a polynucleotide encoding the second fusion protein, wherein the first fusion protein expression vector is at least one selected from the group consisting of (i) a vector comprising a polynucleotide encoding an ND1 domain and a first DNA-binding domain, and (ii) a vector comprising a polynucleotide encoding the ND1 domain and an insertion site for the polynucleotide encoding the first DNA-binding domain,A kit, wherein the second fusion protein expression vector is at least one selected from the group consisting of (iii) a vector comprising a polynucleotide encoding a dND1 domain and a second DNA-binding domain, and (iv) a vector comprising a polynucleotide encoding a dND1 domain and an insertion site for a polynucleotide encoding the second DNA-binding domain.

[0015] According to the present invention, it is possible to provide a method for specifically and highly efficiently inserting (knock-in) a desired nucleotide sequence into a target region using only a protein, a method for producing a cell in which a desired nucleotide sequence has been inserted into a target region (knock-in cell), and kits for use in these methods.

[0016] 1 is a conceptual diagram showing the positional relationship between the PAM sequences on the target genes APC (a) and Rosa26 (b), the cleavage points by nCas9 (D10A), and each TALE and guide RNA used in Test Example 1. Graph showing EGFP fluorescence intensity (RFI) when a combination of TALE47(APC-L)-ND1 and TALE47(APC-R)-ND1, a combination of TALE47(APC-L)-dND1 and TALE47(APC-R)-ND1, or a combination of TALE47(APC-L)-ND1 and TALE47(APC-R)-dND1, or a combination of TALE47(APC-L)-ND1 and TALE47(APC-R)-dND1, or these were further treated with nCas9(D10A) and guide RNA (APC_gRNA-C or APC_gRNA-B) for SSA assay reporter plasmid EGxAPCxFP loaded with APC as the target gene. This graph shows the EGFP fluorescence intensity (RFI) when the SSA assay reporter plasmid EGxRosa26xFP, which is equipped with Rosa26 as the target gene, is treated with a combination of TALE47(Rosa26-L)-ND1 and TALE47(Rosa26-R)-ND1, a combination of TALE47(Rosa26-L)-dND1 and TALE47(Rosa26-R)-ND1, or a combination of TALE47(Rosa26-L)-ND1 and TALE47(Rosa26-R)-dND1, or a combination of TALE47(Rosa26-L)-ND1 and TALE47(Rosa26-R)-dND1, or these, further treated with nCas9(D10A) and guide RNA (Rosa26_gRNA-C or Rosa26_gRNA-B). Fluorescence images of EGFP in the combinations ((a): Mock, (b) SNGD, (c) dND1 + ND1 + PD, (d) dND1 + ND1 + lssDNA) shown in Table 9. Graph showing the average value and standard deviation of the number of EGFP-positive cells / total cell count (%) for each of four wells obtained from the fluorescence images of EGFP in the combinations ((a): Mock, (b) SNGD, (c) dND1 + ND1 + PD, (d) dND1 + ND1 + lssDNA) shown in Table 9.Graph showing the relative value (relative activity) and standard deviation of the average value of the number of EGFP-positive cells / total cell count (%) for four wells obtained from the fluorescent images of each EGFP for the combinations shown in Table 10 ((a): PD-0 (base), (b): 20 ng lssDNA, (c): 40 ng lssDNA, (d): 60 ng lssDNA), where (a) the base value is set to 1.

[0017] The present invention will be described in more detail below by taking preferred embodiments as examples, but the present invention is not limited thereto.

[0018] <Nicking Enzyme> The nicking enzyme of the present invention is an artificial nicking enzyme comprising two domains, ND1 and dND1. The nicking enzyme of the present invention has a single-stranded DNA cleavage activity (nickase activity) that cleaves only one strand of double-stranded DNA.

[0019] (ND1 Domain) The "ND1 domain" according to the present invention is nuclease domain 1, which is one of the nuclease domains discovered by the present inventors through screening of homologous sequences with identities in the range of 35 to 70% to the FokI nuclease domain (Patent Document 1). A typical nucleotide sequence of the ND1 domain is (a) a polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 40, more preferably a polypeptide consisting of the amino acid sequence set forth in SEQ ID NO: 40. The amino acid sequence set forth in SEQ ID NO: 40 has 70% identity with the amino acid sequence of the FokI nuclease domain.

[0020] The "ND1 domain" according to the present invention includes polypeptides comprising an amino acid sequence highly homologous to the amino acid sequence set forth in SEQ ID NO: 40, as long as the polypeptide has double-stranded DNA cleavage activity (nuclease activity) when two molecules of the ND1 domain are used. Such nuclease domains include, for example, ND1 domains derived from other bacteria and modified ND1 domains (natural mutants and artificial mutants).

[0021] Therefore, embodiments of the "ND1 domain" according to the present invention include (b) a polypeptide comprising an amino acid sequence in which one or more amino acids have been substituted, deleted, inserted, and / or added in the amino acid sequence set forth in SEQ ID NO: 40, more preferably a polypeptide consisting of said amino acid sequence. However, since it is necessary that the sequence is not a mutant sequence lacking nuclease activity as described below, in (b), at least the amino acids corresponding to positions 66 and 83 of the amino acid sequence set forth in SEQ ID NO: 40 must be aspartic acid.

[0022] Here, in the amino acid sequence, "a substituted, deleted, inserted, and / or added amino acid sequence" refers to an amino acid sequence in which amino acids (amino acid residues) in the amino acid sequence have been substituted, deleted, inserted, or added, or an amino acid sequence in which two or more of these have been combined. Furthermore, "multiple" refers to an integer of 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2. In the amino acid sequence of the polypeptide (b), the one or more preferably refers to 1 to 30, 1 to 25, or 1 to 20 amino acid residues (e.g., 1 to 10, 1 to 5, 1 to 3, or 2 or less).

[0023] Furthermore, embodiments of the "ND1 domain" according to the present invention also include (c) a polypeptide comprising an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, more preferably a polypeptide consisting of said amino acid sequence. However, even in (c), it is necessary that the sequence is not a nuclease activity-deficient mutant sequence described below, and therefore, at least the amino acids corresponding to positions 66 and 83 of the amino acid sequence set forth in SEQ ID NO: 40 must be aspartic acid.

[0024] Here, when referring to amino acid sequence homology, it means that when a control amino acid sequence (e.g., the amino acid sequence of the polypeptide (a)) and a target amino acid sequence (e.g., the amino acid sequence of the polypeptide (c)) are aligned using amino acid sequence analysis software or the like, the amino acid in the target amino acid sequence (target amino acid) that is at the same position as an amino acid in the control amino acid sequence (control amino acid) may be the same amino acid as the control amino acid, or an amino acid with the same properties as the control amino acid. When referring to amino acid sequence identity, the target amino acid is the same amino acid as the control amino acid. Groups of amino acids with similar properties are well known in the art to which the present invention pertains, and can be classified into, for example, acidic amino acids (aspartic acid and glutamic acid); basic amino acids (lysine, arginine, histidine); and neutral amino acids can be classified into amino acids with hydrocarbon chains (glycine, alanine, valine, leucine, isoleucine, proline), amino acids with hydroxy groups (serine, threonine), amino acids containing sulfur (cysteine, methionine), amino acids with amide groups (asparagine, glutamine), amino acids with imino groups (proline), and amino acids with aromatic groups (phenylalanine, tyrosine, tryptophan).

[0025] The homology and identity of such amino acid sequences are determined by comparing two sequences aligned to maximize sequence identity. Methods for determining the numerical value (%) of sequence homology or identity are known to those skilled in the art. Any algorithm known to those skilled in the art (e.g., the BLAST algorithm, the FASTA algorithm, etc.) can be used to obtain optimal alignment and sequence identity, and the sequence homology or identity of amino acid sequences can be determined using sequence analysis software such as BLASTP or FASTA. Furthermore, the homology of the amino acid sequence of the polypeptide (c) with the amino acid sequence of the polypeptide (a) may be 90% or more, but is preferably 95% or more (e.g., 96% or more, 97% or more, 97.5% or more, 98% or more, 98.5% or more, 99% or more, 99.5% or more, or 99.8% or more). More preferably, the identity is 90% or more, 95% or more (e.g., 96% or more, 97% or more, 97.5% or more, 98% or more, 98.5% or more, 99% or more, 99.5% or more, 99.8% or more).

[0026] Furthermore, in an amino acid sequence, an amino acid that "corresponds to" a specific amino acid refers to an amino acid that is at the same position as the specific amino acid (the reference amino acid (reference amino acid residue), for example, the aspartic acid at position 66 or 83 in the amino acid sequence set forth in SEQ ID NO: 40) when amino acid sequences are aligned using amino acid sequence analysis software (e.g., GENETYX-MAC, Sequencher, ClustalW, etc.) (e.g., parameters: default values ​​(i.e., initial settings)).

[0027] (dND1 domain) The "dND1 domain" according to the present invention is a modified domain in which the amino acid sequence of the ND1 domain has been modified to a mutant sequence (nuclease activity-deficient mutant sequence) that lacks nuclease activity. Such a nuclease activity-deficient mutant sequence of the dND1 domain, when used in combination with the ND1 domain, lacks double-stranded DNA cleavage activity and only needs to have single-stranded DNA cleavage activity. Typically, (da) a polypeptide comprising an amino acid sequence in which the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any other amino acid, more preferably a polypeptide consisting of the amino acid sequence containing the substituted amino acid. The aspartic acid substituted by the substituted amino acid may be either or both of the aspartic acid at position 66 and the aspartic acid at position 83, but preferably contains at least the aspartic acid at position 66.

[0028] Furthermore, the "dND1 domain" of the present invention also includes variants consisting of an amino acid sequence having high homology to such (da). Therefore, embodiments of the "dND1 domain" of the present invention include (db) an amino acid sequence in which one or more amino acids are substituted, deleted, inserted, and / or added in the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acids corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 are substituted with any amino acid other than aspartic acid; and (dc) an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any amino acid other than aspartic acid.

[0029] In the amino acid sequence of the polypeptide (db), the one or more amino acid residues are preferably 1 to 30, 1 to 25, or 1 to 20 (e.g., 1 to 10, 1 to 5, 1 to 3, or 2 or less). Furthermore, in the amino acid sequence of the polypeptide (dc), the homology with the amino acid sequence of (da) may be 90% or more, but is preferably 95% or more (e.g., 96% or more, 97% or more, 97.5% or more, 98% or more, 98.5% or more, 99% or more, 99.5% or more, or 99.8% or more). More preferably, the identity is 90% or more, 95% or more (e.g., 96% or more, 97% or more, 97.5% or more, 98% or more, 98.5% or more, 99% or more, 99.5% or more, or 99.8% or more).

[0030] In (db) and (dc), the aspartic acid substituted with the substituted amino acid may be either or both of the amino acid corresponding to the aspartic acid at position 66 and the amino acid corresponding to the aspartic acid at position 83, but preferably includes at least the amino acid corresponding to the aspartic acid at position 66. Furthermore, in (da), (db), and (dc), the substituted amino acid when substituted with any amino acid other than aspartic acid is preferably an amino acid other than the acidic amino acids, and more preferably alanine or asparagine.

[0031] As a method for substituting the aspartic acid at position 66 and / or the aspartic acid at position 83, or an amino acid corresponding to such aspartic acid, with the substituted amino acid, a conventionally known method or a method similar thereto can be appropriately adopted. For example, a conventionally known method such as site-directed mutagenesis (Kunkel et al., Proc. Natl. Acad. Sci. USA (1985) 82, pp. 488-492), overlap extension PCR, or the like can be appropriately adopted. Furthermore, primers and the like used for the amino acid modification can be appropriately designed based on the amino acid sequence of interest (such as the amino acid sequence set forth in SEQ ID NO: 40) and the desired modification by a conventionally known method or a method similar thereto.

[0032] In the present invention, whether the ND1 domain and dND1 domain, when used in combination with the ND1 domain, possess or lack double-stranded DNA cleavage activity (nuclease activity) can be appropriately confirmed by methods known to those skilled in the art or methods similar thereto. For example, as described in the test examples below, a fusion product in which "TALE-L → domain to be confirmed" is linked in this order from the N-terminus, and a fusion product in which "TALE-R → amino acid sequence set forth in SEQ ID NO: 40 (ND1 domain)" is linked in this order from the N-terminus, are expressed in cells, and confirmation can be made by determining whether double-stranded DNA cleavage occurs in the target DNA corresponding to the TALE. In this method, an SSA assay system targeting a reporter gene on a plasmid (e.g., an EGFP gene into which a TALE recognition sequence has been inserted) may be used, and reporter activity may be used as an indicator for evaluation. For example, when the nuclease activity when the above-mentioned combination of fusions is used is 20% or more of the nuclease activity when a fusion in which "TALE-L → amino acid sequence described in SEQ ID NO: 40 (ND1 domain)" is linked in this order from the N-terminus and a fusion in which "TALE-R → amino acid sequence described in SEQ ID NO: 40 (ND1 domain)" is linked in this order from the N-terminus is expressed in a cell, it can be determined that the nuclease activity is present, and when it is less than 20%, it can be determined that the nuclease activity is lacking.

[0033] Furthermore, in the present invention, whether or not two molecules of the ND1 domain and the dND1 domain have the activity of cleaving single-stranded DNA (nickase activity) can be appropriately confirmed by a method known to those skilled in the art or a method similar thereto. For example, as described in the test examples below, a fusion compound in which "TALE-L → domain 1 to be confirmed" confirmed to have nuclease activity as described above is bound in this order from the N-terminus (or a fusion compound in which "TALE-L → amino acid sequence set forth in SEQ ID NO: 40 (ND1 domain)" is bound in this order from the N-terminus) is combined with a fusion compound in which "TALE-R → domain 2 to be confirmed" confirmed to lack nuclease activity as described above is bound in this order from the N-terminus, and after confirming that nuclease activity is lost, a combination of nCas9, which cleaves single-stranded DNA near and opposite the TALE recognition sequence of the fusion compound confirmed to lack nuclease activity, and a guide RNA is applied to the fusion compound, and it is confirmed whether double-stranded DNA cleavage activity of the target DNA corresponding to the TALE is restored. In this method, the SSA assay system described above may also be used, and reporter activity may be used as an indicator for evaluation.

[0034] <Nicking system> The nicking system of the present invention is a site-specific nicking system that comprises: a first fusion protein comprising a first DNA-binding domain and an ND1 domain; and a second fusion protein comprising a second DNA-binding domain and a dND1 domain, wherein the ND1 domain and the dND1 domain form a dimer to introduce a nick within or near the target region.

[0035] (Fusion Protein) The first fusion protein and second fusion protein according to the present invention (sometimes collectively referred to herein as "fusion proteins") are fusion proteins comprising a DNA-binding domain and the above-described ND1 domain or dND1 domain. The fusion proteins of the present invention each bind to the vicinity of a target region on a target DNA (a DNA-binding domain recognition sequence) via the DNA-binding domain, and the ND1 domain and the dND1 domain form a dimer to cleave (introduce a nick in) one strand of double-stranded DNA within or near the target region, thereby functioning as a site-specific nicking system. The ND1 domain and dND1 domain are as described above, including preferred embodiments thereof.

[0036] (DNA-binding domain) In the present invention, the DNA-binding domain that constitutes the first fusion protein together with the ND1 domain is referred to as the first DNA-binding domain, and the DNA-binding domain that constitutes the second fusion protein together with the dND1 domain is referred to as the second DNA-binding domain, and the first DNA-binding domain and the second DNA-binding domain are sometimes collectively referred to simply as the "DNA-binding domain."

[0037] The "DNA-binding domain" of the present invention is not particularly limited as long as it is a protein domain that can specifically bind to any DNA sequence (DNA-binding domain recognition sequence). Examples of DNA-binding domains include TALE, zinc finger array (ZF), and PPR protein (Pentatricopeptide Repeat Protein). At least one selected from the group consisting of TALE and zinc finger array (ZF) is preferred, and TALE is more preferred. The first DNA-binding domain and the second DNA-binding domain of the present invention may be of the same or different species, but are preferably of the same type (for example, both are TALE).

[0038] [TALE] TALE (Transcription activator-like effector nuclease) is a protein that is typically secreted by proteobacteria of the genus Xanthomonas and activates gene transcription in a host plant, and is a generic term that includes TALE-like proteins possessed by bacteria of the genus Ralstonia. The "TALE" according to the present invention contains at least an N-terminal domain and a TALE repeat domain, and may further contain a C-terminal domain.

[0039] A TALE repeat domain is composed of multiple, for example, 10 to 30, preferably 13 to 25, and more preferably 15 to 20 tandem repeats (TALE repeats) of a TALE sequence that forms a right-handed superhelical coil. A typical TALE repeat unit (one TALE sequence) consists of 33 to 35 amino acids, and recognizes a specific base in DNA by using a repeat variable residue (RVD) consisting of two amino acid residues at the 12th and 13th positions. Examples of RVDs that specifically recognize a base include HD, which recognizes C; NG, which recognizes T; NI, which recognizes A; NN, which recognizes G or A; and NS, which recognizes A, C, G, or T. Based on the DNA recognition mechanism of such TALE repeat domains, by artificially linking TALE sequences that recognize specific bases, it is possible to create TALEs that can recognize and bind to specific nucleotide sequences (TALE recognition sequences) on DNA.

[0040] The TALE sequences of the TALEs of the present invention may be appropriately modified relative to their native amino acid sequences, as long as the TALE repeat domain can recognize and bind to the TALE recognition sequence. For example, the TALE sequences may each independently have a different RVD position, or may have one or more (e.g., 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, more preferably 3 or less or 2 or less) amino acid residues other than the RVD substituted, deleted, inserted, and / or added. Furthermore, the 12th and 13th amino acid residues of the RVD may be substituted with other amino acid residues to enhance specificity for A, T, C, and G.

[0041] The TALE repeat domain of the TALE of the present invention is designed to recognize a nucleotide sequence or its complementary sequence located on the 5' or 3' side of the target region of the target DNA, preferably via a spacer of 0 to 30 bases, i.e., the TALE recognition sequence or its complementary sequence recognized by the TALE is a nucleotide sequence located on the 5' or 3' side of the target region, preferably via a spacer of 0 to 30 bases, more preferably, the distance between the two TALE recognition sequences is 1 to 50 bases. Techniques for designing and producing such desired TALEs are known, and are described, for example, in Miller et al., Nat Biotechnol 29, 2011, pp. 143-148; Sakurama et al., Sci Rep 3, 3379 (2013), and the aforementioned Sakurama et al. (2013) enables the generation of TALEs with high binding activity to the nucleotide sequence.

[0042] The N-terminal domain of the TALE of the present invention can be a naturally occurring amino acid sequence (e.g., the amino acid sequence of the N-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185-32188)). However, the N-terminal domain may be appropriately modified from the naturally occurring amino acid sequence as long as it does not significantly negatively affect the functions of the fusion proteins of the present invention (e.g., binding ability to a TALE recognition sequence, nickase activity, etc.). For example, the N-terminal domain may be modified by substitution, deletion, insertion, and / or addition of one or more amino acid residues (e.g., 50 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, more preferably 3 or less or 2 or less). Furthermore, the N-terminal domain may contain a flag tag for purification or detection, a nuclear localization signal (NLS) for translocating the TALE-ND1 domain and the TALE-dND1 domain to the cell nucleus, etc.

[0043] The chain length of such an N-terminal domain is preferably 49 to 287 amino acid residues, more preferably 80 to 200 amino acid residues, and even more preferably 120 to 180 amino acid residues.

[0044] Specific examples of the N-terminal domain of a TALE according to the present invention include, in addition to the above, the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-136 / 63-VR-HD (ID: 50699) and the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-153 / 47-VR-HD (ID: 50703), but are not limited to these.

[0045] When the TALE of the present invention further comprises a C-terminal domain, the C-terminal domain of the TALE can also be a naturally occurring amino acid sequence (e.g., the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185-32188) (WT: number of amino acid residues = 180)). However, the C-terminal domain may be appropriately modified from the naturally occurring amino acid sequence as long as it does not significantly negatively affect the functions of the second fusion protein of the present invention (e.g., binding ability to the TALE recognition sequence, nickase activity, etc.). For example, the C-terminal domain may be modified by substitution, deletion, insertion, and / or addition of one or more amino acid residues (e.g., 50 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, more preferably 3 or less or 2 or less).

[0046] When the TALE of the present invention further comprises a C-terminal domain, the C-terminal domain of the TALE may be a naturally occurring amino acid sequence from which a portion of the C-terminal amino acid sequence has been removed. The chain length of such a C-terminal domain is preferably 1 to 200 amino acid residues, more preferably 5 to 150 amino acid residues, even more preferably 10 to 125 amino acid residues, and even more preferably 10 to 100 amino acid residues.

[0047] Specific examples of the C-terminal domain include, in addition to the above, the amino acid sequence (number of amino acid residues = 63) of the C-terminal domain contained in Addgene's pTALEN_v2 (ID: 32189 to 32192) and the amino acid sequence (number of amino acid residues = 47) of the C-terminal domain contained in Addgene's ptCMV-153 / 47-VR-NG (ID: 50704), but are not limited thereto.

[0048] A preferred embodiment of the TALE of the present invention is a Platinum TALEN (JP 2015-33365 A), in which amino acids at two specific positions in one TALE repeat unit are changed every four TALE repeat units.

[0049] [Zinc Finger Array] The zinc finger array of the present invention represents the DNA-binding domain of a zinc finger nuclease (ZFN) that contains a DNA-binding domain and a nuclease domain. A zinc finger array typically consists of 2 to 15, preferably 3 to 8, and more preferably 4 to 6 zinc finger proteins, which are linked directly or via linkers. Each zinc finger protein contains one or more zinc atoms and a helix structure that recognizes specific bases, and each zinc finger protein recognizes three specific bases in DNA.

[0050] The zinc finger array of the present invention is not particularly limited as long as it can recognize and bind to a DNA-binding domain recognition sequence, and may be appropriately modified, for example, by substitution, deletion, insertion, and / or addition of one or more amino acid residues (e.g., 50 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, and more preferably 3 or less or 2 or less).

[0051] The zinc finger array of the present invention is designed to recognize a nucleotide sequence or its complementary sequence located on the 5' or 3' side of a target site in a target DNA, preferably via a spacer of 1 to 10 bases, i.e., the zinc finger array recognition sequence or its complementary sequence recognized by the zinc finger array is a nucleotide sequence located on the 5' or 3' side of the target site, respectively, via a spacer of 1 to 10 bases. Techniques for designing and preparing such zinc finger arrays are known, and they can be appropriately designed and prepared with reference to methods described, for example, in Beerli et al., Nat Biotechnol 20, 135-141 (2002) and Sander et al., Nat Methods 8, 67-69 (2011).

[0052] (Linker) In the fusion protein of the present invention, each DNA-binding domain and the ND1 domain or dND1 domain may be linked directly or via a linker. The linker is not particularly limited in length or type, as long as it does not significantly negatively affect the function of each fusion protein of the present invention (such as the ability to bind to the DNA-binding domain recognition sequence or nickase activity). The linker length is typically 2 to 180 amino acid residues, and preferably 2 to 120 amino acid residues. For example, when the DNA-binding domain is a TALE, the chain length of the linker can be adjusted so that the distance between the TALE repeat domain and the ND1 domain or dND1 domain is 1 to 200 amino acid residues, preferably 5 to 150 amino acid residues, more preferably 10 to 125 amino acid residues, and even more preferably 10 to 100 amino acid residues, including the chain length of the N-terminal domain of the TALE (when the order is ND1 domain or dND1 domain → DNA-binding domain from the N-terminus) or C-terminal domain (when the order is DNA-binding domain → ND1 domain or dND1 domain from the N-terminus), more preferably the latter. In this case, the distance between the TALE repeat domain and the ND1 domain or dND1 domain is the number of amino acid residues from the N-terminus to the residue adjacent to the N-terminus of the ND1 domain or dND1 domain, starting from the amino acid residue adjacent to the C-terminus of the TALE repeat domain as the first residue, if the order is DNA-binding domain → ND1 domain or dND1 domain; and from the N-terminus to the residue adjacent to the C-terminus of the ND1 domain or dND1 domain, starting from the amino acid residue adjacent to the N-terminus of the TALE repeat domain as the first residue, if the order is ND1 domain or dND1 domain → DNA-binding domain, it is the number of amino acid residues from the amino acid residue adjacent to the N-terminus of the TALE repeat domain as the first residue, starting from the amino acid residue adjacent to the N-terminus of the TALE repeat domain as the first residue,

[0053] The type of the linker is not particularly limited, and examples thereof include an HTS95 linker consisting of the HTS95 amino acid sequence described in Sun, N., & Zhao, H. (2014) Molecular BioSystems, 10(3), pp. 446-453, a GSS linker, a SAGG linker, a GGGGS linker, an XTEN linker, and the like.

[0054] In the fusion protein of the present invention, the DNA-binding domain and the ND1 domain or dND1 domain may be located either at the N-terminus or the C-terminus, with the linker interposed therebetween as necessary. Among these, the fusion protein of the present invention is preferably one in which the domains are linked from the N-terminus in the order of DNA-binding domain → (the linker as necessary) → ND1 domain or dND1 domain, from the viewpoint of exhibiting excellent nickase activity that is particularly specific to a DNA strand, i.e., specific to a strand having a DNA-binding domain recognition sequence to which the second fusion protein binds.

[0055] In the fusion proteins of the present invention, the DNA-binding domain and the ND1 domain or dND1 domain can be linked at the nucleic acid level and the amino acid level. That is, polynucleotides encoding each of the fusion proteins of the present invention can be prepared by ligating the respective polynucleotides in the order of "polynucleotide encoding the ND1 domain or dND1 domain → polynucleotide encoding the DNA-binding domain" or "polynucleotide encoding the DNA-binding domain → polynucleotide encoding the ND1 domain or dND1 domain." The polynucleotides thus obtained can be inserted into an expression vector and expressed in an appropriate host cell, thereby obtaining each of the fusion proteins of the present invention in which the DNA-binding domain and the ND1 domain are linked at the amino acid level in the order of "ND1 domain or dND1 domain → DNA-binding domain" or "DNA-binding domain → ND1 domain or dND1 domain."

[0056] The expression vector can be appropriately selected from vectors used in the field, and examples thereof include plasmid vectors, virus vectors, phage vectors, phagemid vectors, BAC vectors, YAC vectors, MAC vectors, and HAC vectors. Host cells into which the expression vector is introduced can be appropriately selected in consideration of compatibility with the expression vector, and examples thereof include prokaryotic cells such as Escherichia coli, actinomycetes, and archaea, and eukaryotic cells such as yeast, sea urchin, silkworm, zebrafish, mouse, rat, frog, tobacco, Arabidopsis, and rice. The term "polynucleotide" encompasses both DNA and RNA (mRNA), and each polynucleotide may have its codons optimized for the purpose of increasing expression efficiency in cells.

[0057] The fusion protein of the present invention can also be prepared by artificial synthesis based on amino acid sequence information. Furthermore, the fusion protein of the present invention may be added with an epitope tag (e.g., flag tag, HA tag, etc.) for purification or detection, a labeled protein (e.g., fluorescent protein, etc.), or various transport signals (e.g., nuclear transport signal, plastid transport signal, mitochondrial transport signal, etc.).

[0058] <Method for knocking in a desired nucleotide sequence> The present invention provides a method for inserting a donor sequence into a target region on a target DNA, which comprises the step of contacting the target DNA with donor DNA, which is single-stranded DNA comprising a nucleotide sequence arranged in the following order from the 5' side: 5' homology arm sequence, donor sequence, 3' homology arm sequence; and the nicking system of the present invention (sometimes referred to herein simply as the "knock-in method of the present invention").

[0059] (Target DNA) In the present invention, the region into which a desired nucleotide sequence, i.e., the donor sequence of the donor DNA described below, is inserted is defined as the target region, and DNA containing this target region is referred to as "target DNA." The target DNA according to the present invention is double-stranded DNA, and for the sake of convenience, in order to show its correspondence with the nicking system of the present invention described above, at least one strand is defined as having a structure containing, from the 5' side, a 5' DNA-binding domain recognition sequence or its complementary sequence, optionally the spacer, the target region, optionally the spacer, and a 3' DNA-binding domain recognition sequence or its complementary sequence. The 5' DNA-binding domain recognition sequence and the 3' DNA-binding domain recognition sequence may be configured on the same strand or on opposite strands, but are preferably configured on opposite strands.

[0060] Furthermore, the target DNA of the present invention contains, adjacent to the 5' side of the target region, a 5' side sequence that is identical to the 5' side homology arm sequence (5'HA) of the donor DNA described below, and also contains, adjacent to the 3' side, a 3' side sequence that is identical to the 3' side homology arm sequence (3'HA) of the donor DNA described below. The 5' side sequence and 3' side sequence may overlap with the 5' DNA-binding domain recognition sequence or its complementary sequence, or the 3' DNA-binding domain recognition sequence or its complementary sequence, respectively.

[0061] The "target region" according to the present invention refers to a region into which a donor sequence of the donor DNA described below is inserted or replaced with the donor sequence. The target region is a region containing the cleavage site cleaved by the nicking enzyme or a region in the vicinity thereof. The length of the target region may be a single location between two adjacent bases at the cleavage site or two bases located in the vicinity thereof, a single base located adjacent to the cleavage site or in the vicinity thereof, or two or more bases located in the vicinity of the cleavage site or including the cleavage site. In the present invention, "inserting a donor sequence into a target region" includes not only an embodiment in which the donor sequence is inserted between two bases located on either side of the cleavage site or in the vicinity thereof, but also an embodiment in which one or more bases of the target region are replaced with the donor sequence. The length of the target region of one or more bases is, for example, preferably 1 to 500 bases, more preferably 2 to 300 bases, and even more preferably 2 to 100 bases.

[0062] In the present invention, the term "vicinity" in the "vicinity of the cleavage site" preferably means that the number of bases between the cleavage site and the nucleic acid is within 100 bases, and more preferably within 50 bases.

[0063] The "5' DNA-binding domain recognition sequence" and "3' DNA-binding domain recognition sequence" according to the present invention are sequences recognized by either the first DNA-binding domain or the second DNA-binding domain, respectively. The first DNA-binding domain and the second DNA-binding domain may bind to each other's DNA-binding domain recognition sequence, and either may bind to the 5' DNA-binding domain recognition sequence or the 3' DNA-binding domain recognition sequence.

[0064] The lengths of the "5' DNA-binding domain recognition sequence" and "3' DNA-binding domain recognition sequence" (sometimes collectively referred to herein simply as "DNA-binding domain recognition sequence") according to the present invention are set depending on the type of DNA-binding domain, but for example, when each of the DNA-binding domains is a TALE, each of the DNA-binding domains is preferably independently 10 to 30 bases, preferably 13 to 25 bases, and more preferably 15 to 22 bases. Furthermore, for example, when each of the DNA-binding domains is a zinc finger array, each of the DNA-binding domains is independently preferably 6 to 45 bases, preferably 9 to 24 bases, and more preferably 12 to 18 bases.

[0065] The distance between the DNA-binding domain recognition sequences according to the present invention (the length from the base adjacent to the 5'- or 3'-side of the 5' DNA-binding domain recognition sequence or its complementary sequence as the first base to the base adjacent to the 3'- or 5'-side of the 3' DNA-binding domain recognition sequence or its complementary sequence, including the target region) is preferably 1 to 50 bases, more preferably 2 to 40 bases, and more preferably 3 to 30 bases, from the viewpoints of enabling the nicking enzyme of the present invention to form a dimer and to specifically and highly efficiently introduce nicks within or in the vicinity of the target region.

[0066] The nucleotide sequence of such target DNA is not particularly limited, and the target DNA can be a target for the knock-in method of the present invention by designing the donor DNA described below and each of the DNA-binding domains in the first fusion protein and second fusion protein so that the 5' DNA-binding domain recognition sequence or its complementary sequence, the 3' DNA-binding domain recognition sequence or its complementary sequence, the distance between the DNA-binding domain recognition sequences, the 5' side sequence, and the 3' side sequence satisfy the above-mentioned conditions.

[0067] The target DNA according to the present invention can be DNA present within a cell (endogenous DNA) or DNA present outside a cell, depending on the purpose. The DNA present within a cell may be endogenous DNA or exogenous DNA. Examples of endogenous DNA include intranuclear genomic DNA, chloroplast DNA, and mitochondrial DNA, while examples of exogenous DNA include DNA introduced into a cell. When the target DNA according to the present invention is DNA present within a cell (endogenous DNA), the DNA is more preferably intranuclear genomic DNA or exogenous DNA. The extracellular DNA may be DNA derived from a cell or DNA amplified and synthesized outside the cell.

[0068] (ssDNA) The knock-in method of the present invention is characterized by using single-stranded DNA (ssDNA) containing a donor sequence as the donor DNA.

[0069] In the present invention, the term "donor sequence" refers to a desired nucleotide sequence intended to be inserted into the target region of the target DNA. Any nucleotide sequence can be used without particular limitations. The knock-in method of the present invention replaces the target region with this donor sequence. As an example, the donor sequence can be a sequence in which a mutation (e.g., base substitution, deletion, addition, or insertion) has been introduced into the target region. As another example, the donor sequence can be a sequence containing a foreign DNA sequence (e.g., a recombinase recognition sequence, a recombinase expression cassette, a drug resistance gene expression cassette, a negative selection marker expression cassette, a fluorescent protein expression cassette, etc.). Alternatively, the donor sequence can be a sequence in which the foreign DNA sequence has been introduced into the target region (or a mutated sequence thereof).

[0070] The length of such a donor sequence is not particularly limited and can be, for example, 1 to 10,000 bases, preferably 10 to 10,000 bases, more preferably 30 to 5,000 bases, and even more preferably 50 to 3,000 bases.

[0071] The donor DNA of the present invention also comprises, on the 5' side of the donor sequence, a 5' homology arm sequence (sometimes referred to herein as "5' HA") that is identical to the nucleotide sequence adjacent to the 5' side of the target region (5' side sequence), and, on the 3' side, a 3' homology arm sequence (sometimes referred to herein as "3' HA") that is identical to the nucleotide sequence adjacent to the 3' side of the target region (3' side sequence).

[0072] When the single-stranded DNA of the present invention is designed based on the target region on the sense strand, 5'HA is the same sequence as the nucleotide sequence adjacent to the 5' side of the target region on the sense strand, and 3'HA is the same sequence as the nucleotide sequence adjacent to the 3' side of the target region on the sense strand.Furthermore, when the single-stranded DNA of the present invention is designed based on the target region on the antisense strand, 5'HA is the same sequence as the nucleotide sequence adjacent to the 5' side of the target region on the antisense strand, and 3'HA is the same sequence as the nucleotide sequence adjacent to the 3' side of the target region on the antisense strand.

[0073] The 5'HA and 3'HA of the present invention (sometimes collectively referred to herein as "homology arms") do not have to be completely identical (i.e., 100% identical) to the 5' or 3' sequence on the target DNA as long as they are substantially identical; for example, they may each independently have an identity of 80% or more, preferably 85% or more, more preferably 90% or more, even more preferably 95% or more, and even more preferably 97% or more, 98% or more, or 99% or more.

[0074] The length of each of the homology arms is preferably 10 to 800 bases, more preferably 20 to 750 bases, and even more preferably 30 to 700 bases. Among these, the donor DNA according to the present invention is preferably a long single-stranded DNA (lssDNA) in which the length of each of the homology arms is independently 100 bases or more, and the length of each of the homology arms is more preferably 100 to 700 bases, and even more preferably 100 to 600 bases.

[0075] The donor DNA of the present invention is preferably a single-stranded DNA consisting only of a nucleotide sequence arranged in the following order from the 5' end: 5' HA, donor sequence, 3' HA. However, the donor DNA is not limited to this as long as it does not significantly inhibit the effects of the present invention (such as knock-in efficiency). For example, a nucleotide sequence consisting of any number of bases (preferably 100 bases or less) may be added to the 5' end and / or 3' end.

[0076] The length (total length) of such donor DNA is not particularly limited, but can be, for example, 21 to 11,600 bases, preferably 40 to 10,000 bases, more preferably 60 to 6,000 bases, and even more preferably 100 to 4,000 bases.

[0077] The donor DNA of the present invention can be prepared by known methods for preparing single-stranded DNA or methods similar thereto. For example, it can be prepared using PCR, restriction enzyme cleavage, DNA ligation techniques, etc. In addition, according to the method described in Inoue et al., Cells, 2021, 10, 1076. https: / / doi.org / 10.3390 / cells10051076 (Non-Patent Document 3), a single-stranded DNA sequence is amplified by PCR using a primer set in which one of the primers is phosphorylated at the 5' end, and the strand of the resulting double-stranded DNA with the phosphorylated 5' end is digested from the phosphorylated end, thereby preparing the desired single-stranded DNA even in the case of long single-stranded DNA.

[0078] (Knock-in) In the knock-in method of the present invention, the donor DNA and a nicking system are brought into contact with the target DNA, and the nicking system is used to cleave within or near the target region, thereby inserting the donor sequence into the target region.

[0079] In the present invention, when referring to the position where a nick is to be introduced as "in the vicinity of the target region," "in the vicinity" preferably refers to a position within 100 bases, and more preferably within 50 bases, counting the base adjacent to the target region as the first base.

[0080] When the donor DNA and the nicking system, i.e., the first fusion protein and the second fusion protein, are contacted with a target DNA, the DNA-binding domains of the first fusion protein and the second fusion protein recognize and bind to the corresponding DNA-binding domain recognition sequences on the target DNA, thereby directing the ND1 domain linked to the first DNA-binding domain and the dND1 domain linked to the second DNA-binding domain to the vicinity of the target region on the target DNA. As a result, within or near the target region, a nicking enzyme comprising a dimer of the ND1 domain and the dND1 domain cleaves one strand of the double-stranded DNA to introduce a nick. For example, if the first fusion protein is configured to be linked in the order "TALE → ND1" from the N-terminus, and the second fusion protein is configured to be linked in the order "TALE → dND1" from the N-terminus, a nick can be selectively introduced into the strand containing the DNA-binding domain recognition sequence to which the second fusion protein binds. As a result of the introduction of the nick, the donor sequence is inserted into the target region due to the correlation between the homology arms and the identical sequences (5' sequence, 3' sequence) between the donor DNA and the target DNA.

[0081] The knock-in method of the present invention may be performed intracellularly or in a cell-free system. The "inside a cell" where the knock-in method of the present invention is performed may be a eukaryotic cell or a prokaryotic cell, preferably a eukaryotic cell. Examples of eukaryotic cells include animal cells (e.g., cells of mammals, fish, birds, reptiles, amphibians, insects), plant cells, algae cells, and yeast. Examples of prokaryotic cells include Escherichia coli, Salmonella, Bacillus subtilis, lactic acid bacteria, and extreme thermophiles.

[0082] "Animal cells" include, for example, cells constituting an individual animal, cells constituting organs or tissues removed from an animal, and cultured cells derived from animal tissues. Specific examples include germ cells such as oocytes and sperm; germ cells of various stages of embryos (e.g., 1-cell, 2-cell, 4-cell, 8-cell, 16-cell, and morula stages); stem cells such as induced pluripotent stem (iPS) cells and embryonic stem (ES) cells; and somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells, brain cells, and kidney cells. Pre- and post-fertilization oocytes can be used as oocytes for producing knock-in animals, but post-fertilization oocytes, i.e., fertilized eggs, are preferred. Pronuclear stage fertilized eggs are particularly preferred. Oocytes can be used by thawing cryopreserved oocytes.

[0083] "Plant cells" include, for example, cells that constitute an individual plant, cells that constitute organs or tissues separated from a plant, cultured cells derived from plant tissue, etc. Examples of plant organs and tissues include leaves, stems, shoot tips (growing points), roots, tubers, calli, etc.

[0084] Furthermore, the "cell-free system" that serves as the site for the knock-in method of the present invention refers to a system that does not contain living cells (the eukaryotic cells or prokaryotic cells). The cell-free system according to the present invention is not particularly limited as long as it allows contact between the donor DNA and nicking system and the target DNA, and examples thereof include a buffer solution; a cell lysate of the eukaryotic or prokaryotic cells; and a cell extract.

[0085] The method for contacting the donor DNA and nicking system with the target DNA is not particularly limited. Intracellular contact includes, for example, a method of introducing or expressing the donor DNA and nicking system into cells containing the target DNA, as in the method for producing knock-in cells described below. In a cell-free system, for example, a solution of target DNA may be mixed with a solution of the donor DNA and nicking system. The solvent for these solutions is not particularly limited, but is preferably, for example, a buffer solution such as phosphate buffer, Tris buffer, Good's buffer, or borate buffer.

[0086] <Method for producing knock-in cells> The present invention also provides a method for producing a cell in which a donor sequence has been inserted into a target region on target DNA, the method comprising the steps of introducing or expressing the donor DNA and the nicking system of the present invention into a cell and contacting them with the target DNA (also sometimes referred to in the present specification as the "method for producing knock-in cells of the present invention").

[0087] In the method for producing a knock-in cell of the present invention, the target DNA, donor DNA, and nicking system are each as described above, including their preferred embodiments. The target DNA in the method for producing a knock-in cell of the present invention is intracellular genomic DNA, more preferably intranuclear genomic DNA, and the donor DNA, first fusion protein, and second fusion protein can be designed depending on the purpose of editing the genomic DNA. Examples of the cells include the cells listed above as targets for the knock-in method of the present invention.

[0088] Furthermore, in the method for producing a knock-in cell of the present invention, the donor DNA and nicking system are brought into contact with the target DNA by introducing the donor DNA and nicking system into a cell, or by introducing the donor DNA into a cell and expressing the nicking system in the cell. Therefore, the nicking system of the present invention may be such that each of the fusion proteins constituting the system is introduced into a cell in the form of a protein, or in the form of RNA or DNA (polynucleotide) encoding the protein and expressed in the cell, or in the form of a vector (expression vector) containing the polynucleotide and expressing the protein and expressed in the cell.

[0089] When each of the fusion proteins is introduced into a cell in the form of an expression vector to be expressed in the cell, for example, vectors expressing each of the fusion proteins may be introduced into the cell individually, or a vector expressing a combination of these may be introduced into the cell.

[0090] Furthermore, when the fusion proteins are introduced into cells in the form of an expression vector or polynucleotide and expressed in the cells, the polynucleotides encoding the fusion proteins may be independently codon-optimized as appropriate for the cells to be introduced. Furthermore, the expression vector preferably includes a promoter and / or other regulatory sequence operably linked to the polynucleotide to be expressed. Here, "operably linked" means that the polynucleotide is expressibly linked to the regulatory element. "Regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, polyadenylation signals, etc.). Examples of promoters include Pol III promoters, Pol II promoters, Pol I promoters, and combinations thereof. Those skilled in the art can select an appropriate expression vector depending on the type of cell to be introduced. Furthermore, the expression vector is preferably one that can stably express the encoded protein without being integrated into the host genome. Such fusion protein expression vectors can be constructed according to conventional methods.

[0091] As a method for introducing the donor DNA, each of the fusion proteins, a polynucleotide encoding each of the fusion proteins, or a vector expressing each of the fusion proteins into cells, a known method for introducing proteins, DNA, or RNA fragments into cells can be appropriately adopted depending on the type of cell. Examples of such methods include electroporation, microinjection, particle gun method, calcium phosphate method, polyethyleneimine (PEI) method, liposome method (lipofection method), DEAE-dextran method, cationic lipid-mediated transfection, viruses (adenovirus, lentivirus, adeno-associated virus, baculovirus, etc.), Agrobacterium method, lithium acetate method, spheroplast method, and heat shock method (calcium chloride method, rubidium chloride method). Such methods are described in many standard laboratory manuals, such as Davis et al., Basic methods in molecular biology, New York: Elsevier, 1986.

[0092] When the donor DNA and nicking system are introduced into or expressed in a cell, the donor DNA, each fusion protein, and the target DNA in the cell come into contact, causing insertion of the donor sequence into the target region described in the knock-in method of the present invention above, making it possible to obtain a cell in which the donor sequence has been inserted into the target region.

[0093] The present invention also provides a method for producing a non-human individual containing a cell (knock-in cell) in which a donor sequence has been inserted into a target region of target DNA. This method includes the step of producing a non-human individual from a cell obtained by the above-described method for producing a knock-in cell of the present invention. Examples of the non-human individual include non-human animals and plants. Examples of the non-human animal include mammals (e.g., mice, rats, guinea pigs, hamsters, rabbits, monkeys, pigs, cows, goats, and sheep), fish, birds, reptiles, amphibians, and insects. When producing a model animal, the mammal is preferably a rodent such as a mouse, rat, guinea pig, or hamster, with mice being particularly preferred. Examples of the plant include grains, oilseed crops, forage crops, fruits, and vegetables. Specific examples of crops include rice, corn, banana, peanut, sunflower, tomato, rapeseed, tobacco, wheat, barley, potato, soybean, cotton, and carnation.

[0094] Known methods can be used to create non-human individuals from the knock-in cells. When creating non-human individuals from cells in animals, germ cells or pluripotent stem cells are typically used. For example, the donor DNA and nicking system are microinjected into oocytes, and the resulting oocytes are implanted into the uterus of a pseudopregnant female non-human mammal, after which offspring can be obtained. It has long been known that somatic cells of plants possess totipotency. For example, by microinjecting the donor DNA and nicking system into plant cells and regenerating a plant from the resulting plant cells, a plant with a desired nucleotide sequence inserted into its genome can be obtained. Furthermore, from the resulting non-human individuals, progeny or clones with a desired nucleotide sequence inserted into their genome can also be obtained.

[0095] Confirmation of whether or not the donor sequence has been knocked into the target region on the target DNA and determination of the genotype can be performed based on conventionally known methods, such as PCR, sequencing, Southern blotting, etc.

[0096] <Kit> The present invention also provides a kit for use in the above-mentioned knock-in method of the present invention, the method for producing a knock-in cell of the present invention, or the method for producing a non-human individual of the present invention, which kit comprises: at least one selected from the group consisting of a first fusion protein, a first fusion protein expression vector that expresses the first fusion protein, and a polynucleotide encoding the first fusion protein; and at least one selected from the group consisting of a second fusion protein, a second fusion protein expression vector that expresses the second fusion protein, and a polynucleotide encoding the second fusion protein.

[0097] The first fusion protein and the second fusion protein are as described above in the nicking system of the present invention, including preferred embodiments thereof. These may each independently be in the form of a protein, a polynucleotide encoding the protein, or a vector (expression vector) expressing the protein. The expression vectors for the first fusion protein and the second fusion protein, and the polynucleotides encoding them, are also as described above, including preferred embodiments thereof. The kit of the present invention may further include the donor DNA, as described above in the knock-in method of the present invention, including preferred embodiments thereof. However, the donor DNA to be used in each method may be designed and synthesized by the user as appropriate depending on the target region of the target DNA.

[0098] When the fusion protein of the present invention is in the form of an expression vector, it may be in a form that allows a user to design each fusion protein depending on a target region of a target DNA, and the first fusion protein expression vector can be at least one selected from the group consisting of: (i) a vector comprising a polynucleotide encoding an ND1 domain and a first DNA-binding domain, and (ii) a vector comprising a polynucleotide encoding an ND1 domain and an insertion site for a polynucleotide encoding the first DNA-binding domain, and the second fusion protein expression vector can be at least one selected from the group consisting of: (iii) a vector comprising a polynucleotide encoding a dND1 domain and a second DNA-binding domain, and (iv) a vector comprising a polynucleotide encoding a dND1 domain and an insertion site for a polynucleotide encoding the second DNA-binding domain.

[0099] In the vectors (i) to (iv) above, the order of each component (domain, linker if necessary, insertion site, etc.) can be adjusted appropriately depending on the first and second fusion proteins to be expressed. Furthermore, for example, when the first and second DNA-binding domains are TALEs, the polynucleotide inserted into the insertion site in (ii) or (iv) may consist solely of the TALE repeat sequence. In this case, the vectors (ii) and (iv) may contain a polynucleotide encoding the N-terminal domain of the TALE and, if necessary, a polynucleotide encoding the C-terminal domain. Furthermore, the first and second fusion protein expression vectors may be a single vector that expresses the first and second fusion proteins. Preferably, each of the vectors (i) to (iv) contains an expression unit that enables the expression of the respective polynucleotides.

[0100] The kits of the present invention may further include one or more additional reagents. Examples of such additional reagents include, but are not limited to, a dilution buffer, a reconstitution solution, a wash buffer, a nucleic acid introduction reagent, a protein introduction reagent, and a control reagent. The kits may also include instructions for use in carrying out each method of the present invention.

[0101] The components included in the kit of the present invention may be contained in separate containers or in the same container. Each component may be contained in a single-use amount in a container, or multiple doses may be contained in a single container. Each component may be contained in a container in a dry form, or in a form dissolved in an appropriate solvent (a solvent containing a buffer, stabilizer, preservative, antiseptic, etc.).

[0102] The present invention will be explained in more detail below based on test examples, but the present invention is not limited to the following test examples.

[0103] (Test Example 1) Confirmation of Nickase Activity by Single-Strand Annealing (SSA) Assay [Method] (1) Preparation of TALE Vector Set and TALE Expression Plasmid The sequence and structure of the TALE vector set were constructed according to Sakura et al., Sci. Rep., 3:3379 (2013) so as to correspond to each TALE recognition sequence. First, we modified (particularly the 4th and 32nd) module sequences (non-repeat-variable di-residue (non-RVD)) other than the sequences encoding the four types of variable residues (RVD: HD, NG, NI, NN) consisting of two amino acids at positions 12 and 13, and then synthesized artificial DNA sequences (module sequences) with BsAI restriction enzyme recognition sites added to both ends (16 types in total: 1HD-4HD, 1NG-4NG, 1NI-4NI, 1NN-4NN). These module sequences were then inserted into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan) to prepare a module plasmid set (16 types in total: pEX1HD to pEX4HD, pEX1NG to pEX4NG, pEX1NI to pEX4NI, and pEX1NN to pEX4NN).

[0104] In addition, the DNA sequences FUS2_aXX (total of 7 types: XX = 1a, 2a, 2b, 3a, 3b, 4a, 4b) and FUS2_b (1-4) (total of 4 types) constituting the array plasmid were prepared by artificial DNA synthesis. These were inserted into pCR8 / GW / TOPO (Thermo Fisher Scientific, Waltham, MA, USA) to prepare pCR8_FUS2_aXX and pCR8_FUS_b (1-4). This was then used as a capture vector for the first assembly step (step 1) in the Platinum Gate system described in Sakuma et al. (2013), and an array plasmid with a TALE sequence linked thereto was prepared according to the method described in the same document.

[0105] Next, a destination vector (TALE47 vector) was prepared by inserting one of the module sequences and a sequence encoding the C-terminal domain of the TALE following the 3' end of the N-terminal domain of the TALE, in addition to a sequence encoding the N-terminal domain of the TALE prepared by artificial DNA synthesis, into pcDNA3.1s (prepared by removing the drug resistance gene expression unit from pcDNA3.1(+) (Thermo Fisher Scientific). The C-terminal domain of the TALE used was a C-terminal domain with 47 amino acids (47), and the sequence encoding "47" was synthesized by artificial DNA synthesis with reference to the sequence of the C-terminal domain contained in Addgene's ptCMV-153 / 47-VR-NG (Addgene ID: 50704). Furthermore, the N-terminal domain of the TALE was synthesized artificially by referring to the sequence of the N-terminal domain contained in Addgene's ptCMV-153 / 47-VR-HD (ID: 50703) in correspondence with the C-terminal domain. Using the array plasmid and destination vector prepared above, a TALE expression plasmid was prepared by the Golden Gate method according to the method described in Sakuma et al. (2013).

[0106] (2) Preparation of TALE-ND1 Expression Plasmid and TALE-dND1 Expression Plasmid Using DNA encoding an artificially synthesized ND1 domain (nucleotide sequence number: 39, amino acid sequence number: 40) as a template, a fragment (ND1) was amplified by PCR using the primer set shown in Table 1 below, and using the TALE47 vector prepared in (1) above as a template, a fragment (TALE47) was amplified by PCR using the primer set shown in Table 1 below. These fragments were then ligated by the in-fusion method to prepare a TALE47-ND1 expression plasmid that expresses a fusion protein in which the ND1 domain is fused to the C-terminus of TALE47. Next, using the TALE47-ND1 expression plasmid as a template, fragments amplified by PCR using the two primer sets shown in Table 1 below were ligated by the in-fusion method to create a site-directed mutagenesis method to prepare a TALE47-dND1 expression plasmid that expresses a fusion protein in which a dND1 domain (nucleotide sequence number: 41, amino acid sequence number: 42) in which the D66N mutation was introduced into the ND1 domain was fused to the C-terminus of TALE47.

[0107]

[0108] TALE repeat domains (APC-L, APC-R, Rosa26-L, Rosa26-R) corresponding to the TALE recognition sequences listed in Table 2 below were inserted into TALE47 of the TALE47-ND1 expression plasmid and TALE47-dND1 expression plasmid prepared above, respectively, and subjected to the SSA assay described below.

[0109]

[0110] (3) Quantification of double-stranded DNA cleavage activity by SSA assay To prepare a reporter plasmid for the SSA assay, pcEGxxFP prepared in JP 2023-38023 A was used. As the target genes used in the SSA assay, APC (human adenomatous polyposis coli) was isolated from the human genome, and Rosa26 and HPRT1 were isolated from the hamster genome by PCR. These genes were then inserted into the BamHI / EcoRI restriction enzyme recognition sequence in a linker inserted between the N-terminus and C-terminus of EGFP in pcEGxxFP by in-fusion to prepare the respective reporter plasmids for the SSA assay (pcEGxAPCxFP, pcEGxRosa26xFP).

[0111] Also, in the same manner as in WO 2023-038023, a nickase-type nCas9 (D10A) expression plasmid (pX330 (D10A)_BS) and a guide RNA expression plasmid (pX330_BS-ΔCas9) were prepared, and each guide RNA was designed to correspond to the guide RNA target sequence on each target gene shown in Table 3 below (Table 3 shows the complementary sequence (non-underlined sequence) and PAM sequence (underlined sequence) of the guide RNA target sequence (on the antisense strand). Figure 1 shows the PAM sequence on the target genes APC and Rosa26 (the position corresponding to the PAM sequence on the guide RNA is shown in Figure 1), and a conceptual diagram of the positional relationship between the cleavage point by nCas9 (D10A), each TALE, and the guide RNA.

[0112]

[0113] HEK293T cells grown in DMEM medium containing 10% FBS were cultured at 1 × 10 4One cell per well was seeded into each well of a 96-well plate the day before transfection. 33 ng of TALE-L (APC-L, Rosa26-L) expression plasmid, 33 ng of TALE-R (APC-R, Rosa26-R) expression plasmid, and 33 ng of SSA assay reporter plasmids (pcEGxAPCxFP, pcEGxRosa26xFP) were transfected into HEK293T cells using Lipofectamine 3000 (Thermo Fisher Scientific). When the type or amount of transfected plasmid was reduced, pBluescript II (SK+) (Stratagene) was used to supplement the transfected plasmids, bringing the total amount of transfected plasmid to 100 ng.

[0114] After transfection, the cells were cultured for 48 hours, and the amount of fluorescence per unit area of ​​each well due to EGFP protein was measured using a plate reader. The above combination was also transfected with a combination of nCas9 (D10A) expression plasmid and guide RNA expression plasmid, and measurements were performed in the same manner as above. To reduce variations due to cell localization within the well, measurements were performed at 16 points per well, and the average value was used as the EGFP fluorescence intensity (RFI). The average and standard deviation of the values ​​from three independent wells were graphed.

[0115] The SSA assay reporter plasmid EGxAPCxFP, which is equipped with APC as the target gene, was used in combination with TALE47(APC-L)-ND1 and TALE47(APC-R)-ND1, or in combination with TALE47(APC-L)-dND1 and TALE47(APC-R)-ND1, where one ND1 was dND1, or in combination with TALE47(APC-L)-ND1 and TALE47(APC-R)-dND1, and the results when further treated with nCas9(D10A) and guide RNA (APC_gRNA-C or APC_gRNA-B) are shown in Figure 2.

[0116] Furthermore, the SSA assay reporter plasmid EGxRosa26xFP, which carries Rosa26 as the target gene, was used in combination with TALE47(Rosa26-L)-ND1 and TALE47(Rosa26-R)-ND1, or in combination with TALE47(Rosa26-L)-dND1 and TALE47(Rosa26-R)-ND1, where one ND1 is dND1, or in combination with TALE47(Rosa26-L)-ND1 and TALE47(Rosa26-R)-dND1, or in combination with nCas9(D10A) and guide RNA (Rosa26_gRNA-C or Rosa26_gRNA-B). The results are shown in Figure 3.

[0117] [Results] As shown in Figure 2, when a combination of TALE47(APC-L)-ND1 and TALE47(APC-R)-ND1 was applied to the reporter plasmid EGxAPCxFP, which carries APC as the target gene, double-stranded DNA cleavage activity was confirmed, reaffirming that the ND1 domain functions as a dimeric nuclease. On the other hand, when one of the ND1 domains was replaced with a dND1 domain, double-stranded DNA cleavage activity was confirmed to be lost. Furthermore, by simultaneously expressing nCas9 and a guide RNA corresponding to the guide RNA target sequence on the same strand as the TALE recognition sequence of TALE47-dND1 (APC_gRNA-C for TALE47(APC-L)-dND1, and APC_gRNA-B for TALE47(APC-R)-dND1), double-nicking double-stranded DNA cleavage activity was confirmed, confirming that the combination of the ND1 domain and the dND1 domain has nickase activity.

[0118] Furthermore, as shown in Figure 3, when a combination of TALE47(Rosa26-L)-ND1 and TALE47(Rosa26-R)-ND1 was applied to the reporter plasmid EGxRosa26xFP, which carries the Rosa26 gene as a target gene, double-stranded DNA cleavage activity was confirmed, reaffirming that the ND1 domain functions as a dimeric nuclease. Furthermore, even in this case, when one of the ND1 domains was replaced with a dND1 domain, double-stranded DNA cleavage activity was abolished. Furthermore, in this case as well, by simultaneously expressing nCas9 and a guide RNA corresponding to the guide RNA target sequence on the same strand as the TALE recognition sequence of TALE47-dND1 (Rosa26_gRNA-C for TALE47(Rosa26-L)-dND1, and Rosa26_gRNA-B for TALE47(Rosa26-R)-dND1), double-stranded DNA cleavage activity by double nicking was confirmed, confirming that the combination of the ND1 domain and the dND1 domain has nickase activity.

[0119] Furthermore, these results surprisingly confirmed that in the combination of the ND1 domain and the dND1 domain, a nick is introduced into the strand recognized by the TALE that binds to the dND1 domain.

[0120] (Test Example 2) Evaluation of knock-in efficiency using ND1-dND1 and donor lssDNA [Method] (1) Preparation of a plasmid for establishing a reporter cell line First, following the description of Nakajima et al., Genome Research, 28, 2018, pp. 223-230 (Non-Patent Document 4), mCherry-P2a-EGFP(C321G) / pcDNA3.1Zeo+ was prepared to establish a reporter cell line in which the expression unit of mCherry-P2a-EGFP(C321G) (nucleotide sequence number: 45, amino acid sequence number: 46) was inserted into the genome, resulting in a mutant EGFP protein that is unable to emit fluorescence, by substituting G for the 321st base of the EGFP gene to form a stop codon. The mCherry expression plasmid pmCherry-C1 was purchased from Takara Bio Inc., and the DNA encoding EGFP was fully synthesized. PCR was performed using these DNAs and the expression vector pcDNA3.1Zeo+ (manufactured by Thermo Fisher Scientific) using the primer sets listed in Table 4 below, and the resulting fragments were ligated by the in-fusion method to produce the plasmid mCherry-P2A-EGFP / pcDNA3.1Zeo+ (nucleotide sequence number: 43, amino acid sequence number: 44 for mCherry-P2A-EGFP).

[0121] Next, mCherry-P2A-EGFP(C321G) / pcDNA3.1Zeo+ was prepared as a mutant in which the 321st base C in the nucleotide sequence encoding EGFP on the plasmid mCherry-P2a-EGFP / pcDNA3.1Zeo+ was replaced with G. Using the plasmid mCherry-P2a-EGFP / pcDNA3.1Zeo+ as a template, PCR was performed using the two primer sets listed in Table 4 below, and the resulting fragments were linked by the In-Fusion method using site-directed mutagenesis to prepare the plasmid mCherry-P2A-EGFP(C321G) / pcDNA3.1Zeo+.

[0122]

[0123] (2) Preparation of donor plasmid and each expression plasmid Following the SNGD method described in Non-Patent Document 4, knock-in of the donor plasmid is induced by nicking with nCas9, and EGFP is returned to the wild type. In order to avoid a second attack by nCas9 after this, a guide RNA (sgEGFP_332s (Table 6 below shows the complementary sequence (non-underlined sequence) and PAM sequence (underlined sequence) of the guide RNA target sequence (on the antisense strand)) for inserting a nick on the target gene EGFP) was introduced into EGFP. A plasmid (donor plasmid) expressing donor PD-0 (nucleotide sequence number: 47) into which a mutation without amino acid substitution was introduced to suppress rebinding to the target sequence was prepared. That is, in order to clone the wild-type EGFP translation region sequence into pBluescript II (SK +), PCR was performed using the primer set listed in Table 5 below, using the total synthetic sequence of nucleotides encoding EGFP as a template, and ligated together with pBluescript II (SK +) cleaved with the restriction enzymes EcoRV and BamHI by the In-Fusion method to produce the plasmid EGFP / pBS. Furthermore, PCR was performed using the two primer sets listed in Table 5 below using EGFP / pBS as a template, and the resulting fragments were ligated by the In-Fusion method to produce the donor plasmid PD-0 / pBS.

[0124] Further, as the nCas9 expression plasmid used in the SNGD method, the nickase-type nCas9 (D10A) expression plasmid (pX330 (D10A)_BS) prepared in (3) of the above-mentioned Test Example 1 was used. Furthermore, as the guide RNA expression plasmid used in the SNGD method, the above-mentioned sgEGFP_332s and the guide RNA for causing nicking on the donor plasmid (pBS) (sgUC57N2 (Table 6 below shows the complementary sequence (sequence other than underlined) of the guide RNA target sequence (on the antisense strand) and the PAM sequence (underlined sequence)) were prepared in the same manner as in (3) of the above-mentioned Test Example 1, in which sgEGFP_332s and sgUC57N2 were inserted into the guide RNA expression plasmid (pX330_BS-ΔCas9).

[0125] Furthermore, to evaluate the knock-in efficiency when a combination of TALE47-dND1 and TALE47-ND1 was used instead of nCas9, in order to generate similar nicking near the nicking site by nCas9 and guide RNA (sgEGFP_332s), TALE repeat domains (EGFP-3, EGFP-3R) corresponding to each TALE recognition sequence on the EGFP gene listed in Table 7 below were inserted into TALE47 of the TALE47-dND1 expression plasmid and TALE47-ND1 expression plasmid prepared in Test Example 1 (2) above, to produce a TALE47(EGFP-3)-dND1 expression plasmid and a TALE47(EGFP-3R)-ND1 expression plasmid.

[0126] Furthermore, as a donor plasmid for the combination of TALE47-dND1 and TALE47-ND1, a plasmid (donor plasmid: PD-3 / pBS) expressing donor PD-3 (nucleotide sequence number: 48) in which a mutation without amino acid substitution that suppresses rebinding to the TALE recognition sequence was introduced into EGFP was prepared in the same manner as for the above-mentioned PD-0 / pBS using the two primer sets listed in Table 5 below.

[0127]

[0128]

[0129]

[0130] (3) Establishment of a reporter cell line The plasmid mCherry-P2A-EGFP(C321G) / pcDNA3.1Zeo+ prepared in (1) above was cleaved at the restriction enzyme ScaI cleavage site (present at only one site on the plasmid, on the ampicillin resistance gene), and the linear fragment was isolated and purified by agarose gel electrophoresis. 2.5 × 10 HEK293T cells grown in DMEM medium containing 10% FBS were cultured. 5Cells were seeded onto 6-well plates the day before transfection. 2.5 μg of the above mCherry-P2A-EGFP(C321G) / pcDNA3.1Zeo+ linear fragment was transfected into HEK293T cells using Lipofectamine 3000 (Thermo Fisher Scientific). Two days after transfection, the cells were detached, appropriately diluted with 125 μg / mL Zeocin-containing medium (hereafter referred to as the cell line established in this medium), and seeded onto 10 cm dishes. Approximately two weeks later, single colonies were collected and seeded onto 96-well plates. The 24 strains in which mCherry fluorescence was observed were expanded and subjected to knock-in efficiency evaluation according to the SNGD method described below in (5). Three strains (#7, #8, and #22) with high knock-in efficiency were selected, and single clones were again obtained from these three strains by limiting dilution. The resulting clones #7-1, #7-2, #7-3, #8-1, #8-2, #8-3, #22-1, #22-2, and #22-3 were compared in terms of knock-in efficiency, cell proliferation, etc., and #7-2, #8-1, and #22-3 were selected as strains with equivalent activity and proliferation capacity. For subsequent experiments, #22-3 was used as a reporter cell line (a strain expressing a fluorescent activity-loss mutant) in which the mCherry-P2A-EGFP (C321G) expression unit was inserted into the genome.

[0131] (4) Preparation of lss (long single stranded) DNA As donor lssDNA, Inoue et al., Cells, 2021, 10, 1076. https: / / doi.org / 10.3390 / cells10051076 (Non-Patent Document 3) lssDNA (total length: 720 nt, 5'HA: 320 nt, 3'HA: 399 nt, the same nucleotide sequence as PD-0 or PD-3) was prepared according to the method described therein. First, PCR was performed using the donor plasmid PD-0 / pBS or PD-3 / pBS as a template with the primer set listed in Table 8 below. At this time, one of the primers used was a primer with a phosphorylated 5' end. The resulting PCR product was purified using NucleoSpin Gel and PCR Clean-up (MACHEREY-NAGEL), and then approximately 5 μg was digested from the 5′ phosphorylated end with Lambda exonuclease (New England BioLab). The remaining double-stranded DNA was then digested with Exonuclease III (Takara Bio) and further purified using the above-mentioned NucleoSpin Gel and PCR Clean-up to obtain each lssDNA (PD-0, PD-3).

[0132]

[0133] (5) Evaluation of knock-in efficiency HEK293T cells grown in DMEM medium containing 10% FBS were cultured at 1 × 10 4Each cell was seeded into each well of a poly-L-lysine-coated or collagen-coated 96-well plate the day before transfection. A total of 100 ng of the donor plasmids (PD-0, PD-3), lssDNA (PD-0, PD-3), guide RNA (sgEGFP_332s) expression plasmid, guide RNA (sgUC57N2) expression plasmid, nCas9 expression plasmid, TALE47(EGFP-3)-dND1 expression plasmid, and TALE47(EGFP-3R)-ND1 expression plasmid was transfected into HEK293T cells (#22-3 strain) using Lipofectamine 3000 (Thermo Fisher Scientific) in the combinations and amounts shown in Table 9 below. When the type or amount of introduced plasmid was reduced, pBluescript II (SK+) (Stratagene) was used to supplement the amount, so that the total amount of introduced plasmid and DNA was 100 ng.

[0134]

[0135] After transfection, the cells were cultured for 72 hours and stained with Hoechst 33342 (10 μL of a 50 μg / mL solution was added to each well, 10 minutes, room temperature). Furthermore, 100 μL of 4% paraformaldehyde solution was added and treated at room temperature for 1 hour to fix the cells. After washing with PBS(-), 100 μL of PBS(-) was added to each well, and two types of fluorescence image data (fluorescence from Hoechst 33342 staining, EGFP fluorescence) were acquired from four locations per well using an In Cell Analyzer 6000 (manufactured by Cytiva). The number of fluorescent cells was quantified using image analysis software ImageJ, and the average value and standard deviation of the four locations for the number of EGFP-positive cells / total cell number (%) were graphed. The higher the average value of the obtained number of EGFP-positive cells / total cell number (%), the higher the knock-in efficiency.

[0136] Fluorescence images of EGFP for each of the combinations shown in Table 9 ((a): Mock, (b) SNGD, (c) dND1+ND1+PD, (d) dND1+ND1+lssDNA) are shown in Figure 4, and the number of EGFP-positive cells / total number of cells (%) is shown in Figure 5.

[0137] Furthermore, to compare and verify the case where nCas9 was combined with lssDNA instead of the donor plasmid as the donor DNA, the average number of EGFP-positive cells / total number of cells (%) was calculated in the same manner as above, except that the combinations and amounts of each donor DNA and expression plasmid were as shown in Table 10 below, and the relative activity and standard deviation were graphed, assuming the value (a) when the donor plasmid was used as 1. The results are shown in Figure 6.

[0138]

[0139] [Results] Figures 4 to 5 (b) show the results of confirming the knock-in efficiency by nicking the target genome and donor plasmid using nCas9 on an EGFP fluorescence activity loss mutant expression strain (#22-3 strain) using the SNGD method (SNGD: combination of single nicks in the target gene and donor plasmid) described in Non-Patent Document 4. As shown in Figures 4 to 5, as with the SNGD method described in Non-Patent Document 4, when the donor plasmid PD-0 / pBS, guide RNA (sgEGFP_332s) expression plasmid, and guide RNA (sgUC57N2) expression plasmid, and nCas9 expression plasmid were introduced (b), knock-in was observed in approximately 6% of all cells.

[0140] Figures 4 to 5 (c) show the results of confirming the knock-in efficiency when nicking of the target genome is carried out by a combination of TALE47-dND1 and TALE47-ND1 instead of nCas9. As shown in Figures 4 to 5, when the donor plasmid PD-3 / pBS, guide RNA (sgEGFP_332s) expression plasmid, nCas9 expression plasmid, TALE47 (EGFP-3)-dND1 expression plasmid, and TALE47 (EGFP-3R)-ND1 expression plasmid were introduced (c), knock-in was observed in approximately 6% of all cells, as in (b), and it was confirmed that the nickase activity of the combination of the ND1 domain and dND1 domain of the present invention can replace the nickase activity of nCas9.

[0141] Figures 4 to 5 (d) show the results of confirming the knock-in efficiency when nicking using a combination of the ND1 domain and the dND1 domain was performed using lssDNA instead of a donor plasmid as the donor DNA. As shown in Figures 4 to 5, when lssDNA (PD-3), a TALE47(EGFP-3)-dND1 expression plasmid, and a TALE47(EGFP-3R)-ND1 expression plasmid were introduced ((d): Example), knock-in was observed in approximately 18% of all cells, demonstrating a significantly higher knock-in efficiency, approximately three times higher than when a donor plasmid was used (c).

[0142] On the other hand, as shown in Figure 6, a comparative study was also conducted on nicking by nCas9 when lssDNA was combined in place of the donor plasmid (comparative example). However, in nicking by nCas9, even when the donor DNA was lssDNA (b), or even when the amount of lssDNA was further increased ((c) to (d)), no improvement in knock-in efficiency (number of EGFP-positive cells / total number of cells (%)) was observed compared to when a donor plasmid was used (a). Thus, no significant enhancement in knock-in efficiency was observed, as seen in the combination of the TALE-ND1 of the present invention, TALE-dND1, and lssDNA.

[0143] As described above, the nicking system of the present invention is a site-specific nicking system that comprises a first fusion protein comprising a first DNA-binding domain and an ND1 domain, and a second fusion protein comprising a second DNA-binding domain and a dND1 domain, and thereby exhibits nickase activity that can replace the nickase activity of nickase-type Cas9 (nCas9) using only the proteins. Furthermore, when using such a nicking system in knock-in technology, the knock-in efficiency can be significantly improved by combining it with, in particular, single-stranded DNA as the knock-in donor DNA.

[0144] Thus, the present invention provides a method for specifically and highly efficiently inserting a desired nucleotide sequence into a target region (knock-in) using only a protein, a method for producing cells in which a desired nucleotide sequence has been inserted into a target region (knock-in cells), and kits for use in these methods. Therefore, the present invention is expected to be utilized as an excellent genome editing tool in a wide range of industrial fields, including medicine, agriculture, and industry.

Claims

1. A method for inserting a donor sequence into a target region on a target DNA, comprising the step of contacting the target DNA with: donor DNA, which is a single-stranded DNA comprising a nucleotide sequence arranged in the following order from the 5' side: 5' homology arm sequence, donor sequence, 3' homology arm sequence; and a site-specific nicking system comprising: a first fusion protein comprising a first DNA-binding domain and an ND1 domain, and a second fusion protein comprising a second DNA-binding domain and a dND1 domain, wherein the ND1 domain and the dND1 domain form a dimer to introduce a nick within or near the target region, wherein the ND1 domain is one of the following (a) to (c): (a) a polypeptide comprising the amino acid sequence of SEQ ID NO: 40; or (b) a polypeptide comprising an amino acid sequence in which one or more amino acids in the amino acid sequence of SEQ ID NO: 40 have been substituted, deleted, inserted and / or added, and wherein the amino acids corresponding to positions 66 and 83 of the amino acid sequence of SEQ ID NO: 40 are aspartic acid. (c) an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and comprising an amino acid sequence in which the amino acids corresponding to the 66th and 83rd amino acids of the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid, and the dND1 domain is at least one polypeptide selected from the group consisting of (da) to (dc) below: (da) a polypeptide comprising an amino acid sequence in which the aspartic acid at the 66th position and / or the aspartic acid at the 83rd position of the amino acid sequence set forth in SEQ ID NO: 40 are substituted with any other amino acid; (db) a polypeptide comprising an amino acid sequence in which one or more amino acids are substituted, deleted, inserted and / or added in the amino acid sequence set forth in SEQ ID NO: 40, and wherein the amino acid corresponding to the aspartic acid at the 66th position and / or the aspartic acid at the 83rd position of the amino acid sequence set forth in SEQ ID NO: 40 are substituted with any amino acid other than aspartic acid.(dc) a polypeptide having an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and comprising an amino acid sequence in which the amino acid corresponding to the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any amino acid other than aspartic acid.

2. The method according to claim 1, wherein the donor DNA is a long single-stranded DNA, the lengths of which of the 5' homology arm sequence and the 3' homology arm sequence are independently 100 bases or more.

3. A method for producing a cell in which a donor sequence has been inserted into a target region on a target DNA, the method comprising the steps of introducing or expressing into a cell and contacting with the target DNA: donor DNA, which is a single-stranded DNA comprising a nucleotide sequence arranged in the following order from the 5' side: 5' homology arm sequence, donor sequence, 3' homology arm sequence; and a site-specific nicking system comprising: a first fusion protein comprising a first DNA-binding domain and an ND1 domain; and a second fusion protein comprising a second DNA-binding domain and a dND1 domain, wherein the ND1 domain and the dND1 domain form a dimer to introduce a nick within or near the target region; and the ND1 domain is selected from the group consisting of the following (a) to (c): (a) a polypeptide comprising the amino acid sequence of SEQ ID NO: 40 (b) a polypeptide comprising an amino acid sequence in which one or more amino acids have been substituted, deleted, inserted and / or added in the amino acid sequence set forth in SEQ ID NO: 40, and in which the amino acids corresponding to the 66th and 83rd amino acids in the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid; (c) a polypeptide having 90% or more homology to the amino acid sequence set forth in SEQ ID NO: 40, and comprising an amino acid sequence in which the amino acids corresponding to the 66th and 83rd amino acids in the amino acid sequence set forth in SEQ ID NO: 40 are aspartic acid; and wherein the dND1 domain is one of the following (da) to (dc): (da) a polypeptide comprising an amino acid sequence in which the aspartic acid at position 66 and / or the aspartic acid at position 83 in the amino acid sequence set forth in SEQ ID NO: 40 are substituted with any other amino acid. (db) A polypeptide comprising an amino acid sequence in which one or more amino acids in the amino acid sequence set forth in SEQ ID NO: 40 have been substituted, deleted, inserted, and / or added, and in which the amino acids corresponding to the aspartic acid at position 66 and / or the aspartic acid at position 83 in the amino acid sequence set forth in SEQ ID NO: 40 have been substituted with any amino acid other than aspartic acid.(dc) a polypeptide having an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 40, and comprising an amino acid sequence in which the amino acid corresponding to the aspartic acid at position 66 and / or the aspartic acid at position 83 of the amino acid sequence set forth in SEQ ID NO: 40 is substituted with any amino acid other than aspartic acid.

4. The method according to claim 3, wherein the donor DNA is a long single-stranded DNA, the lengths of which of the 5' homology arm sequence and the 3' homology arm sequence are independently 100 bases or more.

5. A kit for use in the method of any one of claims 1 to 4, comprising at least one selected from the group consisting of a first fusion protein, a first fusion protein expression vector expressing the first fusion protein, and a polynucleotide encoding the first fusion protein; and at least one selected from the group consisting of a second fusion protein, a second fusion protein expression vector expressing the second fusion protein, and a polynucleotide encoding the second fusion protein, wherein the first fusion protein expression vector is at least one selected from the group consisting of (i) a vector comprising a polynucleotide encoding an ND1 domain and a first DNA-binding domain, and (ii) a vector comprising a polynucleotide encoding the ND1 domain and an insertion site for the polynucleotide encoding the first DNA-binding domain; and the second fusion protein expression vector is at least one selected from the group consisting of (iii) a vector comprising a polynucleotide encoding a dND1 domain and a second DNA-binding domain, and (iv) a vector comprising a polynucleotide encoding the dND1 domain and an insertion site for the polynucleotide encoding the second DNA-binding domain.

Citation Information

Patent Citations

  • Nuclease domain forming double-stranded breaks in DNA

    JP2023038023A

  • Genome editing method

    WO2018097257A1

  • Method for modifying target site in double-stranded DNA in cell

    WO2019189147A1

  • Novel nuclease domain and uses thereof

    WO2020045281A1

  • Method for introducing protein in nuclei of plant cells

    WO2020075399A1