Nuclease domains for double-stranded cleavage of DNA

By using the two ND1 nuclease domains linked via the linker as DNA cleavage enzymes, the problem of low double-stranded cleavage DNA activity in the prior art is solved, and efficient genome editing is achieved.

CN120187848APending Publication Date: 2025-06-20SUMITOMO CHEM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280102133.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the nuclease domain activity of double-stranded cleaved DNA is low and it is difficult to effectively perform genome editing in human cells.

Method used

Double-stranded cleavage of DNA is achieved by using the two ND1 nuclease domains linked via the linker as the nuclease domains of the site-specific DNA cleavage enzyme. Adjust the type and chain length of the linker sequence to improve double-strand cleavage activity.

Benefits of technology

Simple and effective DNA double-strand cleavage is achieved, improving the efficiency of genome editing and avoiding the limitation of the nuclease domain requiring two molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present disclosure has found that when two ND1 linked via a linker are used as nuclease domains of a site-specific DNA cleavage enzyme, a target site on DNA is efficiently double-stranded cleavage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a nuclease domain for double-strand cleavage of DNA, and more particularly, to a fusion nuclease domain containing two nuclease domains 1 (ND1) bound via a linker and its applications. Background Art

[0002] The double-strand cleavage of DNA by TALEN, which was developed as a second-generation genome editing technology in 2010 and serves as an artificial restriction enzyme, requires the formation of a dimer of the non-specific DNA cleavage domain of FokI, which is a type IIS restriction enzyme. In the TALEN technology, a pair of paired TALE repeats is required. Since a set of TALENs jointly recognizes a target sequence of 30 to 40 bases and cleaves it, the cleavage specificity is extremely high, and it is an excellent genome editing technology that strongly suppresses off-target occurrence. However, due to the difficulty of simultaneously expressing two molecules of TALEN, which is a huge molecule of 100 kDa or more, in cells, or introducing an equal amount of TALEN protein into cells, and the labor required for its preparation, CRISPR-Cas9, which was published in 2012 in place of TALEN, has rapidly spread.

[0003] On the other hand, research on the improvement of TALEN molecules has also continued, starting with the development of PlatinumTALE, which is a high-activity type TALE, and various derivative technologies have emerged. Among them, there are reports of linking two FokI nuclease domains (scFokI), binding them to zinc fingers or TALEs, and inducing double-strand cleavage (Non-Patent Documents 1 and 2). By this technology, it is not necessary to apply a set of TALEs in the cleavage of the target sequence by TALEN. However, the cleavage activity of scFokI is low, and effective genome editing in human cells is difficult, and further improvement is required.

[0004] In addition, the present inventors have developed two new nuclease domains (nuclease domain 1; ND1 and nuclease domain 2; ND2), which are different from the conventional FokI nuclease domain. By fusing these nuclease domains with ZF or TALE and using them as a set of ZFNs or TALENs, success has been achieved in genome editing at the target site (Patent Document 1).

[0005] Prior Art Documents

[0006] Patent Documents

[0007] Patent Document 1: International Publication No. 2020 / 045281

[0008] Non-Patent Documents

[0009] Non-Patent Document 1: Minczuk et al (2008) Nucleic Acids Research, 36(12), 3926-3938

[0010] Non-Patent Document 2: Mino et al (2009) Journal of Biotechnology, 140(3-4), 156-161 Summary of the Invention

[0011] Problems to be Solved by the Invention

[0012] The present invention has been completed in view of such circumstances, and an object thereof is to provide a nuclease domain capable of simply and effectively double-strand cutting DNA.

[0013] Means for Solving the Problems

[0014] The inventors of the present invention conducted in-depth research to solve the above problems, and as a result, found that when two ND1s linked via a linker are used as the nuclease domain of a site-specific DNA cleavage enzyme (as an example, TALEN), the target site on DNA is effectively double-strand cut. On the other hand, in the case of applying two ND2s linked via a linker as the nuclease domain and in the case of applying two FokI nuclease domains linked via a linker, double-strand cleavage such as that of ND1 was not observed. Based on these facts, it was determined that the excellent double-strand cleavage activity of the linked ND1 on DNA is not an event common to nuclease domains, but a unique action of ND1.

[0015] In addition, the inventors of the present invention conducted research on the linker sequence between ND1 and ND1, and as a result, found that by adjusting the type and chain length of the linker sequence, the double-strand cleavage activity can be further improved, thus completing the present invention.

[0016] Therefore, the present invention relates to a fusion nuclease domain comprising two ND1s bound via a linker and its applications. More specifically, the following are provided.

[0017] (1) A fusion nuclease domain comprising two nuclease domains 1 bound via a linker.

[0018] (2) A site-specific DNA cleavage enzyme comprising a DNA binding domain and the fusion nuclease domain described in (1).

[0019] (3) A nucleic acid encoding the fusion nuclease domain described in (1) or the site-specific DNA cleavage enzyme described in (2).

[0020] (4) A vector comprising the nucleic acid described in (3) or its translation product.

[0021] (5) A method for preparing a cell with modified target DNA, which comprises the step of introducing a molecule selected from the following (a) to (c) into the cell:

[0022] (a) The site-specific DNA cleavage enzyme described in (2)

[0023] (b) A nucleic acid encoding the site-specific DNA cleavage enzyme of (a)

[0024] (c) A vector containing the nucleic acid described in (b) or its translation product.

[0025] (6) A kit for modifying target DNA, which contains a molecule selected from the following (a) to (c):

[0026] (a) The site-specific DNA cleavage enzyme described in (2)

[0027] (b) A nucleic acid encoding the site-specific DNA cleavage enzyme of (a)

[0028] (c) A vector containing the nucleic acid described in (b) or its translation product.

[0029] Advantages of the Invention

[0030] According to the present invention, DNA can be double-strand cleaved by the nuclease domain of one molecule, and it is not necessary to apply the nuclease domains of two molecules for double-strand cleavage of DNA as in the prior art. Therefore, in site-specific DNA cleavage enzymes such as TALEN and ZFN, if the nuclease domain of the present invention is used, site-specific DNA editing can be carried out simply and effectively. Brief Description of the Drawings

[0031] Figure 1 A graph showing the results of detecting the DNA cleavage activity of TALEN (referred to as "TALE-scFokI") containing two FokIs bound via a linker by the SSA assay method. In addition, "95" in the graph represents the HTS95 linker of 95 amino acid residues, and "60" represents the GGGGS×12 linker of 60 amino acid residues (the same hereinafter).

[0032] Figure 2 A graph showing the structures of TALENs (referred to as "TALE-scND1" and "TALE-scND2" respectively) containing two nuclease domains (ND1 or ND2) bound via a linker. As an example of TALE (DNA binding domain), the case of applying TALE63 is shown. In addition, "120" in the graph represents the GGGGS×24 linker of 120 amino acid residues, and "180" represents the GGGGS×36 linker of 180 amino acid residues (the same hereinafter). ​​

[0033] Figure 3 Display application Figure 2 The figure showing the results of detecting the DNA cleavage activity of the shown TALE-scND1 and TALE-scND2 by the SSA assay method.

[0034] Figure 4 The figure showing the results of evaluating the effects of the structure of TALE and the chain length of the linker on the DNA cleavage activity of TALE-scND1 by the SSA assay method. As TALE, TALE63 and TALE47 were applied.

[0035] Figure 5 The figure showing the results of evaluating the effects of the type of target gene and the chain length of the linker on the DNA cleavage activity of TALE-scND1 by the SSA assay method. As the target genes, (A) APC gene, and (B) Rosa26 gene and HPRT1 gene were applied. As TALE, TALE47 was applied.

[0036] Figure 6 The figure showing the results of evaluating the effect of the elongation of the C-terminal domain of TALE on the DNA cleavage activity of TALE-scND1 by the SSA assay method. As the target gene, Rosa26 gene was applied. As TALE, TALE47 was applied. As the linker binding two ND1s, the HTS95 linker of 95 amino acid residues was applied.

[0037] Figure 7 The figure showing the results of evaluating the effect of the shortening of the C-terminal domain of TALE on the DNA cleavage activity of TALE-scND1 by the SSA assay method. As the target genes, Rosa26 gene, APC gene, and HPRT1 gene were applied. As TALE, TALE47 was applied. As the linker binding two ND1s, the HTS95 linker of 95 amino acid residues was applied.

[0038] Figure 8 Show Figure 9 The figure of the structure of TALE-scND1 applied in the experiment of

[0039] Figure 9 The figure showing the results of evaluating the effect of the type of linker on the DNA cleavage activity of TALE-scND1 by the SSA assay method. As the linkers, HTS95 (95 amino acid residues), GSS×32 (96 amino acid residues), SAGG×24 (96 amino acid residues), and GGGGS×19 (95 amino acid residues) were applied. As the target genes, (A) Rosa26 gene, (B) APC gene, and (C) HPRT1 gene were applied. As TALE, TALE24 was applied.​​​​​​​ Detailed implementation mode

[0040] The present invention provides a fusion nuclease domain (scND1) containing two ND1s joined via a linker.

[0041] "ND1" in the present invention is one of the nuclease domains discovered by the present inventors by screening for homologous sequences having an identity with the FokI nuclease domain in the range of 35% to 70% (Patent Document 1).

[0042] The amino acid sequence of the full-length protein containing ND1 (as a representative example, derived from Bacillus sp. SGD-V-76) is shown in SEQ ID NO: 97. ND1 typically corresponds to the partial peptide at positions 391 to 585 of SEQ ID NO: 97 and has 70% identity with the amino acid sequence of the FokI nuclease domain.

[0043] "ND1" in the present invention, when two molecules are joined via a linker, as long as it has DNA double-strand cleavage activity, includes a nuclease domain composed of an amino acid sequence having a high identity with the amino acid sequence corresponding to positions 391 to 585 of SEQ ID NO: 97. Such nuclease domains include, for example, ND1 from other bacterial sources, mutants of ND1 (natural mutants and artificial mutants).

[0044] "High identity" herein means an identity of 85% or more, preferably 90% or more (for example, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more).

[0045] The identity of the amino acid sequences in the present invention is determined by comparing two sequences aligned in a state of maximum sequence consistency. The method for obtaining the numerical value (%) of sequence identity is well known to those skilled in the art. As an algorithm for obtaining the optimal alignment and sequence identity, any algorithm known to those skilled in the art (for example, BLAST algorithm, FASTA algorithm, etc.) can be used. The sequence identity of amino acid sequences is determined using sequence analysis software such as BLASTP, FASTA, etc.

[0046] As mutants of ND1, for example, nuclease domains composed of amino acid sequences in which 1 to several amino acid residues in the amino acid sequence shown by positions 391 to 585 of SEQ ID NO: 97 are substituted, deleted, inserted, or added are listed. Here, "1 to several" is, for example, 1 to 30, preferably 1 to 20 (for example, 1 to 10, 1 to 5, 1 to 3).

[0047] As a "linker" for binding two ND1s, as long as the fusion nuclease domain of the present invention can exhibit double-strand cleavage activity, there are no particular limitations on its chain length and type. The chain length of the linker is usually 30 to 300 amino acid residues, preferably 50 to 200 amino acid residues. As types of linkers, for example, HTS95 linker, GSS linker, SAGG linker, GGGGS linker, HTEN linker, etc. are listed, and the HTS95 linker is preferred. In the case of applying these linkers, multiple linkers can be connected and applied according to the required chain length corresponding to the linker.

[0048] The binding of two ND1s to the linker can be carried out at the nucleic acid level and the amino acid level. That is, by ligating the respective nucleic acids in the order of "nucleic acid encoding ND1 → nucleic acid encoding linker → nucleic acid encoding ND1" through a ligation reaction, a nucleic acid encoding the fusion nuclease domain of the present invention can be prepared. If the nucleic acid thus obtained is inserted into an expression vector and expressed in an appropriate host cell, a fusion nuclease domain that is ligated at the amino acid level in the order of "ND1 → linker → ND1" is formed. Regarding the host cell and the vector, it is the same as in the case of the site-specific DNA cleavage enzyme of the present invention (described later). In addition, the fusion nuclease domain can also be prepared by artificial synthesis based on the amino acid sequence information.

[0049] As described later, whether the fusion nuclease domain has double-strand cleavage activity of DNA can be evaluated by whether the site-specific DNA cleavage enzyme obtained by fusing with the DNA-binding domain exhibits double-strand cleavage activity of site-specific DNA. For example, as described in the examples of the present application, it is only necessary to evaluate whether the site-specific DNA cleavage enzyme obtained by fusing with TALE targeting a specific site of a specific gene exhibits double-strand cleavage activity at that specific site. In this evaluation, an SSA test system targeting a reporter gene on a plasmid (for example, an EGFP gene inserted with a TALE recognition sequence) can be applied, and the evaluation can be carried out using the reporter activity as an index. In addition, a site-specific DNA cleavage enzyme targeting an endogenous gene in a cell can be prepared, and the evaluation can be carried out using the change (mutation) of the base sequence of the target site of this site-specific DNA cleavage enzyme as an index.

[0050] In addition, the present invention provides a site-specific DNA cleavage enzyme comprising a DNA-binding domain and the above-mentioned fusion nuclease domain. The site-specific DNA cleavage enzyme of the present invention binds to a target sequence on DNA via the DNA-binding domain, and double-strand cleaves the DNA at the target cleavage site through the fusion nuclease domain.

[0051] From the perspective of the most effective performance of its characteristics, the "target DNA" of the site-specific DNA cleavage enzyme of the present invention is preferably double-stranded DNA. Examples of double-stranded DNA include, but are not limited to, eukaryotic nuclear genomic DNA, mitochondrial DNA, plastid DNA, prokaryotic genomic DNA, phage DNA, or plasmid DNA.

[0052] The DNA binding domain may be any protein domain that can specifically bind to an arbitrary DNA sequence (target sequence), and there is no particular limitation. For example, TALE (Transcription Activator-Like Effector), zinc finger, PPR (Pentatricopeptide repeat), CRISPR / Cas (a complex of Cas protein and guide RNA), etc. are listed.

[0053] TALE is a protein secreted by bacteria of the genus Xanthomonas and activates the gene transcription of host plants. The TALE used in the present invention may include an N-terminal domain and / or a C-terminal domain in addition to the TALE repeat domain.

[0054] The TALE repeat domain is composed of multiple, for example, 10 to 30, preferably 13 to 25, more preferably 15 to 20 tandem repeats (TALE repeats) of TALE sequences that form a right-handed superhelix. A typical TALE repeat unit (one TALE sequence) consists of 33 to 35 amino acids, and recognizes a specific base of DNA through a variable residue (Repeat Variable Diresidue: RVD) composed of two amino acid residues at its 12th and 13th positions. Examples of RVDs that specifically recognize bases include HD that recognizes C, NG that recognizes T, NI that recognizes A, NN that recognizes G or A, and NS that recognizes A, C, G, or T. Based on the DNA recognition mechanism of this TALE repeat domain, a TALE that can recognize and bind to a desired base sequence (TALE recognition sequence) on DNA can be prepared by artificially linking TALE sequences that recognize specific bases.

[0055] As the TALE, the TALE mode of Platinum TALEN (Japanese Unexamined Patent Application Publication No. 2015-33365), in which the amino acids at two specific positions of one TALE repeat unit change every four TALE repeat units, is preferred.

[0056] As the N-terminal domain of TALE, a natural amino acid sequence can be applied (for example, the amino acid sequence of the N-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185 to 32188)), but as long as it does not have a negative impact on the function of the site-specific DNA cleavage enzyme of the present invention (double-stranded cleavage of the DNA at the target cleavage site), the aforementioned natural amino acid sequence can also be appropriately modified. As specific examples of the N-terminal domain of TALE, for example, the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-136 / 63-VR-HD (ID: 50699), the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-153 / 47-VR-HD (ID: 50703) are listed, but not limited thereto.

[0057] As the C-terminal domain of TALE, a natural amino acid sequence can be applied (for example, the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185 to 32188)), but as long as it does not have a negative impact on the function of the site-specific DNA cleavage enzyme of the present invention (double-stranded cleavage of the DNA at the target cleavage site), the aforementioned natural amino acid sequence can also be appropriately modified. As specific examples of the C-terminal domain, for example, the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185 to 32188) (WT: number of amino acids = 180), the amino acid sequence of the C-terminal domain contained in Addgene's pTALEN_v2 (ID: 32189 to 32192) (number of amino acids = 63), the amino acid sequence of the C-terminal domain contained in Addgene's ptCMV-153 / 47-VR-NG (ID: 50704) (number of amino acids = 47) are listed, but not limited thereto. In fact, in this example, in the case of applying TALE without a C-terminal domain and in the case of applying TALE with a long-chain C-terminal domain of 180 amino acid residues, site-specific double-stranded cleavage activity of DNA was observed ( Figure 6 , 7 ).

[0058] The DNA-binding domain containing zinc fingers is composed of a zinc finger array (hereinafter, also referred to as "ZFA"), and the aforementioned zinc finger array is composed of a plurality of, for example, 3 to 9 zinc fingers. It is known that one zinc finger recognizes a 3-base sequence. For example, zinc fingers that recognize GNN, ANN, or CNN, etc. are known.

[0059] The DNA-binding domain containing PPR is composed of multiple, for example, 10 to 30 tandem repeats (PPR repeats). It is known that one unit of a typical PPR repeat consists of 35 amino acids, and one unit of the PPR repeat recognizes one base. In addition, it is known that TALE specifies the binding base with two amino acids present in one unit of the repeat, whereas PPR specifies the binding base with three amino acids.

[0060] CRISPR-Cas is a complex containing a Cas protein and a guide RNA. For example, CRISPR-Cas9, CRISPR-Cpf1 (Cas12a), CRISPR-Cas12b, CRISPR-CasX (Cas12e), CRISPR-Cas14, CRISPR-Cas3, etc. are listed. The CRISPR-Cas applied in the present invention is preferably CRISPR-Cas composed of a Cas (for example, dCas) in which nuclease activity is inactivated. Cas mutants with inactivated nuclease activity are well known. When the fusion nuclease domain of the present invention is applied to CRISPR-Cas, generally, it binds to the protein constituting CRISPR-Cas. In the case where CRISPR-Cas is a class I system composed of multiple proteins (subunits), the fusion nuclease domain of the present invention can be fused with a subunit without nuclease activity (for example, the subunit constituting Cascade in CRISPR-Cas3). In this case, the subunit with nuclease activity (for example, Cas3 in CRISPR-Cas3) can be excluded from the components of the CRISPR-Cas system.

[0061] In the site-specific DNA-cleaving enzyme of the present invention, the fusion nuclease domain and the DNA-binding domain can bind directly or can bind via a linker. In the case where a linker is present, as long as the fusion nuclease domain of the present invention can exhibit double-strand cleavage activity, its chain length and type are not particularly limited. The chain length of the linker is usually 2 to 180 amino acid residues, preferably 2 to 120 amino acid residues. As types of the linker, for example, HTS95 linker, GSS linker, SAGG linker, GGGGS linker, HTEN linker are listed, but are not limited to these. In the case of applying these linkers, multiple can be connected and applied corresponding to the required chain length of the linker.

[0062] The nuclease domain of the fusion nuclease can bind to the N-terminal side or the C-terminal side of the DNA-binding domain. In addition, DNA-binding domains can be arranged at both ends of the fusion nuclease domain to form a sandwich type (for the sandwich type, refer to Mori et al (2009) Biochemical and Biophysical Research Communications, 390(3), 694-697). Further, flag tags for purification and detection, and various translocation signals (such as nuclear translocation signals, mitochondrial translocation signals, plastid translocation signals, etc.) can be added.

[0063] The target sequence recognized by the DNA-binding domain of the site-specific DNA nuclease of the present invention is any sequence on DNA. In the conventional ZFNs, TALENs, etc. applied in a set (bimolecular), the target sequence needs to be two sequences sandwiching a spacer sequence, but the site-specific DNA nuclease of the present invention can be applied as a single molecule, which is also advantageous in terms of less restriction in targeting.

[0064] The site-specific DNA nuclease of the present invention can be prepared by methods well-known to those skilled in the art. For example, methods of inserting the nucleic acid encoding the site-specific DNA nuclease of the present invention into an appropriate expression vector, introducing the expression vector into a host cell, and expressing the site-specific DNA nuclease in the host cell, and methods of artificial synthesis based on the amino acid sequence information of the site-specific DNA nuclease can be listed. As the expression vector, it can be appropriately selected from the vectors used in this field. For example, plasmid vectors, viral vectors, phage vectors, phagemid vectors, BAC vectors, YAC vectors, MAC vectors, HAC vectors, etc. can be listed. As the host cell for introducing the expression vector, it can be appropriately selected considering the suitability with the expression vector. For example, cells of prokaryotes such as Escherichia coli, Actinomycetes, Archaea, and eukaryotes such as yeast, sea urchin, silkworm, zebrafish, mouse, rat, frog, tobacco, Arabidopsis thaliana, rice, etc. can be listed.

[0065] The present invention provides a nucleic acid encoding the above-mentioned fusion nuclease domain and a nucleic acid encoding the above-mentioned site-specific DNA nuclease. The "nucleic acid" here includes either DNA or RNA (mRNA). In the above nucleic acid, codon optimization can be performed for the purpose of improving the expression efficiency in cells, etc.

[0066] In addition, the present invention provides a method for preparing a cell with the target DNA modified, which includes the step of introducing the site-specific DNA nuclease of the present invention, the nucleic acid encoding the site-specific DNA nuclease, and the vector containing the nucleic acid or its translation product into the cell.

[0067] In the present invention, "alteration" of the target DNA includes deletion, insertion, or substitution of at least 1 nucleotide in the target DNA, or a combination thereof.

[0068] The introduction of the above-mentioned molecule into cells can be physical introduction, or introduction via infection with a virus or an organism, etc., and various methods known in the art can be applied. As physical introduction methods, for example, electroporation, particle gun method, microinjection method, lipofection method, protein transduction method, etc. are listed, but there is no particular limitation. As introduction methods via infection with a virus or an organism, etc., for example, virus transfection, Agrobacterium method, phage infection, conjugation, etc. are listed, but there is no particular limitation.

[0069] The above-mentioned introduction appropriately uses a vector. As the vector, it can be appropriately selected according to the type of the molecule to be introduced and the cell, for example, the above-mentioned plasmid vector, virus vector, phage vector, phagemid vector, BAC vector, YAC vector, MAC vector, HAC vector, etc. are listed, but there is no particular limitation. In addition, vectors such as liposomes and lentiviruses that can carry mRNA (transcription product), protein (translation product), etc., and peptide vectors such as cell-penetrating peptides that can introduce fused or conjugated molecules into cells are listed. Therefore, in the "vector containing a nucleic acid encoding a site-specific DNA-cleaving enzyme or its translation product" in the present invention, it includes not only the above-mentioned plasmid vectors and other vectors into which the DNA encoding the site-specific DNA-cleaving enzyme is inserted, but also the above-mentioned liposomes and other vectors that retain mRNA (transcription product), protein (translation product).

[0070] The cells into which the above-mentioned molecule is introduced can be either prokaryotic cells or eukaryotic cells. For example, bacteria, archaea, yeast, plant cells, insect cells, animal cells (for example, human cells, non-human cells, non-mammalian vertebrate cells, invertebrate cells, etc.) are listed. In addition, the cells can be cells in an organism, or can be isolated cells. In addition, they can be primary cells or cultured cells. In addition, the cells can be somatic cells, germ cells, or stem cells.

[0071] Prokaryotes as the source of the above cells include, for example, Escherichia coli, actinomycetes, archaea, etc. In addition, eukaryotes as the source of the above cells include, for example, fungi such as yeast, mushrooms, and molds; echinoderms such as sea urchins, starfish, and sea cucumbers; insects such as silkworms and flies; fish such as tuna, sea bream, pufferfish, skipjack, and zebrafish; rodents such as mice, rats, guinea pigs, hamsters, and squirrels; even-toed ungulates such as cattle, wild boars, pigs, sheep, and goats; odd-toed ungulates such as horses; reptiles such as lizards; amphibians such as frogs; lagomorphs such as rabbits; carnivores such as dogs, cats, and ferrets; birds such as chickens, ostriches, and quails; plants such as tobacco, Arabidopsis thaliana, rice, corn, bananas, peanuts, sunflowers, tomatoes, melons, rapeseed, wheat, barley, potatoes, soybeans, cotton, morning glories, cyclamens, and carnations. As "animal cells", for example, embryonic cells at various stages (such as 1-cell stage embryo, 2-cell stage embryo, 4-cell stage embryo, 8-cell stage embryo, 16-cell stage embryo, morula stage embryo, etc.) are listed; stem cells such as induced pluripotent stem (iPS) cells, embryonic stem (ES) cells, and hematopoietic stem cells; somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells, brain cells, and kidney cells; fertilized eggs, etc.

[0072] The target DNA can be any gene or extra-gene region DNA in the genomic DNA of the above cells. In order to cleave the target site in the target DNA, the site-specific DNA cleavage enzyme of the present invention is designed such that the DNA-binding domain contained in the site-specific DNA cleavage enzyme of the present invention binds to the sequence near the target cleavage site (selected as the target sequence of the DNA-binding domain). As a result, the sequence of the target cleavage site in the target DNA is cleaved, for example, the expression of the gene is reduced or disappears, or the function of the gene is reduced or not expressed.

[0073] In the method of the present invention, in addition to the above site-specific DNA cleavage enzyme, a donor polynucleotide can also be introduced into the cell. The donor polynucleotide contains at least 1 donor sequence containing the alteration to be introduced into the target cleavage site. The donor polynucleotide can be appropriately designed by those skilled in the art based on techniques known in the art. In the case where a donor polynucleotide exists in the method of the present invention, homologous recombination repair occurs at the target cleavage site, and the donor polynucleotide is inserted into the site, or the site is replaced by the donor sequence. As a result, the desired alteration is introduced at the target cleavage site.

[0074] In the case where no donor polynucleotide exists in the method of the present invention, the target cleavage site is mainly repaired by non-homologous end joining (NHEJ). Since NHEJ is prone to errors, deletions, insertions, or substitutions of at least 1 nucleotide, or combinations thereof, can occur in the repair of the cleavage. Thus, the target DNA is altered at the target cleavage site.

[0075] In addition, the present invention provides a kit for altering a target DNA, which comprises: the above-described site-specific DNA-cleaving enzyme, a nucleic acid encoding the site-specific DNA-cleaving enzyme, or a vector containing the nucleic acid or its translation product. The kit may further include one or more additional reagents. Examples of the additional reagents include, but are not limited to, a dilution buffer, a reconstitution solution, a wash buffer, a nucleic acid introduction reagent, a protein introduction reagent, and a control reagent. Usually, the kit is accompanied by an instruction manual.

[0076] Examples

[0077] 1. Method

[0078] (1) Preparation of TALE vector set and TALE expression plasmid

[0079] Regarding the sequence and composition of the TALE vector set, materials were de novo constructed according to the literature (Sakuma et al (2013) Scientific Reports, 3, 1-8). First, for the modular plasmids (a total of 16 types) corresponding to various RVDs of HD, NG, NI, and NN while retaining a variant called non-repetitive variable diresidue (non-RVD), sequences with recognition sites for the restriction enzyme BsAI added to both ends of each base sequence were prepared by artificial gene synthesis. These were inserted into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan) to prepare a set of modular plasmids (pEX1HD-pEX4HD, pEX1NG-pEX4NG, pEX1NI-pEX4NI, and pEX1NN-pEX4NN). Then, the DNA sequences FUS2_axx (7 types) and b(1-4) (4 types) constituting the array plasmid were prepared by artificial gene synthesis. These were inserted into pCR8 / GW / TOPO (Thermo Fisher Scientific, Waltham, MA, USA) to prepare pCR8_FUS2_axx and pCR8_FUS b(1-4), which were used as capture vectors in the initial assembly step of the Platinum Gate system described in the above-mentioned literature by Sakuma et al. The final vector was prepared by inserting three TALE constructs WT, +136 / +63, +153 / +47 and the final repeat, which were prepared by artificial gene synthesis, into pcDNA3.1s obtained by removing the drug resistance gene expression unit from pcDNA3.1(+) (Thermo Fisher Scientific). Using the above plasmids, TALE expression plasmids were prepared by the Golden Gate method according to the above-mentioned literature by Sakuma et al.

[0080] (2) Preparation of single-chain FokI, single-chain ND1, and single-chain ND2 expression plasmids

[0081] Synthesize a DNA sequence encoding the HTS95 amino acid sequence (Sun, N., & Zhao, H. (2014) MolecularBioSystems, 10(3), 446 - 453), and ligate an artificial gene (FokI-95-FokI) that connects two partial genes encoding the nuclease domain of the FokI gene. In addition, synthesize an artificial gene (FokI-GGGGSx12-FokI) in which the DNA sequence encoding the HTS95 amino acid sequence is replaced with a linker encoding GGGGSx12 (a total of 60 amino acid residues). For the ND1 gene and ND2 gene, artificially synthesize ND1-60-ND1 in which a linker of GGGGSx12 (a total of 60 amino acid residues) is inserted between partial genes encoding the nuclease domain of ND1, and ND2-95-ND2 in which the HTS95 amino acid sequence is inserted between partial genes encoding the nuclease domain of ND2. Additionally, during artificial synthesis, restriction enzyme AleI / XmaI sites are added as a common sequence before the linker sequences of GGGGSx12 and HTS95, and restriction enzyme PshAI / SacII sites are added as a common sequence after them for synthesis. Insert these artificially synthesized genes into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan).

[0082] Then, ND1-60-ND1 / pEX-A2J2 and ND2-95-ND2 / pEX-A2J2 were digested with restriction enzymes AleI and PshAI, and the 295 bp DNA fragment (95aa) encoding the HTS95 amino acid sequence, the 190 bp DNA fragment (60aa) encoding the linker of GGGGSx12, and the fragments of each remaining vector part were separated. After the 60aa was ligated by a ligation reaction with T4 DNA ligase, it was digested with restriction enzymes XmaI and SacII, and the resulting DNA fragments were separated by agarose gel electrophoresis, thereby separating and purifying the 380 bp DNA fragment (120aa) and the 570 bp DNA fragment (180bp). Fragments of 95aa, 120aa, and 180aa were respectively inserted into the vector part fragments obtained by digesting ND1-60-ND1 / pEX-A2J2 with restriction enzymes AleI and PshAI. In addition, fragments of 60aa, 120aa, and 180aa were respectively inserted into the vector part fragments obtained by digesting ND2-95-ND2 / pEX-A2J2 with restriction enzymes AleI and PshAI, to obtain plasmids ND1-95-ND1 / pEX-A2J2, ND1-120-ND1 / pEX-A2J2, ND1-180-ND1 / pEX-A2J2, ND2-60-ND2 / pEX-A2J2, ND2-120-ND2 / pEX-A2J2, and ND2-180-ND2 / pEX-A2J2.

[0083] Then, the TALE63-scFokI expression plasmid, TALE63-scND1 expression plasmid, and TALE63-scND2 expression plasmid, which connect the sequences of the above scFokI, scND1, and scND2 to the downstream of TALE(+136 / +63), were prepared by the following method. Using the FokI-95-FokI / pEX-A2J2, FokI-GGGGSx12-FokI / pEX-A2J2, ND1-95-ND1 / pEX-A2J2, ND1-60-ND1 / pEX-A2J2, ND1-120-ND1 / pEX-A2J2, ND1-180-ND1 / pEX-A2J2, ND2-95-ND2 / pEX-A2J2, ND2-60-ND2 / pEX-A2J2, ND2-120-ND2 / pEX-A2J2, ND2-180-ND2 / pEX-A2J2 of the above plasmids as templates, the sequences of scFokI, scND1, and scND2 were amplified by PCR respectively. They were inserted into the downstream of +136 / +63 of the platinum TALE structure, which is a TALE expression plasmid, by the In-Fusion method (TaKaRa Bio Inc, Shiga, Japan) to prepare the TALE63-scFokI expression plasmid, TALE63-scND1 expression plasmid, and TALE63-scND2 expression plasmid. The primers used in their preparation are shown in Table 1.

[0084]

[0085] In addition, separately, the base sequences and amino acid sequences of the N-terminal domain and C-terminal domain of TALE+136 / +63 are shown in SEQ ID NOs: 64 to 67, the base sequences and amino acid sequences of each scFokI are shown in SEQ ID NOs: 60 to 63, the base sequences and amino acid sequences of each scND1 are shown in SEQ ID NOs: 72 to 79, and the base sequences and amino acid sequences of each scND2 are shown in SEQ ID NOs: 80 to 87. The amino acid sequence of the C-terminal domain of TALE+136 / +63 (SEQ ID NO: 67) adds two amino acids "LK" on the C-terminal side when scFokI is used as the fusion nuclease domain.

[0086] (3) Preparation of the reporter plasmid for single-strand annealing (SSA) test

[0087] In the SSA test report plasmid, first, according to the sequence information described in the reference (Mashiko et al (2013) Scientific Reports, 3, 1 - 6), two amplified products of the N-terminal side and C-terminal side of the EGxxFP sequence were obtained by PCR using the fully synthesized EGFP cDNA as a template. Subsequently, after PCR was performed again with primers adding the linker sequence between the N-terminal side and C-terminal side to obtain the amplified products, they were inserted into the restriction enzyme BamHI / EcoRV sites of pcDNA3.1s by the In-Fusion method to prepare (the obtained plasmid is called "pcEGxxFP"). The primers used in the preparation of pcEGxxFP are shown in Table 2.

[0088]

[0089] Each target sequence used in the SSA test was isolated from the genome by PCR, and each SSA test report plasmid was prepared by inserting them into the restriction enzyme BamHI / EcoRI sites within the linker inserted between the N-terminal side and C-terminal side of EGFP in pcEGxxFP by the In-Fusion method. The target sequences (APC, Rosa26, HPRT1) used in the SSA test evaluation are shown in SEQ ID NOs: 94 - 96.

[0090] (4) Preparation of plasmids for TALE C-terminal length study

[0091] The study on the linker length between scND1 and the TALE repeats was carried out by changing the C-terminal length starting from the C-terminal of TALE +153 / +47. In the case of longer lengths, PCR was performed using TALE WT as a template to prepare 9 TALE C-terminal regions with linker lengths from 63 residues to 180 residues, and they were ligated to the N-terminal region of TALE +153 / +47. In the case of shorter lengths, PCR was performed using TALE +153 / +47 as a template to prepare 4 TALEs with linker lengths from 0 residues to 36 residues. The primers used in the preparation of expression plasmids for each TALE with different C-terminal lengths are shown in Table 3. In addition, the nucleotide sequences and amino acid sequences of the N-terminal domain and C-terminal domain of TALE +153 / +47 are shown in SEQ ID NOs: 68 - 71.

[0092]

[0093] (5) Preparation of flexible linker

[0094] Three flexible linkers, GSS, SAGG, and GGGGS, were prepared so that they could be inserted into the target sequence in a form where the number of repeats could be adjusted. First, linker sequences, GSSx4 adaptor, SAGGx3 adaptor, and GGGGSx3 adaptor (each forward and reverse), with a 5'-overhanging cleavage sequence configured with restriction enzyme XmaI at the 5'-end and BamHI or MroI at the 3'-end, were synthesized. The primers used in their preparation are shown in Table 4.

[0095]

[0096] After phosphorylating the 5'-ends of the respective oligomers and annealing them, they were inserted into plasmids having XmaI, BamHI, and MroI at each one position (for example, in the 60aa linker sequence of TALE63-ND1-60-ND1, since XmaI, BamHI, and MroI each exist at one position, TALE63-ND1-60-ND1 / pcDNA3.1s was used as a dummy vector in this example). They were the GSS adaptor vector, SAGG adaptor vector, and GGGGS adaptor vector, respectively.

[0097] In addition, as flexible linker repeats to be inserted into each adaptor vector, linker sequence oligomers, GSSx4, GSSx7, SAGGx3, SAGGx4, GGGGSx3, and GGGGSx4 (each forward and reverse), with a 5'-overhanging cleavage sequence configured with restriction enzyme BglII at the 5'-end and BamHI at the 3'-end, or with XmaI at the 5'-end and MroI at the 3'-end, were synthesized (Table 4).

[0098] After phosphorylating the 5'-ends of the respective oligomers and annealing them, a ligation reaction was carried out, and double cleavage of BglII and BamHI or XmaI and MroI was performed, and fragments of the target number of repeats were recovered by agarose gel electrophoresis. For example, in the case of GSSx4 repeats, repeats such as GSSx8, GSSx12, and GSSx16 were obtained, and by mixing with GSSx7 and carrying out a ligation reaction, DNA fragments with various numbers of repeats such as GSSx11, GSSx14, and GSSx15 could be obtained.

[0099] By inserting DNA fragments with each target number of repeats (the final target number of repeats - the number of repeats inserted into the adaptor vector) into the BamHI site or MorI site of the GSS adaptor vector, SAGG adaptor vector, and GGGGS adaptor vector, the linker sequences of the final target number of repeats were obtained. These linker sequences were separated by double cleavage of XmaI and BamHI or XmaI and MorI.

[0100] To insert the final repeat fragment in the form of replacement with the 95aa linker of the ND1-95-ND1 sequence, plasmids TALE24-scND1(bam) / pcDNA3.1s and TALE24-scND1(mro) / pcDNA3.1s were prepared by inserting restriction enzyme sites of BamHI or MorI at the C-terminal side before ND1 in the TALE24-ND1-95-ND1 expression plasmid. The primers used in the preparation of these plasmids are shown in Table 5.

[0101]

[0102] Then, after double digestion of each plasmid with XmaI and BamHI or XmaI and MorI, the 5'-ends were dephosphorylated with alkaline phosphatase. Subsequently, a ligation reaction was carried out together with the linker sequences of the final target repeat number prepared previously, and finally TALE24-ND1-GSSx32-ND1, TALE24-ND1-SAGGx24-ND1, and TALE24-ND1-GGGGSx19-ND1 were obtained. The base sequences and amino acid sequences of each fusion nuclease domain are shown in SEQ ID NOs: 88 to 93. TALE repeats for each target were inserted into the respective plasmids and used for the SSA test.

[0103] (6) Quantification of the double-strand cleavage activity of DNA by the SSA test

[0104] HEK293T cells grown in DMEM medium containing 10% FBS were seeded into each well of a 96-well plate at a density of 1x10 4 cells per well one day before transfection. Using Lipofectamine 3000 (Thermo Fisher Scientific), 33 ng of the TALE-L expression plasmid, 33 ng of the TALE-R expression plasmid, and 33 ng of the EGxxFP reporter plasmid were introduced into HEK293T cells. In cases where the types or amounts of the introduced plasmids were reduced, they were supplemented with pBluescript II (SK+)(Stratagene, La Jolla, CA, USA), and the total amount of the introduced plasmids was 100 ng.

[0105] After transfection, the cells were cultured for 48 hours, and the fluorescence intensity caused by EGFP protein per unit area of each well was measured with a microplate reader. To reduce the deviation in the localization of cells in the wells, fluorescence measurements were taken at 16 points in each well, and the average value was used as the fluorescence intensity. The average value and standard deviation of the values of three independent wells were graphed.

[0106] 2. Results

[0107] (1) Confirmation of the activity of scFokI

[0108] First, as a monomeric nuclease lacking base recognition activity, the scFokI activity of the reported molecule was confirmed. As the TALE target sequence, the Right TALE sequence (tACAGAAGCGGGCAAAGG / corresponding to positions 119 to 136 of SEQ ID NO: 94), Left TALE sequence (tATGTACGCCTCCCTGGG / corresponding to positions 85 to 102 of SEQ ID NO: 94) of the TALEN targeting the human adenomatous polyposis coli (APC) gene, whose intracellular activity had been confirmed in non-patent literature (the above-mentioned literature by Sakuma et al.) were applied. As a positive control, Platinum TALEN (APC-R) + Platinum TALEN (APC-L) was placed. When evaluated by the SSA assay in HEK293T cells, the activities of Platinum TALE (+136 / +63 scaffold, APC-R)-scFokI and Platinum TALE (+136 / +63 scaffold, APC-L)-scFokI were not significantly different from those of the reporter only in the SSA assay ( Figure 1 ). Therefore, in anticipation of improved activity, Platinum TALE (+136 / +63 scaffold, APC-R)-scFokI-12 and Platinum TALE (+136 / +63 scaffold, APC-L)-scFokI-12 with the linker sequence between FokI changed to a flexible linker of GGGGSx12 were prepared and evaluated in the same way, but the same activity as scFokI could not be confirmed ( Figure 1 ). It was judged that the cleavage activity of scFokI was very low and it was difficult to use for effective genome editing in human cells.

[0109] (2) Preparation and activity confirmation of scND1

[0110] Then, while following the structure of scFokI, the same study was carried out using ND1 / ND2, which was discovered as an alternative factor for FokI.

[0111] Using +136 / +63 Platinum TALE as a scaffold, expression plasmids of TALE(APC-L)ND1mono, TALE(APC-R)ND1mono, TALE(APC-L)ND1-60-ND1, TALE(APC-L)ND1-95-ND1, TALE(APC-L)ND1-120-ND1, TALE(APC-L)ND1-180-ND1, TALE(APC-L)ND2 mono, TALE(APC-R)ND2 mono, TALE(APC-L)ND2-60-ND2, TALE(APC-L)ND2-95-ND2, TALE(APC-L)ND2-120-ND2, and TALE(APC-L)ND2-180-ND2 were prepared ( Figure 2 ) and evaluated by the SSA assay in 293T cells. As a result, the activity of the conventional TALE-ND1 (TALE(APC-L)ND1mono + TALE(APC-R)ND1mono) was significantly lower than that of the conventional TALE-ND2 (TALE(APC-L)ND2mono + TALE(APC-R)ND2mono). However, in the case of ligation and single-strand formation, surprisingly, all ND1 showed excellent activity, while all ND2 did not show activity ( Figure 3 ). Furthermore, in the case of changing the TALE to a +153 / +47 scaffold, compared with the +136 / +63 scaffold, TALE-scND1 showed overall decreased activity, but TALE-ND1-95-ND1 maintained higher activity ( Figure 4) The above knowledge about scND1 was obtained by repeating the TALE for APC-L. However, the same is true for the other five target sequences, specifically APC-R, the Right TALE sequence of the TALEN designed at the human hypoxanthine phosphoribosyltransferase 1 (HPRT1) locus (tCGACATAGTGATTAGGA / corresponding to positions 357-374 of Sequence No. 96) and the Left TALE sequence (tGAACCAGGCTATGACC / corresponding to positions 324-340 of Sequence No. 96) (the above-mentioned literature by Sakuma et al.), and the Right TALE sequence of the TALEN designed at the human Rosa26 locus (tGCCCAGAAGACTCCCG / corresponding to positions 163-179 of Sequence No. 95) and the Left TALE sequence (tGATCTGCAAGTCGAGGC / corresponding to positions 130-147 of Sequence No. 95) (Sato et al (2015) Stem Cell Reports, 14;5(1), 75-82). TALE-ND1-95-ND1 shows activity ( Figure 5 )). Based on the above results, it is determined that the linker sequence between ND1s preferably has a sequence of 95 amino acids found in HTS. In subsequent experiments, the TALE of the +153 / +47 scaffold, which exhibits particularly high activity in combination with ND1-95-ND1, was mainly used.

[0112] (3) Optimization of the TALE-scND1 structure

[0113] Then, a study was conducted on the effect of the length of the C-terminal side sequence of the TALE scaffold +153 / +47 on the activity. First, when the C-terminal side sequence (63, 75, 90, 105, 120, 135, 150, 165, 180 residues) was extended in TALE (Rosa26-L) scND1 and TALE (Rosa26-R) scND1, no change in activity corresponding to the number of residues was observed ( Figure 6 )). On the contrary, when the number of residues was reduced to 36, 24, 12, 0, there were cases where the activity decreased among 0, 12, 36 residues depending on the target sequence. On the other hand, in the case of 24 residues, it was more stable than the original 47 residues and reached the same level or higher ( Figure 7 ).

[0114] In the "(2) Preparation and activity confirmation of scND1", the linker length between ND1s was changed to 60, 95, 120, and 180 residues, and the result that 95 residues were the most suitable was obtained. However, the linker of 95 residues is a unique sequence obtained from HTS, and the other 60, 120, and 180 residues are composed of repetitive flexible linkers of the GGGGS sequence. It is not clear whether the chain length of 95 residues is the most suitable as a linker or whether the HTS sequence is the most suitable. Therefore, three types of flexible linkers (GSSx32, SAGGx24, GGGGSx19) were prepared and inserted as the linker of scND1 used in the "(2) Preparation and activity confirmation of scND1" ( Figure 8 ), and the activity was compared with the case where the HTS95 sequence was used as the linker ( Figure 9 ). As a result, in the case of applying the HTS95 sequence, good activity was seen as a whole. In the case of using the APC gene as a target, the same activity as the HTS95 sequence was also seen for the other linkers.

[0115] Industrial applicability

[0116] The artificial nuclease incorporating the fusion nuclease domain of the present invention has double-strand cleavage activity for site-specific DNA per molecule, and can easily perform DNA editing as compared with a conventional TALEN in which two molecules binding to respective strands of DNA are used as a set. The artificial nuclease incorporating the fusion nuclease domain of the present invention can be used as an excellent genome editing tool in a wide range of industrial fields such as medicine, agriculture, and industry.

Claims

1. A fusion nuclease domain comprising two nuclease domains 1 bound via a linker.

2. A site-specific DNA cleavage enzyme comprising a DNA-binding domain and the fusion nuclease domain according to claim 1.

3. A nucleic acid encoding the fusion nuclease domain according to claim 1 or the site-specific DNA cleavage enzyme according to claim 2.

4. A vector comprising the nucleic acid according to claim 3 or its translation product.

5. A method for preparing a cell with altered target DNA, comprising the step of introducing a molecule selected from the following (a) to (c) into a cell: (a) The site-specific DNA cleavage enzyme according to claim 2 (b) A nucleic acid encoding the site-specific DNA cleavage enzyme of (a) (c) A vector comprising the nucleic acid of (b) or its translation product.

6. A kit for altering target DNA, comprising a molecule selected from the following (a) to (c): (a) The site-specific DNA cleavage enzyme according to claim 2 (b) A nucleic acid encoding the site-specific DNA cleavage enzyme of (a) (c) A vector comprising the nucleic acid of (b) or its translation product.

Citation Information

Patent Citations

  • Polypeptide containing DNA-binding domain

    JP2015033365A

  • Novel nuclease domain and uses thereof

    WO2020045281A1