Nuclease domain that makes double-strand breaks in DNA
The fusion nuclease domain with linked ND1 molecules addresses inefficiencies in existing genome editing technologies by enhancing double-strand cleavage activity, enabling efficient site-specific DNA editing with a single molecule.
Patent Information
- Application Number
- JP2021144906
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-06
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2041-09-06
AI Technical Summary
Existing genome editing technologies, such as TALENs and CRISPR-Cas9, face challenges in efficiently cleaving double-strand DNA due to the need for two nuclease domain molecules and difficulties in expression and dimerization, leading to inefficiencies in site-specific DNA editing.
A fusion nuclease domain comprising two ND1 nuclease domains linked via a linker, which enhances double-strand cleavage activity, allowing for efficient site-specific DNA editing with a single nuclease molecule.
The fusion nuclease domain enables efficient and simple site-specific DNA editing by eliminating the need for two nuclease molecules, improving the efficacy of genome editing processes.
Smart Images

Figure 0007720053000006 
Figure 0007720053000007 
Figure 0007720053000008
Abstract
Description
[Technical Field]
[0001] The present invention relates to a nuclease domain that causes double-strand cleavage in DNA, and more particularly to a fusion nuclease domain comprising two nuclease domains 1 (ND1) linked via a linker, and uses thereof. [Background technology]
[0002] TALENs, an artificial restriction enzyme developed as a second-generation genome editing technology in 2010, require the dimerization of the nonspecific DNA cleavage domain of the type IIS restriction enzyme FokI to cleave double-stranded DNA. TALEN technology requires a pair of paired TALE repeats. A pair of TALENs recognizes and cleaves a target sequence of 30–40 bases in total, resulting in extremely high cleavage specificity and strong suppression of off-target events. However, due to the difficulty of simultaneously expressing two TALENs (over 100 kDa) in cells or introducing equal amounts of TALEN protein into cells, as well as the labor required to create them, CRISPR-Cas9, which was released in 2012, has rapidly become a popular alternative to TALENs.
[0003] Meanwhile, research into improving TALEN molecules continues, resulting in the development of various derivative technologies, including the highly active Platinum TALE. Among these, there are reports of linking two FokI nuclease domains (scFokI) and binding them to zinc fingers or TALEs to induce double-strand breaks (Non-Patent Documents 1 and 2). This technology eliminates the need to use a pair of TALEs to cleave target sequences using TALENs. However, scFokI has low cleavage activity, making efficient genome editing in human cells difficult, and further improvement was necessary.
[0004] In addition, the inventors have developed two novel nuclease domains (nuclease domain 1; ND1 and nuclease domain 2; ND2) that are different from the conventional FokI nuclease domain, and have succeeded in genome editing of target sites by fusing these nuclease domains with ZF or TALE and using them as a set of ZFN or TALEN (Patent Document 1). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2020 / 045281 [Non-patent literature]
[0006] [Non-Patent Document 1] Minczuk et al (2008) Nucleic Acids Research, 36(12), 3926-3938 [Non-patent document 2] Mino et al (2009) Journal of Biotechnology, 140(3-4), 156-161 Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a nuclease domain that is capable of easily and efficiently cleaving double strands of DNA. [Means for solving the problem]
[0008] The present inventors conducted extensive research to solve the above-mentioned problems, and found that when two ND1s linked via a linker were used as the nuclease domain of a site-specific DNA cleaving enzyme (e.g., TALEN), they efficiently cleaved the target site in DNA in double strands. On the other hand, when two ND2s linked via a linker were used as the nuclease domain, or when two FokI nuclease domains linked via a linker were used, double-strand cleavage like that observed with ND1 was not observed. These facts revealed that the excellent double-strand cleavage activity of linked ND1 on DNA is not a phenomenon common to nuclease domains, but is an action specific to ND1.
[0009] Furthermore, the present inventors investigated the linker sequence between ND1 and ND1 and found that the double-strand cleavage activity can be further enhanced by adjusting the type and length of the linker sequence, thereby completing the present invention.
[0010] Therefore, the present invention relates to a fusion nuclease domain comprising two ND1s linked via a linker, and uses thereof, and more specifically provides the following.
[0011] (1) A fusion nuclease domain containing two nuclease domains 1 connected via a linker.
[0012] (2) A site-specific DNA cleavage enzyme comprising a DNA binding domain and the fusion nuclease domain described in (1).
[0013] (3) A nucleic acid encoding the fusion nuclease domain described in (1) or the site-specific DNA cleaving enzyme described in (2).
[0014] (4) A vector containing the nucleic acid according to (3) or its translation product.
[0015] (5) A method for producing a cell in which target DNA has been modified, comprising introducing into the cell a molecule selected from the following (a) to (c): (a) (2) The site-specific DNA cleaving enzyme (b) a nucleic acid encoding the site-specific DNA cleavage enzyme of (a) (c) A vector containing the nucleic acid according to (b) or its translation product.
[0016] (6) A kit for modifying target DNA, comprising a molecule selected from the following (a) to (c): (a) (2) The site-specific DNA cleaving enzyme (b) a nucleic acid encoding the site-specific DNA cleavage enzyme of (a) (c) A vector containing the nucleic acid according to (b) or its translation product. [Effects of the Invention]
[0017] According to the present invention, double-stranded DNA can be cleaved with a single nuclease domain molecule, eliminating the need for two nuclease domain molecules for double-stranded DNA cleavage, as in the past. Therefore, by using the nuclease domain of the present invention in site-specific DNA cleavage enzymes such as TALEN and ZFN, site-specific DNA editing can be performed simply and efficiently. [Brief explanation of the drawings]
[0018] [Figure 1] 1 shows the results of the SSA assay for the DNA cleavage activity of a TALEN containing two FokIs linked via a linker (referred to as "TALE-scFokI"). In the figure, "95" indicates the 95-amino acid HTS95 linker, and "60" indicates the 60-amino acid GGGGSx12 linker (the same applies below). [Figure 2]This figure shows the structure of a TALEN (referred to as "TALE-scND1" and "TALE-scND2," respectively) containing two nuclease domains (ND1 or ND2) linked via a linker. TALE63 is used as an example of a TALE (DNA binding domain). Note that "120" in the figure indicates a 120 amino acid residue GGGGS x 24 linker, and "180" indicates a 180 amino acid residue GGGGS x 36 linker (the same applies below). [Figure 3] This is a graph showing the results of detecting DNA cleavage activity by SSA assay using TALE-scND1 and TALE-scND2 shown in Figure 2. [Figure 4] This is a graph showing the results of an SSA assay evaluating the effects of TALE structure and linker length on the DNA cleavage activity of TALE-scND1. TALE63 and TALE47 were used as TALEs. [Figure 5] This graph shows the results of an SSA assay evaluating the effects of the type of target gene and the length of the linker on the DNA cleavage activity of TALE-scND1. The target genes used were (A) the APC gene, and (B) the Rosa26 gene and the HPRT1 gene. TALE47 was used as the TALE. [Figure 6] This graph shows the results of an SSA assay evaluating the effect of lengthening the C-terminal domain of TALE on the DNA cleavage activity of TALE-scND1. The target gene used was the Rosa26 gene. TALE47 was used as the TALE. The linker connecting the two ND1s was the 95-amino acid HTS95 linker. [Figure 7] This graph shows the results of an SSA assay evaluating the effect of shortening the C-terminal domain of TALE on the DNA cleavage activity of TALE-scND1. The target genes used were the Rosa26 gene, the APC gene, and the HPRT1 gene. TALE47 was used as the TALE. The linker connecting the two ND1s was the 95-amino acid HTS95 linker. [Figure 8]This figure shows the structure of TALE-scND1 used in the experiment in Figure 9. [Figure 9] This graph shows the results of an SSA assay to evaluate the effect of linker type on the DNA cleavage activity of TALE-scND1. The linkers used were HTS95 (95 amino acid residues), GSSx32 (96 amino acid residues), SAGGx24 (96 amino acid residues), and GGGGSx19 (95 amino acid residues). The target genes used were (A) the Rosa26 gene, (B) the APC gene, and (C) the HPRT1 gene. TALE24 was used as the TALE. DETAILED DESCRIPTION OF THE INVENTION
[0019] The present invention provides a fusion nuclease domain (scND1) comprising two ND1s linked via a linker.
[0020] "ND1" in the present invention is one of the nuclease domains discovered by the present inventors through screening of homologous sequences that have an identity in the range of 35% to 70% with the FokI nuclease domain (Patent Document 1).
[0021] The amino acid sequence of a full-length protein containing ND1 (as a representative example, derived from Bacillus SGD-V-76) is shown in SEQ ID NO: 97. ND1 is typically a partial peptide corresponding to positions 391 to 585 of SEQ ID NO: 97, and has 70% identity with the amino acid sequence of the FokI nuclease domain.
[0022] "ND1" in the present invention includes a nuclease domain consisting of an amino acid sequence highly identical to the amino acid sequence corresponding to positions 391 to 585 of SEQ ID NO: 97, so long as it has double-stranded DNA cleavage activity when two molecules are linked via a linker. Such nuclease domains include, for example, ND1 derived from other bacteria and ND1 mutants (natural mutants and artificial mutants).
[0023] Here, "high identity" means 85% or more identity, preferably 90% or more (for example, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more) identity.
[0024] In the present invention, the identity of amino acid sequences is determined by comparing two sequences aligned to maximize sequence identity. Methods for determining the numerical value (%) of sequence identity are known to those skilled in the art. As an algorithm for obtaining optimal alignment and sequence identity, any algorithm known to those skilled in the art (e.g., BLAST algorithm, FASTA algorithm, etc.) may be used. The sequence identity of amino acid sequences is determined, for example, using sequence analysis software such as BLASTP or FASTA.
[0025] An example of an ND1 mutant is a nuclease domain consisting of an amino acid sequence in which one to several amino acid residues have been substituted, deleted, inserted, or added in the amino acid sequence shown at positions 391 to 585 of SEQ ID NO: 97. Here, "one to several" means, for example, 1 to 30, preferably 1 to 20 (e.g., 1 to 10, 1 to 5, or 1 to 3).
[0026] The "linker" used to link two ND1s is not particularly limited in length or type, as long as the fusion nuclease domain of the present invention can exhibit double-strand cleavage activity. The linker length is usually 30 to 300 amino acid residues, and preferably 50 to 200 amino acid residues. Examples of linker types include the HTS95 linker, GSS linker, SAGG linker, GGGGS linker, and HTEN linker, with the HTS95 linker being preferred. When using these linkers, multiple linkers may be linked together depending on the required linker length.
[0027] The binding of two ND1s and a linker can be carried out at the nucleic acid level and the amino acid level. That is, a nucleic acid encoding the fusion nuclease domain of the present invention can be prepared by linking the respective nucleic acids by a ligation reaction in the order of "nucleic acid encoding ND1 → nucleic acid encoding linker → nucleic acid encoding ND1." The nucleic acid thus obtained is inserted into an expression vector and expressed in an appropriate host cell, resulting in a fusion nuclease domain in which the ND1s are linked at the amino acid level in the order of "ND1 → linker → ND1." The host cell and vector are the same as those for the site-specific DNA cleaving enzyme of the present invention (described below). The fusion nuclease domain can also be prepared by artificial synthesis based on amino acid sequence information.
[0028] Whether a fusion nuclease domain has DNA double-strand cleavage activity can be evaluated by determining whether a site-specific DNA cleavage enzyme obtained by fusing the fusion nuclease domain with a DNA-binding domain exhibits site-specific DNA double-strand cleavage activity, as described below. For example, as described in the Examples of the present application, a site-specific DNA cleavage enzyme obtained by fusing the fusion nuclease domain with a TALE that targets a specific site in a specific gene can be evaluated by determining whether it exhibits double-strand cleavage activity at the specific site. This evaluation can be performed using an SSA assay system targeting a reporter gene on a plasmid (e.g., an EGFP gene into which a TALE recognition sequence has been inserted) and using reporter activity as an indicator. Alternatively, a site-specific DNA cleavage enzyme targeting an endogenous gene in a cell can be prepared and evaluation can be performed using modification (mutation) of the base sequence at the target site by the site-specific DNA cleavage enzyme as an indicator.
[0029] The present invention also provides a site-specific DNA cleaving enzyme comprising a DNA-binding domain and the above-described fusion nuclease domain. The site-specific DNA cleaving enzyme of the present invention binds to a target sequence on DNA via the DNA-binding domain and causes double-strand cleavage of the DNA at the target cleavage site via the fusion nuclease domain.
[0030] The "target DNA" of the site-specific DNA cleaving enzyme of the present invention is preferably double-stranded DNA, since this allows the enzyme to most effectively exhibit its properties. Examples of double-stranded DNA include, but are not limited to, eukaryotic nuclear genomic DNA, mitochondrial DNA, plastid DNA, prokaryotic genomic DNA, phage DNA, and plasmid DNA.
[0031] The DNA-binding domain is not particularly limited as long as it is a protein domain that can specifically bind to any DNA sequence (target sequence), and examples include TALE (Transcription Activator-Like Effector), zinc finger, PPR (Pentatricopeptide repeat), and CRISPR / Cas (a complex of Cas protein and guide RNA).
[0032] TALEs are proteins secreted by Xanthomonas proteobacteria that activate gene transcription in host plants. The TALEs used in the present invention may contain an N-terminal domain and / or a C-terminal domain in addition to the TALE repeat domain.
[0033] A TALE repeat domain is composed of multiple, for example, 10 to 30, preferably 13 to 25, and more preferably 15 to 20 tandem repeats (TALE repeats) of TALE sequences that form a right-handed superhelical structure. A typical TALE repeat unit (one TALE sequence) consists of 33 to 35 amino acids, and recognizes specific bases in DNA using a repeat variable residue (RVD) consisting of two amino acid residues at positions 12 and 13. Examples of RVDs that specifically recognize bases include HD, which recognizes C; NG, which recognizes T; NI, which recognizes A; NN, which recognizes G or A; and NS, which recognizes A, C, G, or T. Based on the DNA recognition mechanism of such a TALE repeat domain, TALEs that can recognize and bind to desired base sequences (TALE recognition sequences) on DNA can be created by artificially linking TALE sequences that recognize specific bases.
[0034] The TALE is preferably a Platinum TALEN (JP Patent Publication No. 2015-33365) TALE in which amino acids at two specific positions in one TALE repeat unit are changed every four TALE repeat units.
[0035] The N-terminal domain of a TALE can be a naturally occurring amino acid sequence (e.g., the amino acid sequence of the N-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185-32188)), but it may also be appropriately modified from the naturally occurring amino acid sequence as long as it does not negatively affect the function of the site-specific DNA cleavage enzyme of the present invention (double-stranded DNA cleavage at the target cleavage site). Specific examples of the N-terminal domain of a TALE include, but are not limited to, the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-136 / 63-VR-HD (ID: 50699) and the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-153 / 47-VR-HD (ID: 50703).
[0036] The C-terminal domain of a TALE can be a naturally occurring amino acid sequence (e.g., the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185-32188)), but may also be appropriately modified from the naturally occurring amino acid sequence as long as it does not negatively affect the function of the site-specific DNA cleavage enzyme of the present invention (double-stranded DNA cleavage at the target cleavage site). Specific examples of the C-terminal domain include, but are not limited to, the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185-32188) (WT: 180 amino acids), the amino acid sequence of the C-terminal domain contained in Addgene's pTALEN_v2 (ID: 32189-32192) (63 amino acids), and the amino acid sequence of the C-terminal domain contained in Addgene's ptCMV-153 / 47-VR-NG (ID: 50704) (47 amino acids). In fact, in this example, site-specific DNA double-strand cleavage activity was observed both when a TALE without a C-terminal domain was used and when a TALE with a long C-terminal domain of 180 amino acid residues was used (Figures 6 and 7).
[0037] A DNA-binding domain containing zinc fingers is composed of a zinc finger array (hereinafter also referred to as "ZFA") consisting of multiple, for example, 3 to 9 zinc fingers. Each zinc finger is known to recognize a three-base sequence, and zinc fingers that recognize, for example, GNN, ANN, or CNN are known.
[0038] A DNA-binding domain containing PPR is composed of multiple, e.g., 10 to 30 tandem repeats (PPR repeats). A typical PPR repeat unit consists of 35 amino acids, and each PPR repeat unit is known to recognize one base. Furthermore, while TALEs specify the binding base with two amino acids present in one repeat unit, PPRs are known to specify the binding base with three amino acids.
[0039] CRISPR-Cas is a complex containing a Cas protein and a guide RNA, and examples include CRISPR-Cas9, CRISPR-Cpf1 (Cas12a), CRISPR-Cas12b, CRISPR-CasX (Cas12e), CRISPR-Cas14, and CRISPR-Cas3. The CRISPR-Cas used in the present invention is preferably a CRISPR-Cas containing a Cas with inactivated nuclease activity (e.g., dCas). Cas mutants with inactivated nuclease activity are known. When the fused nuclease domain of the present invention is applied to CRISPR-Cas, it is typically bound to a protein that constitutes CRISPR-Cas. When CRISPR-Cas is a Type I system consisting of multiple proteins (subunits), the fused nuclease domain of the present invention may be fused to a subunit that does not have nuclease activity (e.g., a subunit that constitutes Cascade in CRISPR-Cas3). In this case, the subunit with nuclease activity (e.g., Cas3 in CRISPR-Cas3) can be excluded from the components of the CRISPR-Cas system.
[0040] In the site-specific DNA cleaving enzyme of the present invention, the fusion nuclease domain and the DNA-binding domain may be linked directly or via a linker. When a linker is present, there are no particular limitations on its length or type, as long as the fusion nuclease domain of the present invention can exhibit double-strand cleavage activity. The linker length is generally 2 to 180 amino acid residues, preferably 2 to 120 amino acid residues. Examples of linker types include, but are not limited to, the HTS95 linker, GSS linker, SAGG linker, GGGGS linker, and HTEN linker. When these linkers are used, multiple linkers can be linked together depending on the required linker length.
[0041] The fusion nuclease domain may be bound to the N-terminus or C-terminus of the DNA-binding domain. Alternatively, a sandwich structure may be formed by placing DNA-binding domains at both ends of the fusion nuclease domain (see Mori et al. (2009) Biochemical and Biophysical Research Communications, 390(3), 694-697 for details on sandwich structures). Furthermore, a flag tag for purification or detection, or various transport signals (e.g., nuclear transport signal, mitochondrial transport signal, plastid transport signal, etc.) may be added.
[0042] The target sequence recognized by the DNA-binding domain of the site-specific DNA cleaving enzyme of the present invention is any sequence on DNA. Conventional ZFNs and TALENs, which are used in pairs (two molecules), require the target sequence to consist of two sequences sandwiched between a spacer sequence, but the site-specific DNA cleaving enzyme of the present invention can be used in a single molecule, which is advantageous in that there are fewer restrictions on targeting.
[0043] The site-specific DNA cleaving enzyme of the present invention can be produced by methods known to those skilled in the art. Examples include a method in which a nucleic acid encoding the site-specific DNA cleaving enzyme of the present invention is inserted into an appropriate expression vector, and the expression vector is then introduced into a host cell to express the site-specific DNA cleaving enzyme in the host cell. Alternatively, a method in which the site-specific DNA cleaving enzyme is artificially synthesized based on the amino acid sequence information of the site-specific DNA cleaving enzyme is synthesized. The expression vector may be appropriately selected from vectors used in the art, such as plasmid vectors, virus vectors, phage vectors, phagemid vectors, BAC vectors, YAC vectors, MAC vectors, and HAC vectors. The host cell into which the expression vector is introduced may be appropriately selected in consideration of compatibility with the expression vector, such as prokaryotic cells, such as Escherichia coli, actinomycetes, and archaea, and eukaryotic cells, such as yeast, sea urchin, silkworm, zebrafish, mouse, rat, frog, tobacco, Arabidopsis, and rice.
[0044] The present invention provides a nucleic acid encoding the fusion nuclease domain and a nucleic acid encoding the site-specific DNA cleaving enzyme. Here, "nucleic acid" encompasses both DNA and RNA (mRNA). The nucleic acid may be codon-optimized for the purpose of increasing intracellular expression efficiency.
[0045] The present invention also provides a method for producing a cell in which target DNA has been modified, which method comprises introducing into a cell the site-specific DNA cleaving enzyme of the present invention, a nucleic acid encoding the site-specific DNA cleaving enzyme, or a vector containing the nucleic acid or a translation product thereof.
[0046] In the present invention, "modification" of a target DNA includes deletion, insertion, or substitution of at least one nucleotide in the target DNA, or a combination thereof.
[0047] Introduction of the above-mentioned molecules into cells may be physical introduction or introduction via viral or biological infection, and various methods known in the art can be used. Physical introduction methods are not particularly limited, and examples include electroporation, particle gun method, microinjection method, lipofection method, and protein transduction method. Introduction methods via viral or biological infection are not particularly limited, and examples include viral transduction, Agrobacterium method, phage infection, and conjugation.
[0048] A vector is used as appropriate for the introduction. The vector may be appropriately selected depending on the type of molecule and cell to be introduced and is not particularly limited. Examples include the above-mentioned plasmid vectors, viral vectors, phage vectors, phagemid vectors, BAC vectors, YAC vectors, MAC vectors, and HAC vectors. Other examples include vectors such as liposomes and lentiviruses that can transport mRNA (transcription products) and proteins (translation products), as well as peptide vectors such as cell-penetrating peptides that can introduce fused or conjugated molecules into cells. Therefore, the "vector containing a nucleic acid encoding a site-specific DNA cleaving enzyme or its translation product" in the present invention includes not only vectors such as the above-mentioned plasmid vectors into which DNA encoding the site-specific DNA cleaving enzyme has been inserted, but also vectors such as the above-mentioned liposomes that carry mRNA (transcription products) and proteins (translation products).
[0049] The cells into which the molecules are introduced may be either prokaryotic or eukaryotic. Examples include bacteria, archaea, yeast, plant cells, insect cells, and animal cells (e.g., human cells, non-human cells, non-mammalian vertebrate cells, invertebrate cells, etc.). The cells may be in vivo cells, isolated cells, primary cells, or cultured cells. The cells may also be somatic cells, germ cells, or stem cells.
[0050] Examples of prokaryotes from which the above cells are derived include Escherichia coli, actinomycetes, and archaea. Examples of eukaryotic organisms from which the cells are derived include fungi such as yeast, mushrooms, and mold; echinoderms such as sea urchins, starfish, and sea cucumbers; insects such as silkworms and flies; fish such as tuna, sea bream, pufferfish, barracuda, and zebrafish; rodents such as mice, rats, guinea pigs, hamsters, and squirrels; even-toed ungulates such as cattle, wild boars, pigs, sheep, and goats; perissodactyls such as horses; reptiles such as lizards; amphibians such as frogs; lagomorphs such as rabbits; carnivora such as dogs, cats, and ferrets; birds such as chickens, ostriches, and quails; and plants such as tobacco, Arabidopsis, rice, corn, bananas, peanuts, sunflowers, tomatoes, melons, rapeseed, wheat, barley, potatoes, soybeans, cotton, morning glory, cyclamen, and carnations. Examples of "animal cells" include embryonic cells of various stages of embryos (e.g., 1-cell embryo, 2-cell embryo, 4-cell embryo, 8-cell embryo, 16-cell embryo, morula embryo, etc.); stem cells such as induced pluripotent stem (iPS) cells, embryonic stem (ES) cells, and hematopoietic stem cells; somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells, brain cells, and kidney cells; and fertilized eggs.
[0051] The target DNA may be any gene or extragenic DNA in the genomic DNA of the cell. To cleave a target site in the target DNA, the site-specific DNA cleaving enzyme of the present invention is designed so that the DNA-binding domain contained in the site-specific DNA cleaving enzyme binds to a sequence near the target cleavage site (selected as the target sequence for the DNA-binding domain). As a result, the sequence at the target cleavage site in the target DNA is cleaved, resulting in, for example, reduced or eliminated expression of the gene, or reduced or absent function of the gene.
[0052] In the method of the present invention, in addition to the site-specific DNA cleavage enzyme, a donor polynucleotide may be introduced into cells. The donor polynucleotide comprises at least one donor sequence containing the modification to be introduced into the target cleavage site. The donor polynucleotide can be appropriately designed by those skilled in the art based on techniques known in the art. When the donor polynucleotide is present in the method of the present invention, homologous recombination repair occurs at the target cleavage site, and the donor polynucleotide is inserted into the site or the site is replaced by the donor sequence. As a result, the desired modification is introduced into the target cleavage site.
[0053] In the method of the present invention, when there is no donor polynucleotide, the target cleavage site is mainly repaired by non-homologous end joining (NHEJ).Because NHEJ is error-prone, at least one nucleotide deletion, insertion, or substitution, or a combination thereof, may occur during the repair of the cleavage.Thus, the target DNA is modified at the target cleavage site.
[0054] The present invention also provides a kit for modifying target DNA, comprising the site-specific DNA cleaving enzyme, a nucleic acid encoding the site-specific DNA cleaving enzyme, or a vector containing the nucleic acid or its translation product. The kit may further comprise one or more additional reagents, including, but not limited to, a dilution buffer, a reconstitution solution, a washing buffer, a nucleic acid introduction reagent, a protein introduction reagent, and a control reagent. Typically, the kit comes with an instruction manual. [Example]
[0055] 1. Method (1) Preparation of TALE vector set and TALE expression plasmid The sequence and structure of the TALE vector set were constructed from scratch according to the literature (Sakuma et al. (2013) Scientific Reports, 3, 1-8). First, we constructed 16 modular plasmids corresponding to the HD, NG, NI, and NN RVDs, while retaining variations known as non-repeat-variable di-residues (non-RVDs). Each nucleotide sequence was enriched with BsAI restriction enzyme recognition sites at both ends. These were then inserted into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan) to create modular plasmid sets (pEX1HD-pEX4HD, pEX1NG-pEX4NG, pEX1NI-pEX4NI, and pEX1NN-pEX4NN). Next, we constructed the DNA sequences FUS2_axx (7 types) and b(1-4) (4 types) that comprise the array plasmids by artificial gene synthesis. These vectors were inserted into pCR8 / GW / TOPO (Thermo Fisher Scientific, Waltham, MA, USA) to generate pCR8_FUS2_axx and pCR8_FUS b(1-4), which were used as capture vectors for the first assembly step in the Platinum Gate system described in Sakuma et al. The final vectors were constructed by inserting three TALE constructs (WT, +136 / +63, +153 / +47, and the final repeat) into pcDNA3.1s, which was prepared by removing the drug resistance gene expression unit from pcDNA3.1(+) (Thermo Fisher Scientific). Using these plasmids, TALE expression plasmids were constructed by the Golden Gate method according to Sakuma et al.
[0056] (2) Construction of single-chain FokI, single-chain ND1, and single-chain ND2 expression plasmids We synthesized an artificial gene (FokI-95-FokI) by linking two partial genes encoding the nuclease domain of the FokI gene via a DNA sequence encoding the HTS95 amino acid sequence (Sun, N., & Zhao, H. (2014) Molecular BioSystems, 10(3), 446-453). We also synthesized an artificial gene (FokI-GGGGSx12-FokI) in which the HTS95 amino acid sequence was replaced with a DNA sequence encoding a linker of GGGGSx12 (60 amino acid residues in total). For the ND1 and ND2 genes, we artificially synthesized ND1-60-ND1, in which a linker of GGGGSx12 (60 amino acid residues in total) was inserted between the partial genes encoding the ND1 nuclease domain, and ND2-95-ND2, in which the HTS95 amino acid sequence was inserted between the partial genes encoding the ND2 nuclease domain. During artificial synthesis, the AleI / XmaI restriction enzyme site was added before the linker sequence of GGGGSx12 and HTS95, and the PshAI / SacII restriction enzyme site was added after the linker sequence as a consensus sequence. These artificially synthesized genes were inserted into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan).
[0057] Next, ND1-60-ND1 / pEX-A2J2 and ND2-95-ND2 / pEX-A2J2 were digested with the restriction enzymes AleI and PshAI to isolate a 295-bp DNA fragment (95 aa) encoding the HTS95 amino acid sequence, a 190-bp DNA fragment (60 aa) encoding the GGGGSx12 linker, and the remaining vector fragments. The 60-aa fragment was ligated using T4 DNA ligase and then digested with the restriction enzymes XmaI and SacII. The resulting DNA fragments were separated by agarose gel electrophoresis to isolate and purify a 380-bp DNA fragment (120 aa) and a 570-bp DNA fragment (180 bp). ND1-60-ND1 / pEX-A2J2 was cleaved with the restriction enzymes AleI and PshAI to obtain a vector fragment, and the 95aa, 120aa, and 180aa fragments were inserted into the vector fragment, respectively. Also, ND2-95-ND2 / pEX-A2J2 was cleaved with the restriction enzymes AleI and PshAI to obtain a vector fragment, and the 60aa, 120aa, and 180aa fragments were inserted into the vector fragment, respectively. Thus, plasmids ND1-95-ND1 / pEX-A2J2, ND1-120-ND1 / pEX-A2J2, ND1-180-ND1 / pEX-A2J2, ND2-60-ND2 / pEX-A2J2, ND2-120-ND2 / pEX-A2J2, and ND2-180-ND2 / pEX-A2J2 were obtained.
[0058] Next, the TALE63-scFokI expression plasmid, TALE63-scND1 expression plasmid, and TALE63-scND2 expression plasmid were constructed by linking the above scFokI, scND1, and scND2 sequences downstream of TALE(+136 / +63) using the following method. Using the above-mentioned plasmids FokI-95-FokI / pEX-A2J2, FokI-GGGGSx12-FokI / pEX-A2J2, ND1-95-ND1 / pEX-A2J2, ND1-60-ND1 / pEX-A2J2, ND1-120-ND1 / pEX-A2J2, ND1-180-ND1 / pEX-A2J2, ND2-95-ND2 / pEX-A2J2, ND2-60-ND2 / pEX-A2J2, ND2-120-ND2 / pEX-A2J2, and ND2-180-ND2 / pEX-A2J2 as templates, the scFokI, scND1, and scND2 sequences were amplified by PCR. These were inserted downstream of the platinum TALE structure +136 / +63 of the TALE expression plasmids using the In-Fusion method (TaKaRa Bio Inc, Shiga, Japan) to generate the TALE63-scFokI expression plasmid, TALE63-scND1 expression plasmid, and TALE63-scND2 expression plasmid. The primers used for their construction are shown in Table 1.
[0059] [Table 1]
[0060] The nucleotide and amino acid sequences of the N-terminal and C-terminal domains of TALE+136 / +63 are shown in SEQ ID NOs: 64 to 67, the nucleotide and amino acid sequences of each scFokI are shown in SEQ ID NOs: 60 to 63, the nucleotide and amino acid sequences of each scND1 are shown in SEQ ID NOs: 72 to 79, and the nucleotide and amino acid sequences of each scND2 are shown in SEQ ID NOs: 80 to 87. When scFokI is used as the fusion nuclease domain, the amino acid sequence of the C-terminal domain of TALE+136 / +63 (SEQ ID NO: 67) has two additional amino acids, "LK," added to the C-terminus.
[0061] (3) Preparation of reporter plasmid for single-strand annealing (SSA) assay For the reporter plasmid for SSA assays, first, referring to the sequence information described in the literature (Mashiko et al. (2013) Scientific Reports, 3, 1-6), a PCR was performed using fully synthesized EGFP cDNA as a template to obtain two amplification products, one for the N-terminal and one for the C-terminal ends of the EGxxFP sequence. Next, PCR was performed again using primers with linker sequences added between the N- and C-terminal ends to obtain amplification products, which were then inserted into the BamHI / EcoRV restriction enzyme sites of pcDNA3.1s using the In-Fusion method (the resulting plasmid is referred to as "pcEGxxFP"). The primers used to generate pcEGxxFP are listed in Table 2.
[0062] [Table 2]
[0063] Each target sequence used in the SSA assay was isolated from the genome by PCR and inserted into the BamHI / EcoRI restriction enzyme sites in the linker inserted between the N- and C-termini of EGFP in pcEGxxFP by the In-Fusion method to prepare reporter plasmids for the SSA assay. The target sequences (APC, Rosa26, and HPRT1) used in the SSA assay evaluation are shown in SEQ ID NOs: 94 to 96.
[0064] (4) Preparation of plasmid for TALE C-terminal length analysis The length between scND1 and the TALE repeat was examined by varying the C-terminal length starting from the C-terminus of TALE+153 / +47. To lengthen the repeat, PCR was performed using TALE WT as a template to create nine TALE C-terminal regions with lengths ranging from 63 to 180 residues, which were then linked to the N-terminal region of TALE+153 / +47. To shorten the repeat, PCR was performed using TALE+153 / +47 as a template to create four TALEs with lengths ranging from 0 to 36 residues. The primers used to create expression plasmids for each TALE with different C-terminal lengths are shown in Table 3. The nucleotide and amino acid sequences of the N- and C-terminal domains of TALE+153 / +47 are shown in SEQ ID NOs: 68 to 71.
[0065] [Table 3]
[0066] (5) Preparation of flexible linkers Three flexible linkers, GSS, SAGG, and GGGGS, were constructed so that the repeat number could be adjusted for insertion into the target sequence. First, the linker sequences GSSx4, SAGGx3, and GGGGSx3 (forward and reverse, respectively) were synthesized, each containing the restriction enzyme XmaI at the 5' end and a 5' overhanging cleavage sequence for BamHI or MroI at the 3' end. The primers used for their construction are listed in Table 4.
[0067] [Table 4]
[0068] After phosphorylating the 5' end of each oligo and annealing, one XmaI, one BamHI, and one MroI site were inserted into a plasmid (for example, the 60aa linker sequence of TALE63-ND1-60-ND1 contains one XmaI, one BamHI, and one MroI site, so in this example, TALE63-ND1-60-ND1 / pcDNA3.1s was used as a temporary vector). These were designated as GSS adapter vector, SAGG adapter vector, and GGGGS adapter vector, respectively.
[0069] In addition, we synthesized linker sequence oligos GSSx4, GSSx7, SAGGx3, SAGGx4, GGGGSx3, and GGGGSx4 (forward and reverse, respectively) as flexible linker repeat sequences to be inserted into each adapter vector, each with a 5'-overhanging restriction enzyme BglII at the 5' end and BamHI at the 3' end, or XmaI at the 5' end and MroI at the 3' end (Table 4).
[0070] The 5' end of each oligo was phosphorylated and annealed, followed by ligation. The oligos were then double-cleaved with BglII and BamHI or XmaI and MroI, and fragments with the desired repeat number were recovered by agarose gel electrophoresis. For example, in the case of a GSSx4 repeat, repeat sequences such as GSSx8, GSSx12, and GSSx16 were obtained. By mixing this with GSSx7 and ligating it, DNA fragments with various repeat numbers, such as GSSx11, GSSx14, and GSSx15, could be obtained.
[0071] DNA fragments with the desired number of repeats (final target number of repeats minus the number of repeats inserted into the adapter vector) were inserted into the BamHI or MorI site of the GSS, SAGG, or GGGGS adapter vectors to obtain linker sequences with the desired number of repeats. These linker sequences were isolated by double digestion with XmaI and BamHI or XmaI and MorI.
[0072] To insert the final repeat fragment in place of the 95aa linker in the ND1-95-ND1 sequence, we constructed the plasmids TALE24-scND1(bam) / pcDNA3.1s and TALE24-scND1(mro) / pcDNA3.1s by inserting a BamHI or MorI restriction enzyme site immediately before the C-terminal ND1 in the TALE24-ND1-95-ND1 expression plasmid. The primers used to construct these plasmids are shown in Table 5.
[0073] [Table 5]
[0074] Next, each plasmid was double-cleaved with XmaI and BamHI or XmaI and MorI, and the 5' end was dephosphorylated with alkaline phosphatase. Subsequently, a ligation reaction was performed with the linker sequence corresponding to the final target repeat number prepared previously, ultimately yielding TALE24-ND1-GSSx32-ND1, TALE24-ND1-SAGGx24-ND1, and TALE24-ND1-GGGGSx19-ND1. The nucleotide and amino acid sequences of each fusion nuclease domain are shown in SEQ ID NOs: 88 to 93. TALE repeats for each target were inserted into each plasmid and subjected to SSA assays.
[0075] (6) Quantification of DNA double-strand break activity by SSA assay 1x10 HEK293T cells grown in DMEM medium containing 10% FBS 4Cells were seeded into each well of a 96-well plate the day before transfection. 33 ng of the TALE-L expression plasmid, 33 ng of the TALE-R expression plasmid, and 33 ng of the EGxxFP reporter plasmid were transfected into HEK293T cells using Lipofectamine 3000 (Thermo Fisher Scientific). When the type or amount of transfected plasmid was reduced, pBluescript II (SK+) (Stratagene, La Jolla, CA, USA) was used to supplement the total amount of transfected plasmid to 100 ng.
[0076] After 48 hours of culture following transfection, the amount of EGFP protein fluorescence per unit area in each well was measured using a plate reader. To reduce variations due to cell localization within the well, fluorescence measurements were taken at 16 points in each well, and the average was used as the fluorescence intensity. The average and standard deviation of the values from three independent wells were plotted.
[0077] 2.Results (1) Confirmation of scFokI activity First, we confirmed the scFokI activity of previously reported molecules as monomeric nucleases lacking base recognition activity. The TALE target sequences used were the Right TALE sequence (tACAGAAGCGGGCAAAGG, corresponding to positions 119-136 of SEQ ID NO: 94) and the Left TALE sequence (tATGTACGCCTCCCTGGG, corresponding to positions 85-102 of SEQ ID NO: 94) of TALENs targeting the human adenomatous polyposis coli (APC) gene, whose intracellular activity has already been confirmed in non-patent literature (Sakuma et al., supra). As a positive control, Platinum TALEN (APC-R) + Platinum TALEN (APC-L) was used, and evaluation was performed using the SSA assay in HEK293T cells.The activity of Platinum TALE (+136 / +63 scaffold, APC-R)-scFokI and Platinum TALE (+136 / +63 scaffold, APC-L)-scFokI was not significantly different from that of the SSA assay reporter alone (Figure 1). Therefore, in hopes of improving activity, we created Platinum TALE(+136 / +63 scaffold, APC-R)-scFokI-12 and Platinum TALE(+136 / +63 scaffold, APC-L)-scFokI-12, in which the linker sequence between FokIs was changed to a flexible GGGGSx12 linker, and evaluated them in the same way. However, we were unable to confirm the same activity as scFokI (Figure 1). scFokI has very low cleavage activity, and it was determined that its use for efficient genome editing in human cells would be difficult.
[0078] (2) Creation of scND1 and confirmation of its activity Next, while following the structure of scFokI, a similar study was carried out on ND1 / ND2, which were found to be alternative factors to FokI.
[0079] Using the +136 / +63 Platinum TALE as a scaffold, expression plasmids for TALE(APC-L)ND1mono, TALE(APC-R)ND1mono, TALE(APC-L)ND1-60-ND1, TALE(APC-L)ND1-95-ND1, TALE(APC-L)ND1-120-ND1, TALE(APC-L)ND1-180-ND1, TALE(APC-L)ND2mono, TALE(APC-R)ND2mono, TALE(APC-L)ND2-60-ND2, TALE(APC-L)ND2-95-ND2, TALE(APC-L)ND2-120-ND2, and TALE(APC-L)ND2-180-ND2 were constructed (Figure 2), and the expression was evaluated by SSA assay in 293T cells. As a result, the activity of the normal TALE-ND1 (TALE(APC-L)ND1mono + TALE(APC-R)ND1mono) was significantly lower than that of the normal TALE-ND2 (TALE(APC-L)ND2mono + TALE(APC-R)ND2mono). However, when they were linked to form single strands, surprisingly, all ND1s showed excellent activity, whereas none of the ND2s showed activity (Figure 3). Furthermore, when the TALE was changed to the +153 / +47 scaffold, TALE-scND1 showed overall reduced activity compared to the +136 / +63 scaffold, while TALE-ND1-95-ND1 retained higher activity (Figure 4).The above findings regarding scND1 were obtained using TALE repeats for APC-L. However, five other target sequences were also identified, specifically, the Right TALE sequence (tCGACATAGTGATTAGGA / corresponding to positions 357-374 of SEQ ID NO: 96) and the Left TALE sequence (tGAACCAGGCTATGACC / corresponding to positions 324-340 of SEQ ID NO: 96) of TALENs designed on the APC-R, human hypoxanthine phosphoribosyl transferase 1 (HPRT1) locus (Sakuma et al., supra), and the Right TALE sequence (tGCCCAGAAGACTCCCG / corresponding to positions 163-179 of SEQ ID NO: 95) and the Left TALE sequence (tGATCTGCAAGTCGAGGC / corresponding to positions 130-147 of SEQ ID NO: 95) of TALENs designed on the human Rosa26 locus (Sato et al., (2015) Stem Cell Reports, 14;5(1), 75-82), TALE-ND1-95-ND1 also showed activity (Figure 5). These results indicate that the 95 amino acid sequence found by HTS is preferable for the linker sequence between ND1s. In subsequent experiments, we mainly used the +153 / +47 scaffold TALE, which exhibits particularly high activity in combination with ND1-95-ND1.
[0080] (3) TALE-scND1 structure optimization Next, we investigated the effect of the C-terminal sequence length of the TALE scaffold +153 / +47 on activity. First, we extended the C-terminal sequence of TALE(Rosa26-L)scND1 and TALE(Rosa26-R)scND1 (to 63, 75, 90, 105, 120, 135, 150, 165, and 180 residues). No change in activity was observed depending on the number of residues (Figure 6). Conversely, when the number of residues was reduced to 36, 24, 12, and 0, activity was sometimes weakened at 0, 12, and 36 residues depending on the target sequence. However, activity at 24 residues was stable and equal to or greater than the original 47 residues (Figure 7).
[0081] In "(2) Construction of scND1 and Confirmation of Activity," we varied the linker length between ND1s to 60, 95, 120, and 180 residues, and found that 95 residues was optimal. However, the 95-residue linker is a unique sequence obtained from HTS, and the other 60, 120, and 180 residues consist of a flexible linker with repeated GGGGS sequences. Therefore, it is unclear whether the 95-residue chain length or the HTS sequence is optimal. Therefore, we constructed three flexible linkers (GSSx32, SAGGx24, and GGGGSx19) and inserted them as linkers for the scND1 used in "(2) Construction of scND1 and Confirmation of Activity" (Figure 8). We compared the activity with that when the HTS95 sequence was used as the linker (Figure 9). As a result, the HTS95 sequence showed good overall activity. When targeting the APC gene, other linkers also showed activity equivalent to that of the HTS95 sequence. [Industrial Applicability]
[0082] The artificial nucleic acid cleaving enzyme using the fusion nuclease domain of the present invention has the activity of site-specifically cleaving double-stranded DNA with a single molecule, and can edit DNA more easily than conventional TALENs, which use a pair of two molecules that bind to each strand of DNA.The artificial nucleic acid cleaving enzyme containing the fusion nuclease domain of the present invention can be used as an excellent genome editing tool in a wide range of industrial fields, such as medicine, agriculture, and industry.
Claims
1. A fusion nuclease domain comprising two nuclease domains 1 linked via a linker, wherein the nuclease domain 1 is a nuclease domain consisting of an amino acid sequence having 90% or more identity to the amino acid sequence shown at positions 391 to 585 of SEQ ID NO:
97.
2. A site-specific DNA cleavage enzyme comprising a DNA binding domain and the fusion nuclease domain of claim 1.
3. A nucleic acid encoding the fusion nuclease domain of claim 1 or the site-specific DNA cleaving enzyme of claim 2.
4. A vector comprising the nucleic acid or its translation product according to claim 3.
5. A method for producing a cell in which target DNA has been modified, comprising introducing into the cell a molecule selected from the following (a) to (c) (provided that the cell is not present in a human individual): (a) the site-specific DNA cleaving enzyme according to claim 2 (b) a nucleic acid encoding the site-specific DNA cleavage enzyme of (a) (c) A vector containing the nucleic acid according to (b) or its translation product.
6. A kit for modifying a target DNA, comprising a molecule selected from the following (a) to (c): (a) the site-specific DNA cleaving enzyme according to claim 2 (b) a nucleic acid encoding the site-specific DNA cleavage enzyme of (a) (c) A vector containing the nucleic acid according to (b) or its translation product.
Citation Information
Patent Citations
Novel nuclease domain and uses thereof
WO2020045281A1