Nicking enzyme, dna editing system, method for editing target dna, and method for producing cells in which target dna has been edited
By using ND1 linker-bound nickase and TALE fusion protein, the safety and versatility issues of DNA editing in existing technologies are resolved, and efficient editing of target DNA other than mitochondrial DNA is achieved.
Patent Information
- Application Number
- CN202480014755.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-28
- Filing Date
- 2024-02-27
- Publication Date
- 2025-09-19
AI Technical Summary
When existing technologies rely on the CRISPR-Cas system for DNA editing, guide RNA needs to be introduced into cells, which poses a safety risk. In addition, the MutH nickase has difficulty editing nuclear DNA and exogenous DNA other than mitochondrial DNA in cells.
A nicking enzyme containing two ND1s and ND1n bound via an ND1 linker is combined with a fusion protein of TALE and a nuclease base transferase. By changing the N-terminus or C-terminus of ND1 to a nuclease activity-deficient mutant sequence, a site-specific nicking enzyme is formed, achieving efficient editing of target DNA other than mitochondrial DNA.
It achieves efficient and specific DNA editing that does not rely on CRISPR-Cas systems such as Cas9, and is able to edit target DNA other than mitochondrial DNA in cells, improving the safety and versatility of editing.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to a nicking enzyme, a DNA editing system using the same, a method for editing a target DNA, and a method for preparing cells in which the target DNA has been edited. Background Art
[0002] In recent years, DNA editing technology using the CRISPR-Cas system, especially the CRISPR-Cas9 system comprising a Cas9 protein (dCas9, nCas9) that has lost some or all of its nuclease activity and its guide RNA, has attracted attention. For example, the Base Editor (BE) system (Komor et al., Nature 533, 2016, p. 420-424 (Non-patent Document 1)) and the Target-AID system (Nishida et al., Science 353, aaf8729, 2016 (Non-patent Document 2)) have been developed so far. In addition, the present inventors have developed a DNA editing technology using the aforementioned CRISPR-Cas9 system and a fusion protein comprising a TALE and a nucleic acid base transferase (such as a deaminase) (International Publication No. 2020 / 050377 Specification (Patent Document 1)). However, these previous technologies rely on the CRISPR-Cas system, and in order to apply them for genome editing, the guide RNA must also be introduced into the cell. Therefore, a technology that can edit the genome only by protein and is safer is desired.
[0003] As a technology that can edit genomes solely using proteins, TALEN (Transcription activator-like effector nuclease), an artificial restriction enzyme, was developed in 2010 as a second-generation genome editing technology. TALEN is a fusion protein that combines the nuclease domain of FokI, a type IIS restriction enzyme (FokI nuclease domain), with a TALE as a DNA binding domain. A pair of TALEN TALEs binds to opposite strands of the target DNA, and the FokI nuclease domains form dimers, thereby exerting site-specific double-strand cleavage activity (nuclease activity) on the DNA. Regarding the FokI nuclease domain, the following reports have been reported so far: two FokI nuclease domains were linked and combined with zinc finger arrays (ZF) and TALE to induce double-stranded cleavage (Minczuk et al., Nucleic Acids Research, 36(12), 2008, p. 3926-3938 (non-patent document 3)); and a structure in which a mutation was added to one of the FokI nuclease domains to induce single-stranded cleavage (nicking) (Yan Luo et al., Scientific Reports, 6: 20657, 2016, doi: 10.1038 / srep20657 (non-patent document 4)).
[0004] In addition, the present inventors have developed two new nuclease domains (nuclease domain 1: ND1, and nuclease domain 2: ND2) that are different from the previous FokI nuclease domain. By using an artificial nucleic acid-cleaving enzyme containing these nuclease domains and DNA binding domains such as ZF and TALE, they successfully edited the target site on the target DNA (International Publication No. 2020 / 045281 (Patent Document 2)).
[0005] Prior art literature
[0006] Patent Literature
[0007] Patent Document 1: International Publication No. 2020 / 050377
[0008] Patent Document 2: International Publication No. 2020 / 045281
[0009] Non-patent literature
[0010] Non-Patent Literature 1: Komor et al., Nature 533, 2016, pp. 420-424
[0011] Non-patent document 2: Nishida et al., Science 353, aaf8729, 2016
[0012] Non-patent document 3: Minczuk et al., Nucleic Acids Research, 36(12), 2008, pp. 3926-3938
[0013] Non-patent literature 4: Yan Luo et al., Scientific Reports, 6: 20657, 2016, doi: 10.1038 / srep20657 Summary of the Invention
[0014] Problems to be solved by the invention
[0015] According to the above-mentioned prior art, it is important to introduce a nick (single-strand cut) into double-stranded DNA for effective DNA editing. However, in the above-mentioned BE system, the formation of the R-loop of the CRISPR-Cas system and the exposure of the single-stranded DNA accompanying it are necessary, and the introduction of the nick remains in an auxiliary role. Therefore, the present inventors believe that by such a mechanism different from the role of the nick in the prior art, that is, the relaxation of the higher-order structure of the genome caused by the introduction of the nick and the generation of partial single-stranded DNA regions, a DNA editing technology that does not rely on the R-loop can be developed.
[0016] Furthermore, for example, Zongyi Yi et al., Nature Biotechnology, https: / / doi.org / 10.1038 / s41587-023-01791-y (Document 1), published on May 22, 2023, describes a technique for editing target sites on mitochondrial DNA using a combination of a fusion protein comprising a nickase such as MutH or Nt.BspD6I(C) and a TALE, and a deaminase. However, MutH requires a specific recognition motif (GATC), and therefore has a problem of lack of versatility. As a result of research, the present inventors found that the technique using Nt.BspD6I(C) described in Document 1 is difficult to edit nuclear DNA other than mitochondrial DNA or exogenous DNA in cells.
[0017] The present invention was completed in view of the problems existing in the above-mentioned prior art, and its purpose is to provide a method for editing target DNA in cells other than mitochondrial DNA specifically and efficiently through nucleobase transferases using only proteins, independent of CRISPR-Cas systems such as Cas9, as well as a nicking enzyme useful therefor.
[0018] Means for solving problems
[0019] To achieve the above-mentioned objectives, the present inventors believe that by using a single molecule of nuclease domain 1 (ND1) as a surrogate factor for the FokI nuclease domain, a highly unique structure capable of introducing an incision can be formed. Thus, the nicking of a site-specific DNA cleavage enzyme (DNA double-stranded cleavage enzyme) developed in international application PCT / JP2022 / 044398, comprising a fusion nuclease domain (scND1) and a DNA binding domain comprising two ND1s bound via a linker, was studied. By changing the N-terminal or C-terminal side of the two ND1s to a nuclease activity-deficient mutant sequence, a site-specific nicking enzyme (nicking enzyme active form) was successfully formed. In addition, based on non-patent document 4 in which the FokI nuclease domain is modified, for example, in the case where ND1 on the N-terminal side is set as a nuclease activity-deficient mutant sequence (for example, in the case where the nuclease activity-deficient mutant sequence is set as TALE→ND1 from the N-terminal side→ND1), it is expected that a nick is introduced into the opposite chain (TALE non-recognition chain) of the chain having the DNA binding domain recognition sequence. However, surprisingly, with respect to the above-mentioned nicking enzyme in which ND1 is modified, it is clear that a nick is introduced into the chain (TALE recognition chain) having the DNA binding domain recognition sequence. It was further discovered that by combining the nicking enzyme and a fusion protein comprising TALE and a nucleic acid base transferase (such as a deaminase), the bases of the target site of the target DNA in the cell other than mitochondrial DNA such as nuclear DNA and exogenous DNA can be efficiently replaced and edited, thereby completing the present invention.
[0020] The aspects of the present invention provided by this finding are as follows. [1]
[0022] An artificial nickase comprising two domains ND1 and ND1n bound via an ND1 linker, wherein
[0023] ND1 is at least one polypeptide selected from the following (a) to (c):
[0024] (a) a polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 98,
[0025] (b) a polypeptide comprising the following amino acid sequence: one or more of the amino acid sequences described in SEQ ID NO: 98 are substituted, deleted, inserted, and / or added, and the amino acids corresponding to positions 66 and 83 of the amino acid sequence described in SEQ ID NO: 98 are aspartic acid,
[0026] (c) a polypeptide comprising an amino acid sequence having 80% or more homology to the amino acid sequence of SEQ ID NO: 98, wherein the amino acids corresponding to amino acids 66 and 83 of the amino acid sequence of SEQ ID NO: 98 are aspartic acid,
[0027] and,
[0028] ND1n is at least one polypeptide selected from the following (an) to (cn):
[0029] (an) a polypeptide comprising the following amino acid sequence: aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted with any other amino acid,
[0030] (bn) A polypeptide comprising the following amino acid sequence: one or more of the amino acid sequences described in SEQ ID NO: 98 are substituted, deleted, inserted, and / or added, and the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted with any amino acid other than aspartic acid;
[0031] (cn) A polypeptide comprising the following amino acid sequence: having more than 80% homology with the amino acid sequence recorded in sequence number 98, and the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence recorded in sequence number 98 is a substituted amino acid substituted by any amino acid other than aspartic acid. [2]
[0033] [1] The nickase described in claim 1, wherein the length of the ND1 linker is 30 to 300 amino acid residues. [3]
[0035] A fusion protein comprising a first DNA binding domain and the nicking enzyme described in [1] or [2]. [4]
[0037] A DNA editing system comprising:
[0038] The first fusion protein as the fusion protein described in [3], and
[0039] A second fusion protein comprises a second DNA binding domain and a nucleobase transferase. [5]
[0041] [4] The DNA editing system described in claim 4, further comprising:
[0042] A third fusion protein comprises a third DNA binding domain and a transcriptional regulator. [6]
[0044] A method for editing target DNA, comprising the steps of:
[0045] The DNA editing system described in [4] or [5] is brought into contact with a target DNA, and the bases at the target site of the target DNA are edited by the activity of the aforementioned nucleobase transferase, wherein:
[0046] The distance between the first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain and the second DNA binding domain recognition sequence or its complementary sequence recognized by the second DNA binding domain is 8 to 48 bases. [7]
[0048] A method for preparing a cell in which target DNA has been edited comprises the following steps:
[0049] The DNA editing system described in [4] or [5] is introduced into a cell or expressed in a cell and brought into contact with a target DNA in the cell other than mitochondrial DNA, and the bases at the target site of the target DNA are edited by the activity of the aforementioned nucleobase transferase, wherein:
[0050] The distance between the first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain and the second DNA binding domain recognition sequence or its complementary sequence recognized by the second DNA binding domain is 8 to 48 bases. [8]
[0052] The method according to [6] or [7], wherein
[0053] The first DNA binding domain recognition sequence or its complementary sequence is present on the 5' side or 3' side of the aforementioned target site via a first spacer of 4 to 16 bases, and
[0054] The second DNA-binding domain recognition sequence or its complementary sequence is present on the opposite side of the target site to the first DNA-binding domain recognition sequence or its complementary sequence via a second spacer of 3 to 31 bases. [9]
[0056] A kit for use in the method described in [6], [7], or [8], comprising: at least one selected from the group consisting of a first fusion protein, a first fusion protein expression vector, and a polynucleotide encoding the first fusion protein, and at least one selected from the group consisting of a second fusion protein, a second fusion protein expression vector, and a polynucleotide encoding the second fusion protein, wherein:
[0057] The first fusion protein expression vector is at least one selected from the group consisting of: (i) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, and a first DNA binding domain, and (ii) a vector comprising a polynucleotide encoding ND1, ND1n, and an ND1 linker and an insertion site for a polynucleotide encoding the first DNA binding domain,
[0058] The second fusion protein expression vector is at least one selected from the following: (iii) a vector comprising a polynucleotide encoding a nucleic acid base transferase and a second DNA binding domain, and (iv) a vector comprising an insertion site for a polynucleotide encoding a nucleic acid base transferase and a polynucleotide encoding a second DNA binding domain.
[10]
[0060] [9] The kit further comprises at least one selected from the following: a third fusion protein comprising a third DNA binding domain and a transcriptional regulatory factor, a third fusion protein expression vector, and a polynucleotide encoding the third fusion protein, wherein:
[0061] The third fusion protein expression vector is at least one selected from the following: (vii) a vector comprising a polynucleotide encoding a transcriptional regulatory factor and a third DNA binding domain, and (viii) a vector comprising an insertion site for a polynucleotide encoding a transcriptional regulatory factor and a polynucleotide encoding a third DNA binding domain.
[11]
[0063] [3] The fusion protein described in further comprising a nucleobase transferase.
[12]
[0065] A method for editing target DNA, comprising the steps of:
[0066] The fusion protein described in
[11] is brought into contact with the target DNA, and the bases at the target site of the target DNA are edited by the activity of the nucleic acid base transferase.
[13]
[0068] A method for preparing a cell in which target DNA has been edited comprises the following steps:
[0069] The fusion protein described in
[11] is introduced into a cell or expressed in a cell and brought into contact with a target DNA in the cell other than mitochondrial DNA, and the bases at the target site of the target DNA are edited by the activity of the aforementioned nucleobase transferase.
[14]
[0071] A kit for use in the method described in
[12] or
[13] , comprising:
[0072] At least one selected from the group consisting of the aforementioned fusion protein, fusion protein expression vector, and a polynucleotide encoding a fusion protein, wherein:
[0073] The aforementioned fusion protein expression vector is at least one selected from the following: (v) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, a nucleic acid base transferase, and a first DNA binding domain, and (vi) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, a nucleic acid base transferase, and an insertion site for a polynucleotide encoding the first DNA binding domain.
[0074] Effects of the Invention
[0075] According to the present invention, a method for editing target DNA in cells other than mitochondrial DNA specifically and efficiently using a nucleobase transferase using only proteins, independent of a CRISPR-Cas system such as Cas9, and a nicking enzyme useful therefor can be provided. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] [ Figure 1 ] shows the results of the SSA assay for double-stranded cleavage activity of a TALE-scFokI comprising two FokI nuclease domains bound via a linker. In addition, "95" in the figure represents a 95-amino acid residue HTS95 linker, and "60" represents a 60-amino acid residue GGGGS×12 linker (the same applies hereinafter).
[0077] [ Figure 2 ] A conceptual diagram showing the structures of TALE-scND1 and TALE-scND2, which contain two nuclease domains (ND1 or ND2) bound via a linker, and TALE-ND1mono and TALE-ND2mono, which contain one nuclease domain. The case where TALE63 is applied as an example of TALE (DNA binding domain) is shown. In addition, "120" in the figure represents a GGGGS×24 linker of 120 amino acid residues, and "180" represents a GGGGS×36 linker of 180 amino acid residues (the same applies below).
[0078] [ Figure 3 Show apps Figure 2 Shown are graphs showing the results of detecting double-strand cleavage activity of TALE-scND1 and TALE-ND1mono (ND1 derivatives), and TALE-scND2 and TALE-ND2mono (ND2 derivatives) using the SSA assay.
[0079] [ Figure 4] A diagram showing the results of detecting the effect of linker length on the double-strand cleavage activity of TALE-scND1 by the SSA test method in the case where TALE47 and TALE63 were used as TALE (DNA binding domain).
[0080] [ Figure 5A ] A diagram showing the results of evaluating the effects of target gene type and linker length on the double-strand cleavage activity of TALE-scND1 by the SSA assay. APC was used as the target gene, and TALE47 was used as the TALE.
[0081] [ Figure 5B ] A diagram showing the results of evaluating the effects of target gene type and linker length on the double-strand cleavage activity of TALE-scND1 using the SSA assay. Rosa26 or HPRT1 was used as the target gene, and TALE47 was used as the TALE.
[0082] [ Figure 6 ] A diagram showing the results of evaluating the effect of lengthening the C-terminal domain of TALE on the double-stranded cleavage activity of TALE-scND1 using the SSA assay. Rosa26 was used as the target gene, TALE47 was used as the TALE, and the HTS95 linker was used as the linker for binding two ND1s.
[0083] [ Figure 7A ] A diagram showing the results of evaluating the effect of shortening the C-terminal domain of TALE on the double-stranded cleavage activity of TALE-scND1 using the SSA assay. Rosa26 or APC was used as the target gene, TALE47 was used as the TALE, and the HTS95 linker was used as the linker for binding two ND1s.
[0084] [ Figure 7B ] A diagram showing the results of evaluating the effect of shortening the C-terminal domain of TALE on the double-stranded cleavage activity of TALE-scND1 using the SSA assay. HPRT1 was used as the target gene, TALE47 was used as the TALE, and the HTS95 linker was used as the linker for binding two ND1s.
[0085] [ Figure 8 ] A conceptual diagram showing the structure of TALE-scND1 comprising two ND1s bound via different types of linkers.
[0086] [ Figure 9 Show apps Figure 8Figure 1 shows the results of detecting double-stranded cleavage activity of TALE-scND1 by the SSA test method. As linkers, HTS95 linker (95 amino acid residues), GSS×32 linker (96 amino acid residues), SAGG×24 linker (96 amino acid residues), and GGGGS×19 linker (95 amino acid residues) were used. As target genes, (A) Rosa26, (B) APC, and (C) HPRT1 were used, and TALE24 was used as a TALE.
[0087] [ Figure 10 ] A conceptual diagram showing the positional relationship between the PAM sequence on each target gene and the cleavage point by Cas9 (D10A), as well as the TALE and guide RNA in (3) (i) of [Test Example 1] <2. Results>. (A) Rosa26, (B) APC, and (C) HPRT1 were used as target genes.
[0088] [ Figure 11A ] Figure showing the results of the SSA assay for DNA cleavage activity of various ND1 mutants of TALE-scND1 in (3)(i) of [Test Example 1] <2. Results>. gRNA-A or gRNA-B designed for the target gene Rosa26 was co-expressed with nCas9 (D10A) to evaluate which strand of the double-stranded DNA was cleaved. TALE12 (Rosa26-L) was used as the TALE.
[0089] [ Figure 11B ] Figure showing the results of the SSA assay for DNA cleavage activity of various ND1 mutants of TALE-scND1 in (3)(i) of [Test Example 1] <2. Results>. gRNA-C or gRNA-D designed for the target gene Rosa26 was co-expressed with nCas9 (D10A) to evaluate which strand of the double-stranded DNA was cleaved. TALE12 (Rosa26-R) was used as the TALE.
[0090] [ Figure 12 ] A figure showing the results of detecting the DNA cleavage activity of various ND1 mutants of TALE-scND1 by the SSA test method in (3) (i) of [Test Example 1] <2. Results>. APC was used as the target gene, and gRNA-A or gRNA-B designed on APC was co-expressed with nCas9 (D10A), and which chain of the double-stranded DNA was cut was evaluated. As a TALE, TALE12 (APC-L) was used. In addition, gRNA-C or gRNA-D designed on APC was co-expressed with nCas9 (D10A), and which chain of the double-stranded DNA was cut was evaluated. As a TALE, TALE12 (APC-R) was used.
[0091] [ Figure 13 ] A diagram showing the results of detecting the DNA cleavage activity of various ND1 mutants of TALE-scND1 by the SSA test method in (3) (i) of [Test Example 1] <2. Results>. HPRT1 was used as the target gene, and gRNA-A or gRNA-B designed on HPRT1 was co-expressed with nCas9 (D10A), and which chain of the double-stranded DNA was cut was evaluated. As a TALE, TALE12 (HPRT1-L) was used. In addition, gRNA-C or gRNA-D designed on HPRT1 was co-expressed with nCas9 (D10A), and which chain of the double-stranded DNA was cut was evaluated. As a TALE, TALE12 (HPRT1-R) was used.
[0092] [ Figure 14 ][Test Example 1]<2. Results> In (3)(ii), (A) a conceptual diagram showing the positional relationship between the PAM sequence, the cleavage point by nCas9, the guide RNA, and the TALE of TALE12(off4-L)-scND1n, and (B) an agarose gel showing the results of the T7E1 test, which was introduced into cells together with Cas9 for the purpose of confirming that the guide RNA (gRNA-off4-L3, gRNA-off4-L4, gRNA-off4-L5) designed on the APC used in the confirmation of nickase activity functions normally. Photographs of gel electrophoresis, (C) TALE12(off4-L)-scND1n was co-expressed with any one of gRNA-off4-L3, gRNA-off4-L4, and gRNA-off4-L5, and any one of nCas9(D10A), nCas9(H840A), and dCas9, and the results of the T7E1 test were shown by introducing these combinations into cells for the purpose of confirming which of nCas9(D10A) and nCas9(H840A) co-existed and thereby producing mutations.
[0093] [ Figure 15 ] Conceptual diagram showing the structures of (A) TALE-AID and (B) ABE8e-TALE in (4) of [Test Example 1] <2. Results>.
[0094] [ Figure 16 ](A) A conceptual diagram showing one method of base editing through collaboration between TALE-scND1n and TALE-AID, (B) A conceptual diagram showing one method of base editing through collaboration between TALE-scND1n and ABE8e-TALE.
[0095] [ Figure 17] Conceptual diagrams showing the structure of the NanoLuc expression reporter used to evaluate the base substitution activity prepared in (12) of [Test Example 1] <1. Method> in (4) of [Test Example 1] <2. Results>. (A) shows one embodiment in which TALE-AID is used as a TALE-deaminase, and (B) shows one embodiment in which ABE8e-TALE is used as a TALE-deaminase.
[0096] [ Figure 18 ] A diagram showing the evaluation results of a reporter assay for confirming the base substitution activity of TALE-scND1n in collaboration with a TALE-deaminase. TALE47(APC-R)-12-AID was used as the TALE-deaminase, and TALE12(Rosa26-L)-scND1n was used as the TALE-scND1n.
[0097] [ Figure 19 ] A graph showing the evaluation results of a reporter test for investigating the optimal length of spacer 2 in the base substitution activity of the collaboration between TALE-scND1n and TALE-AID. A reporter with a spacer 2 length of 7 to 16 bases per base was used. TALE47(APC-R)-12-AID was used as the TALE-AID, and TALE12(Rosa26-LM.2)-scND1n was used as the TALE-scND1n.
[0098] [ Figure 20 ] A graph (heat map) showing the evaluation results of the base substitution activity tested by a reporter when both the length of spacer 1 and the length of spacer 2 are changed in the base substitution activity of the collaboration between TALE-scND1n and TALE-AID. A reporter with a length of 9 to 13 bases per 1 base is used. As TALE12-scND1n, a TALE repeat domain corresponding to the nucleotide sequence of 8 TALE recognition sequences on the reporter: Rosa26-LM3.2, Rosa26-LM2.2, Rosa26-LM1.2, Rosa26-LM.2, Rosa26-LM.2+1, Rosa26-LM.2+2, Rosa26-LM.2+3, Rosa26-LM.2+4 is applied. As a TALE-AID, TALE47(APC-R)-12-AID is applied.
[0099] [ Figure 21] A graph showing the evaluation results of a reporter test to investigate the optimal length of Spacer 2 in the collaborative base substitution activity of TALE-scND1n and ABE8e-TALE. A reporter with a Spacer 2 length of 7 to 16 bases per base was used. ABE8e-32-TALE47 (APC-R) was used as the ABE8e-TALE, and TALE12 (Rosa26-LM1.2)-scND1n was used as the TALE-scND1n.
[0100] [ Figure 22 ] A graph (heat map) showing the evaluation results of the base substitution activity tested by a reporter when both the length of spacer 1 and the length of spacer 2 are varied in the base substitution activity of the collaboration between TALE-scND1n and ABE8e-TALE. A reporter with a length of 9 to 14 bases per spacer 2 was used. As TALE12-scND1n, a TALE repeat domain corresponding to the nucleotide sequence of 9 TALE recognition sequences on the reporter: Rosa26-LM4.2, Rosa26-LM3.2, Rosa26-LM2.2, Rosa26-LM1.2, Rosa26-LM.2, Rosa26-LM.2+1, Rosa26-LM.2+2, Rosa26-LM.2+3, and Rosa26-LM.2+4 was applied. As ABE8e-TALE, ABE8e-32-TALE47 (APC-R) was applied.
[0101] [ Figure 23 ] Graph showing the ratio of T in each target base analyzed by EditR in the base substitution activity on endogenous DNA through the collaboration of TALE-scND1n and TALE-AID. TALE47(APC-R)-12-AID was used as the TALE-AID, and TALE12(APC-L2-1)-scND1n, TALE12(APC-L2-2)-scND1n, and TALE12(APC-L2-3)-scND1n were used as the TALE-scND1n. (A) is a conceptual diagram showing the recognition positions of the TALEs of TALE47(APC-R)-12-AID and three types of TALE-scND1n on APC, as well as the positions of the target bases, and (B) is a graph showing the ratio of T in each C that is the target base.
[0102] [ Figure 24] A conceptual diagram showing the structure of TALE-BspD6I used for nicking activity evaluation using the SSA test of the nickase BspD6I. In the same TALE structure as TALE12-scND1n, BspD6I is linked instead of scND1n, and a TALE repeat domain corresponding to the nucleotide sequence of Rosa26-L and Rosa26-R is inserted.
[0103] [ Figure 25 ] Figure showing the results of detecting the nickase activity of nuclear TALE12-BspD6I by the SSA assay in (1) of [Test Example 2] <2. Results>. Double nicking activity was evaluated using each combination of TALE12(Rosa26-L)-scND1n, TALE12(Rosa26-R)-scND1n, TALE12(Rosa26-L)-BspD6I, TALE12(Rosa26-R)-BspD6I, nCas9(D10A), and guide RNAs (gRNA-A, B, D).
[0104] [ Figure 26 ] Conceptual diagram showing the structures of various structures of TALE-ABE8e linked via different types of linkers.
[0105] [ Figure 27 ] A conceptual diagram showing one method of base editing through collaboration between TALE-scND1n and TALE-ABE8e.
[0106] [ Figure 28 ] A diagram showing the evaluation results of a reporter test for investigating the optimal length of spacer 2 in the collaborative base substitution activity of various structures of TALE-scND1n and TALE-ABE8e in (2) of [Test Example 2] <2. Results>.
[0107] [ Figure 29] A graph (heat map) showing the evaluation results of the base substitution activity tested by the reporter when both the length of spacer 1 and the length of spacer 2 are changed in the base substitution activity of the collaboration between TALE-scND1n and TALE-ABE8e. A reporter with a length of 7 to 16 bases per 1 base is used. As TALE12-scND1n, a TALE repeat domain corresponding to the nucleotide sequence of 8 TALE recognition sequences on the reporter: Rosa26-LM4.2, Rosa26-LM3.2, Rosa26-LM2.2, Rosa26-LM1.2, Rosa26-LM.2, Rosa26-LM.2+1, Rosa26-LM.2+2, Rosa26-LM.2+3 is applied. As TALE-ABE8e, TALE47(APC-R)-12-ABE8e is applied.
[0108] [ Figure 30 ] Conceptual diagram showing the positional relationship between the TALE recognition sequence of TALE-scND1n and the TALE recognition sequence of TALE47-12-AID or TALE-ABE8e at each target site ((A) HEK1, (B) VEGFA3, (C) ABE-site9, (D) BCL11A).
[0109] [ Figure 31 ] Graph showing the ratio of T to each C in each target base at the HEK1 site analyzed by EditR, in the base substitution activity on endogenous DNA by the collaboration of TALE-scND1n and TALE-AID, which have different distances between TALE recognition sequences. TALE47(HEK1-L)-12-AID was used as the TALE-AID, and TALE12(HEK1-R1 to R10)-scND1n was used as the TALE-scND1n.
[0110] [ Figure 32 ] Graph showing the ratio of T to C in each target base at the VEGFA3 site analyzed by EditR, in the base substitution activity on endogenous DNA by the collaboration of TALE-scND1n and TALE-AID, which differ in the distance between TALE recognition sequences. TALE47(VEGFA3-L)-12-AID was used as the TALE-AID, and TALE12(VEGFA3-R1-R11)-scND1n was used as the TALE-scND1n.
[0111] [ Figure 33] Graph showing the ratio of T to C in each target base at the BCL11A site analyzed by EditR, in the base substitution activity on endogenous DNA by the collaboration of TALE-scND1n and TALE-AID, which differ in the distance between TALE recognition sequences. TALE47(BCL11A-L)-12-AID was used as the TALE-AID, and TALE12(BCL11A-R1-R6)-scND1n was used as the TALE-scND1n.
[0112] [ Figure 34 ] A graph showing the ratio of Gs per As at each target base at the ABE-site 9 site analyzed by EditR in base substitutions in endogenous DNA by collaboration between TALE-scND1n and TALE-ABE8e at different inter-TALE distances. TALE47(ABE-site9-L)-12-ABE8e was used as TALE-ABE8e, and TALE12(ABE-site9-R1-R6)-scND1n was used as TALE-scND1n.
[0113] [ Figure 35 ] A graph showing the ratio of Gs per A at each target base at the BCL11A site analyzed by EditR in base substitutions in endogenous DNA by collaboration between TALE-scND1n and TALE-ABE8e at different inter-TALE distances. TALE47(BCL11A-L)-12-ABE8e was used as the TALE-ABE8e, and TALE12(BCL11A-R1-R6)-scND1n was used as the TALE-scND1n.
[0114] [ Figure 36 ] Conceptual diagram showing the positional relationship between the TALE recognition sequence of TALE-scND1n and the TALE recognition sequence of TALE47-12-AID or TALE47-12-ABE8e at each target site ((A) AAVS1-2, (B) APC, (C) MALTA1, (D) RPCI, (E) ABE-site5, (F) HEK2, (G) DYRK1A).
[0115] [ Figure 37] A graph showing the ratio of T in each C of each target base at the AAVS1-2, APC, MALTA1, and RPCI sites analyzed by EditR in the base substitution activity on endogenous DNA through the collaboration of TALE-scND1n and TALE-AID. Respectively, at the AAVS1-2 site, TALE47(AAVS1-2-L)-12-AID was applied as the TALE-AID, and TALE12(AAVS1-2-R1)-scND1n was applied as the TALE-scND1n; at the APC site, TALE47(APC-R3)-12-AID was applied as the TALE-AID, and TALE12(APC-L2-3)-scND1n was applied as the TALE-scND1n; at the MALTA1 site, TALE47(MALTA1-L)-12-AID was applied as the TALE-AID, and TALE12(MALTA1-R1)-scND1n was applied as the TALE-scND1n; at the RPCI site, TALE47(RPCI-L)-12-AID was applied as the TALE-scND1n, and TALE12(RPCI-R1)-scND1n was applied as the TALE-AID.
[0116] [ Figure 38 ] A graph showing the ratio of G in each A of each target base at ABE-site5, HEK2, and DYRK1A sites analyzed by EditR in the base substitution activity on endogenous DNA by the collaboration of TALE-scND1n and TALE-ABE8e. At the ABE-site5 site, TALE47(ABE-site5-L)-12-ABE8e was used as the TALE-ABE8e, TALE12(ABE-site5-R1)-scND1n or TALE12(ABE-site5-R2)-scND1n was used as the TALE-scND1n, and at the HEK2 site, TALE47(HEK2-L)-12-ABE8e was used as the TALE-ABE8e. , as TALE-scND1n, TALE12(HEK2-R1)-scND1n or TALE12(HEK2-R2)-scND1n was used, at the DYRK1A site, TALE47(DYRK1A-L)-12-ABE8e was used as the TALE-ABE8e, and as the TALE-scND1n, TALE12(DYRK1A-R1)-scND1n or TALE12(DYRK1A-R2)-scND1n was used.
[0117] [ Figure 39] A conceptual diagram showing the positional relationship at the APC site between the TALE recognition sequences of TALE-deaminase and TALE-scND1n prepared in (7) and (8) of [Test Example 2] <1. Method> and the TALE recognition sequence of TALE-VPR prepared in (9) of [Test Example 2] <1. Method>.
[0118] [ Figure 40 ] A graph showing the ratio of T to C in each target base at the APC site analyzed by EditR in the base substitution activity on endogenous DNA by the collaboration of TALE-scND1n, TALE-AID, and TALE-VPR. TALE47(APC-R3)-12-AID was used as the TALE-AID, TALE12(APC-L2-3)-scND1n was used as the TALE-scND1n, and TALE47(APC-TA-L1~L3))-12-VPR was used as the TALE-VPR.
[0119] [ Figure 41 ] A graph showing the ratio of Gs per A at each target base at the APC site analyzed by EditR in the base substitution activity of TALE-scND1n, TALE-ABE8e, and TALE-VPR on endogenous DNA. TALE47(APC-R3)-12-ABE8e was used as the TALE-ABE8e, TALE12(APC-L2-3)-scND1n was used as the TALE-scND1n, and TALE47(APC-TA-L1-L3))-12-VPR was used as the TALE-VPR. DETAILED DESCRIPTION
[0120] Hereinafter, the present invention will be described in more detail by taking preferred embodiments of the present invention as examples, but the present invention is not limited thereto.
[0121] <Nickase>
[0122] The nickase of the present invention is an artificial nickase comprising two domains, ND1 and ND1n, bound via an ND1 linker (referred to herein as "single-stranded ND1n" or "scND1n" depending on the context). The nickase of the present invention has the activity of cleaving only one strand of double-stranded DNA (nickase activity).
[0123] (ND1)
[0124] The "ND1" of the present invention is a nuclease domain 1 (Patent Document 2), one of the nuclease domains discovered by the present inventors through screening of homologous sequences with an identity of 35% to 70% with the FokI nuclease domain. The amino acid sequence of the full-length protein comprising ND1 (as a representative example, derived from Bacillus SGD-V-76) is shown in SEQ ID NO: 97. ND1 is typically (a) a polypeptide comprising the amino acid sequence described in SEQ ID NO: 98 (amino acid residues 391 to 585 of the amino acid sequence of SEQ ID NO: 97), more preferably a polypeptide consisting of the amino acid sequence described in SEQ ID NO: 98. In addition, the amino acid sequence described in SEQ ID NO: 98 has 70% identity with the amino acid sequence of the FokI nuclease domain.
[0125] The "ND1" of the present invention, when two molecules are bound via an ND1 linker, comprises a nuclease domain consisting of an amino acid sequence having high homology to the amino acid sequence described in SEQ ID NO: 98, as long as it has double-strand DNA cleavage activity (nuclease activity). Such a nuclease domain includes, for example, ND1 derived from other bacteria and variants of ND1 (natural mutants and artificial mutants).
[0126] Therefore, the embodiment of "ND1" according to the present invention also includes (b) an amino acid sequence in which one or more amino acids are substituted, deleted, inserted, and / or added in the amino acid sequence described in SEQ ID NO: 98. However, the nuclease activity-deficient mutant sequence described below is not essential, and therefore, in (b), it is essential that at least the amino acids corresponding to positions 66 and 83 of the amino acid sequence described in SEQ ID NO: 98 are aspartic acid.
[0127] Here, in the amino acid sequence, "substituted, deleted, inserted, and / or added amino acid sequence" means an amino acid sequence in which the amino acids (amino acid residues) in the amino acid sequence are substituted, deleted, inserted, or added, or an amino acid sequence in which a combination of two or more of these is performed. In addition, "plurality" means an integer of 30, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2. In the amino acid sequence of the polypeptide of (b), as one or more, preferably the amino acid residues are 1 to 30, 1 to 20 (for example, 1 to 10, 1 to 5, 1 to 3, or 2 or less).
[0128] Furthermore, an embodiment of "ND1" according to the present invention also includes (c) an amino acid sequence having 80% or more identity with the amino acid sequence described in SEQ ID NO: 98. However, in (c), it is also necessary that it is not a nuclease activity-deficient mutant sequence described below, and therefore it is necessary that at least the amino acids corresponding to positions 66 and 83 of the amino acid sequence described in SEQ ID NO: 98 are aspartic acid.
[0129] Here, in the context of amino acid sequence homology, when a reference amino acid sequence is aligned with a target amino acid sequence using amino acid sequence analysis software, the amino acid in the target amino acid sequence (target amino acid) at the same position as the amino acid in the reference amino acid sequence (reference amino acid) can be the same amino acid as the reference amino acid or an amino acid having the same properties as the reference amino acid. In the context of amino acid sequence identity, the target amino acid is the same amino acid as the reference amino acid. Groups of amino acids having the same properties are well known in the art to which the present invention pertains, and include, for example, acidic amino acids (aspartic acid and glutamic acid); basic amino acids (lysine / arginine / histidine); and neutral amino acids, which can be categorized by hydrocarbon chain-containing amino acids (glycine / alanine / valine / leucine / isoleucine / proline), hydroxyl group-containing amino acids (serine / threonine), sulfur-containing amino acids (cysteine / methionine), amide group-containing amino acids (asparagine / glutamine), imino group-containing amino acids (proline), and aromatic group-containing amino acids (phenylalanine / tyrosine / tryptophan).
[0130] The homology and identity of such amino acid sequences are determined by comparing two sequences aligned in a state where the consistency of the sequence becomes the maximum. The method for obtaining the numerical value (%) of sequence homology or identity is well known to those skilled in the art. As an algorithm for obtaining the most suitable comparison and sequence identity, any algorithm known to those skilled in the art (e.g., BLAST algorithm, FASTA algorithm, etc.) can be utilized. The sequence homology or identity of the amino acid sequence, for example, can be determined using sequence analysis software such as BLASTP, FASTA. In addition, in the amino acid sequence of the polypeptide of (c), the homology with the amino acid sequence of (a) can be as long as it is more than 80%, preferably more than 85%, more than 90%, more than 95% (e.g., more than 96%, more than 97%, more than 98%, more than 99%). More preferably, with identity, it is more than 80%, more than 85%, more than 90%, more than 95% (e.g., more than 96%, more than 97%, more than 98%, more than 99%).
[0131] In addition, in an amino acid sequence, the amino acid "corresponding to" a specific amino acid means an amino acid that is at the same position as the aforementioned specific amino acid (the aforementioned control amino acid (amino acid residue), for example, aspartic acid at position 66 or 83 of the amino acid sequence recorded in SEQ ID NO: 98) when the amino acid sequences are aligned using amino acid sequence analysis software (for example, GENETYX-MAC, Sequencher, ClustalW, etc.) (for example, parameters: default values (i.e., initial setting values)).
[0132] (ND1n)
[0133] The "ND1n" of the present invention is an altered domain in which the amino acid sequence of ND1 is altered to a mutant sequence (nuclease-deficient mutant sequence) in a manner that impairs nuclease activity. Such nuclease-deficient mutant sequences of ND1n are not particularly limited as long as their double-stranded DNA cleavage activity is impaired when bound to ND1 via an ND1 linker. Typically, they are polypeptides comprising the following amino acid sequence: Aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted with any other amino acid, more preferably a polypeptide comprising the aforementioned amino acid sequence containing the aforementioned substituted amino acids. The substituted aspartic acid in the aforementioned substituted amino acids may be either or both aspartic acid at position 66 and aspartic acid at position 83, preferably containing at least aspartic acid at position 66.
[0134] Furthermore, the "ND1n" of the present invention also includes variants consisting of an amino acid sequence having high homology to such (an). Therefore, the embodiments of the "ND1n" of the present invention also include: (bn) a polypeptide comprising the following amino acid sequence: one or more of the amino acid sequences described in SEQ ID NO: 98 are substituted, deleted, inserted, and / or added, and the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted amino acid by any amino acid other than aspartic acid; and (cn) a polypeptide comprising the following amino acid sequence: having 80% or more homology to the amino acid sequence described in SEQ ID NO: 98, and the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted amino acid by any amino acid other than aspartic acid.
[0135] In the amino acid sequence of the polypeptide of (bn), the number of one or more amino acid residues is preferably 1 to 30, 1 to 20 (e.g., 1 to 10, 1 to 5, 1 to 3, or 2 or less). Furthermore, in the amino acid sequence of the polypeptide of (cn), the homology with the amino acid sequence of (an) is sufficient as long as it is 80% or more, preferably 85% or more, 90% or more, or 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more). More preferably, the identity is 80% or more, 85% or more, 90% or more, or 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more).
[0136] In (bn) and (cn), the substituted amino acid substituted with any amino acid other than aspartic acid is preferably an amino acid other than the aforementioned acidic amino acids, more preferably alanine or asparagine. Furthermore, the aspartic acid substituted with the aforementioned substituted amino acid may be either or both of the amino acid corresponding to aspartic acid at position 66 and the amino acid corresponding to aspartic acid at position 83, but preferably includes at least the amino acid corresponding to aspartic acid at position 66.
[0137] As a method for substituting the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 with the aforementioned substituted amino acid, conventionally known methods or methods based thereon can be appropriately employed. For example, conventionally known methods such as site-specific mutagenesis (Kunkel et al. (Proc. Natl. Acad. Sci. USA (1985) 82, pp. 488-492)) and repeat extension PCR can be appropriately employed. Furthermore, primers and the like used for the aforementioned amino acid modification can be appropriately designed based on the target amino acid sequence (such as the amino acid sequence described in SEQ ID NO: 98) and the desired modification by conventionally known methods or methods based thereon.
[0138] In addition, in the present invention, whether ND1 and ND1n have double-stranded cutting activity (nuclease activity) of DNA when bound to ND1 via the ND1 linker can be appropriately confirmed by methods known to those skilled in the art. For example, as described in the following test example, the fusion body in which "the following TALE→the structural domain of the object→the following ND1 linker→the amino acid sequence described in sequence number 98 (ND1)" is bound in this order from the N-terminal side is expressed in the cell, and it can be confirmed by whether the target DNA corresponding to the TALE is double-stranded cut. In this method, an SSA test system with a reporter gene on a plasmid (for example, an EGFP gene with a TALE recognition sequence inserted) as the target is used, and evaluation can be performed using reporter activity as an indicator. For example, as long as the nuclease activity in the case of applying the aforementioned fusion is more than 50% of the nuclease activity in the case where "TALE→amino acid sequence (ND1) recorded in sequence number 98→ND1 linker→amino acid sequence (ND1) recorded in sequence number 98" are combined and expressed in this order from the N-terminal side, it can be judged that the nuclease activity is present; as long as it is less than 50%, it can be judged that the nuclease activity is absent.
[0139] (ND1 connector)
[0140] The "ND1 linker" of the present invention is a polypeptide that functions to link the two domains ND1 and ND1n described above. The length and type of such an ND1 linker are not particularly limited, as long as the nickase of the present invention can exert nickase activity. Typically, the length is 30 to 300 amino acid residues, preferably 50 to 200 amino acid residues, and more preferably 70 to 100 amino acid residues.
[0141] The type of ND1 linker involved in the present invention is not particularly limited as long as the nickase of the present invention can exert nickase activity. For example, an HTS95 linker composed of the HTS95 amino acid sequence described in Sun, N., & Zhao, H. (2014) Molecular BioSystems, 10(3), pp. 446-453 (amino acid residues 197 to 292 of the amino acid sequence of SEQ ID NO: 61), a GSS linker composed of GSS or its repeats, a SAGG linker composed of SAGG (SEQ ID NO: 183) or its repeats, a GGGGS linker composed of GGGGS (SEQ ID NO: 99) or its repeats, an XTEN linker (SEQ ID NO: 184), etc. are mentioned, and the HTS95 linker is preferred.
[0142] In the nicking enzyme of the present invention, ND1 and ND1n are linked via an ND1 linker, either on the N-terminal side or the C-terminal side. However, in the case of a fusion protein with a TALE, which is a DNA-binding protein described below, it is preferred to bind in the order of ND1n → ND1 linker → ND1 from the N-terminal side, particularly from the perspective of exhibiting excellent specificity for the DNA chain, i.e., specificity for the TALE recognition chain.
[0143] In the nickases of the present invention, ND1 and ND1n can be bound to the ND1 linker at both the nucleic acid and amino acid levels. Specifically, a polynucleotide encoding the nickase of the present invention can be prepared by linking the respective polynucleotides via a ligation reaction in the order of "polynucleotide encoding ND1n → polynucleotide encoding ND1 linker → polynucleotide encoding ND1" or "polynucleotide encoding ND1 → polynucleotide encoding ND1 linker → polynucleotide encoding ND1n." By inserting the resulting polynucleotide into an expression vector and expressing it in a suitable host cell, the nickase of the present invention can be obtained, bound at the amino acid level in the order of "ND1n → ND1 linker → ND1" or "ND1 → ND1 linker → ND1n."
[0144] As the aforementioned expression vector, it can be suitably selected from the vectors used in this field, for example, plasmid vectors, viral vectors, phage vectors, phagemid vectors, BAC vectors, YAC vectors, MAC vectors, and HAC vectors are enumerated. As the host cell for importing the aforementioned expression vector, it can be considered that the compatibility with the expression vector is suitably selected, for example, cells of prokaryotes such as Escherichia coli, actinomycetes, and archaebacteria and eukaryotic organisms such as yeast, sea urchin, silkworm, zebrafish, mouse, rat, frog, tobacco, Arabidopsis, and rice are enumerated. In addition, "polynucleotide" comprises any one of DNA and RNA (mRNA), and in each polynucleotide, codon optimization can be performed to improve the expression efficiency in the cell.
[0145] In addition, the nickase of the present invention can be artificially synthesized and formulated based on the amino acid sequence information. Furthermore, the nickase of the present invention can be added with epitope tags (e.g., Flag tags, HA tags, etc.) for purification and detection, and various transfer signals (e.g., nuclear transfer signals, plastid transfer signals, etc.).
[0146] <First fusion protein>
[0147] The fusion protein of the present invention is a fusion protein comprising a first DNA binding domain and the nickase of the present invention (hereinafter referred to as "first fusion protein" as appropriate). The first fusion protein of the present invention binds to the vicinity of the target site on the target DNA via the first DNA binding domain (first DNA binding domain recognition sequence), and cleaves one strand of the double-stranded DNA (introducing a nick) through the nickase, thereby functioning as a site-specific nickase. The aforementioned nickase includes preferred embodiments thereof as described above.
[0148] (First DNA binding domain)
[0149] The "first DNA binding domain" of the present invention is not particularly limited as long as it is a protein domain that can specifically bind to any DNA sequence (first DNA binding domain recognition sequence). Examples of the first DNA binding domain include TALE, zinc finger array (ZF), and PPR protein (Pentatricopeptide Repeat Protein), preferably at least one selected from TALE and zinc finger array (ZF), more preferably TALE.
[0150] [TALE]
[0151] TALEs (Transcription activator-like effector nucleases) are proteins typically secreted by Proteobacteria of the genus Xanthomonas that activate gene transcription in host plants. It is a general term for TALE-like proteins found in bacteria of the genus Ralstonia. The "TALEs" of the present invention comprise at least an N-terminal domain and a TALE repeat domain, and may further include a C-terminal domain.
[0152] The TALE repeat domain is composed of multiple, for example 10 to 30, preferably 13 to 25, and more preferably 15 to 20 tandem repeats (TALE repeats) of the TALE sequence that forms a right-handed superhelix. A typical TALE repeat unit (one TALE sequence) consists of 33 to 35 amino acids, and recognizes specific bases of DNA through variable residues (repeat variable diresidue: RVD) composed of two amino acid residues at positions 12 and 13. Examples of RVDs that specifically recognize bases include HD that recognizes C, NG that recognizes T, NI that recognizes A, NN that recognizes G or A, and NS that recognizes A, C, G, or T. Based on the DNA recognition mechanism of the TALE repeat domain, by artificially linking TALE sequences that recognize specific bases, a TALE that can recognize and bind to a desired nucleotide sequence (TALE recognition sequence) on DNA can be prepared.
[0153] The aforementioned TALE sequences of the TALEs involved in the present invention can be appropriately modified for the natural amino acid sequence as long as the TALE repeat domain can recognize and bind to the TALE recognition sequence. For example, in the aforementioned TALE sequences, the position of the RVD can be changed, and one or more (for example, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, further preferably 3 or less or 2 or less) amino acid residues other than the RVD can be substituted, deleted, inserted and / or added. In addition, the two amino acid residues at positions 12 and 13 of the RVD can be substituted with other amino acid residues to enhance the specificity for A, T, C and G.
[0154] The TALE repeat domain of the TALE involved in the present invention is designed to recognize a nucleotide sequence or its complementary sequence that is present on the 5' side or 3' side of the target site of the target DNA via a first spacer of 4 to 16 bases in length when the following DNA editing system is applied. That is, the TALE recognition sequence or its complementary sequence recognized by the TALE is designed to be a nucleotide sequence that is present on the 5' side or 3' side of the aforementioned target site via a first spacer of 4 to 16 bases in length. Such techniques for designing and preparing desired TALEs are well known. For example, TALEs having high binding activity to the aforementioned nucleotide sequences can be prepared using the Platinum Gate system described in Miller et al., Nat Biotechnol 29, 2011, p. 143-148; Sakuma et al., Sci Rep 3, 3379 (2013), and the aforementioned Sakuma et al. (2013).
[0155] As the N-terminal domain of the TALE involved in the present invention, a natural amino acid sequence (for example, the amino acid sequence of the N-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185~32188)) can be applied. As long as it does not have a negative impact on the function of the first fusion protein of the present invention (the ability to bind to the TALE recognition sequence, etc.), the aforementioned natural amino acid sequence can be appropriately changed. For example, one or more (for example, less than 50, less than 30, less than 20, less than 10, less than 9, less than 8, less than 7, less than 6, preferably less than 5 or less than 4, further preferably less than 3 or less than 2) amino acid residues can be substituted, deleted, inserted and / or added. In addition, it also includes a Flag tag for purification and detection, a nuclear localization signal (NLS) for transferring the TALE-nickase to the cell nucleus, etc.
[0156] The chain length of such an N-terminal domain is preferably 49 to 287 amino acid residues, more preferably 80 to 200 amino acid residues, and even more preferably 120 to 180 amino acid residues.
[0157] As specific examples of the N-terminal domain of the TALE involved in the present invention, in addition to the above, for example, the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-136 / 63-VR-HD (ID: 50699) and the amino acid sequence of the N-terminal domain contained in Addgene's ptCMV-153 / 47-VR-HD (ID: 50703) are listed, but are not limited to these.
[0158] In the case where the TALE involved in the present invention further includes a C-terminal domain, a natural amino acid sequence (e.g., the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185-32188) (WT: number of amino acid residues = 180)) can be applied as the C-terminal domain of the TALE. As long as it does not negatively affect the function of the first fusion protein of the present invention (such as the ability to bind to the TALE recognition sequence), the natural amino acid sequence can be appropriately changed. For example, one or more (e.g., 50 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, more preferably 3 or less or 2 or less) amino acid residues may be substituted, deleted, inserted and / or added.
[0159] When the TALE of the present invention further comprises a C-terminal domain, the C-terminal domain of the TALE may be a C-terminal domain in which a portion of the amino acid sequence on the C-terminal side of the native amino acid sequence is removed. Such a C-terminal domain preferably has a chain length of 1 to 200 amino acid residues, more preferably 10 to 200 amino acid residues, even more preferably 10 to 190 amino acid residues, and even more preferably 20 to 180 amino acid residues.
[0160] As specific examples of the aforementioned C-terminal domain, in addition to the above, for example, the amino acid sequence of the C-terminal domain contained in Addgene's pTALEN_v2 (ID: 32189~32192) (amino acid residue number = 63) and the amino acid sequence of the C-terminal domain contained in Addgene's ptCMV-153 / 47-VR-NG (ID: 50704) (amino acid residue number = 47) are listed, but are not limited to these.
[0161] As a preferred embodiment of the TALE according to the present invention, a preferred embodiment of the TALE is Platinum TALEN (Japanese Patent Application Laid-Open No. 2015-33365), in which the amino acids at two specific positions of one TALE repeat unit are changed four times per TALE repeat unit.
[0162] [Zinc finger array]
[0163] The zinc finger arrays described herein represent the DNA-binding region of zinc finger nucleases (ZFNs), which comprise a DNA-binding region and a nuclease domain. Zinc finger arrays typically consist of 2 to 15, preferably 3 to 8, and more preferably 4 to 6, zinc finger proteins, each of which is bound directly or via a linker. Each zinc finger protein contains one or more zinc atoms and a helical structure that recognizes specific bases, with each zinc finger protein recognizing three specific bases on DNA.
[0164] The zinc finger arrays of the present invention are not particularly limited as long as they can recognize and bind to the DNA binding domain recognition sequence, and can be modified appropriately. For example, one or more (e.g., 50 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, preferably 5 or less or 4 or less, more preferably 3 or less or 2 or less) amino acid residues can be substituted, deleted, inserted, and / or added.
[0165] The zinc finger array of the present invention is designed to recognize a nucleotide sequence or its complementary sequence that is present on the 5' side or 3' side of the target site of the target DNA via a first spacer with a chain length of 4 to 16 bases, or a complementary sequence thereof, when used in the following DNA editing system. That is, the zinc finger array recognition sequence or its complementary sequence recognized by the zinc finger array is designed to be a nucleotide sequence that is present on the 5' side or 3' side of the aforementioned target site via a first spacer with a chain length of 4 to 16 bases. Such techniques for designing and preparing zinc finger arrays are well known and can be appropriately designed and prepared, for example, by referring to the methods described in Beerli et al., Nat Biotechnol 20, 135-141 (2002), Sander et al., Nat Methods 8, 67-69 (2011), etc.
[0166] (First connector)
[0167] In the first fusion protein of the present invention, the nickase and the first DNA binding domain can be directly bound or can be bound via a linker. In the case where a linker is present, it will be referred to as the "first linker" according to the circumstances. As the first linker, as long as it does not inhibit the effect of the present invention (binding ability to the first DNA binding domain recognition sequence, nickase activity of the nickase, etc.), there is no particular limitation in its chain length or type. As the chain length of the first linker, it is generally 2 to 180 amino acid residues, preferably 2 to 120 amino acid residues. For example, in the case where the aforementioned first DNA binding domain is TALE, the chain length of the first linker can be adjusted so that the distance between the TALE repeat domain and the aforementioned nickase is 1 to 200 amino acid residues, preferably 10 to 200 amino acid residues, more preferably 10 to 190 amino acid residues, and further preferably 20 to 180 amino acid residues, including the N-terminal domain of TALE (from the N-terminal side, in the order of nickase → first DNA binding domain) or the C-terminal domain (from the N-terminal side, in the order of first DNA binding domain → nickase), more preferably the chain length of the C-terminal domain. In this case, the distance between the TALE repeat domain and the aforementioned nickase is the number of amino acid residues from the amino acid residue adjacent to the C-terminal side of the C-terminus of the TALE repeat domain as the first residue to the residue adjacent to the N-terminal side of the nickase in the order from the N-terminal side, first DNA-binding domain → nickase, and from the amino acid residue adjacent to the N-terminal side of the TALE repeat domain as the first residue to the residue adjacent to the C-terminal side of the nickase in the order from the N-terminal side, nickase → first DNA-binding domain. The type of the first linker according to the present invention is not particularly limited, and for example, the ND1 linker can be appropriately used.
[0168] In the first fusion protein, the nickase and the first DNA binding domain are connected via the first linker as needed, either on the N-terminal side or on the C-terminal side. In addition, the first DNA binding domain can be configured at both ends of the nickase to form a sandwich type (for sandwich type, refer to Mori et al (2009) Biochemical and Biophysical Research Communications, 390 (3), 694-697). Among them, as the first fusion protein of the present invention, in the case of combining with the second fusion protein described below, in particular, from the perspective of showing specificity for the DNA chain, that is, excellent specificity for the TALE recognition chain, it is preferably combined from the N-terminal side in the order of the first DNA binding domain → (first linker as needed) → nickase.
[0169] In the first fusion protein, the binding of the nickase to the first DNA binding domain can be carried out at the nucleic acid level and the amino acid level. That is, by linking the respective polynucleotides through a ligation reaction in the order of "polynucleotide encoding nickase → polynucleotide encoding the first DNA binding domain" or "polynucleotide encoding the first DNA binding domain → polynucleotide encoding nickase", a polynucleotide encoding the first fusion protein of the present invention can be prepared. By inserting the polynucleotide obtained in this way into an expression vector and expressing it in an appropriate host cell, the first fusion protein of the present invention that is bound at the amino acid level in the order of "nickase → first DNA binding domain" or "first DNA binding domain → nickase" can be obtained. The aforementioned expression vector and host cell are the same as those listed for the nickase of the present invention.
[0170] In addition, the first fusion protein of the present invention can be artificially synthesized based on the amino acid sequence information. Furthermore, the first fusion protein of the present invention can be supplemented with epitope tags (e.g., Flag tags, HA tags, etc.) for purification and detection, as well as various translocation signals (e.g., nuclear translocation signals, chromatin translocation signals, etc.).
[0171] DNA Editing System
[0172] The DNA editing system of the present invention comprises the first fusion protein of the present invention described above, and a second fusion protein containing a second DNA binding domain and a nucleic acid base transferase. As the first fusion protein, its preferred embodiment is also included, as described above. In addition, as the DNA editing system of the present invention, from the perspective of promoting base conversion by the collaboration of the first fusion protein and the second fusion protein, it is preferably further included a third fusion protein containing a third DNA binding domain and a transcriptional regulator.
[0173] (Second fusion protein)
[0174] The second fusion protein of the present invention is a fusion protein comprising a second DNA-binding domain and a nucleobase transferase. The second DNA-binding domain of the present invention also includes preferred embodiments thereof, such as those described for the first DNA-binding domain. When the first DNA-binding domain is a TALE, the second DNA-binding domain is preferably at least one selected from a TALE and a zinc finger array, more preferably a TALE.
[0175] [Nucleic acid base transferase]
[0176] In the present invention, a nucleobase transferase means an enzyme that can convert a target base into another base without cleaving the DNA chain by catalyzing a reaction that converts a substituent on a purine or pyrimidine ring of a DNA base into another group or atom, or causes it to be removed.
[0177] As such a nucleic acid base transfer enzyme, there is no particular limitation as long as it can catalyze the above-mentioned reaction, and for example, deaminases and glycosylases are listed, and it can be only one of them or a combination of two or more thereof. Among them, as the nucleic acid base transfer enzyme related to the present invention, deaminases are preferred.
[0178] The aforementioned deaminases are enzymes that catalyze deamination reactions that convert the amino group of a base to a carboxyl group and belong to the nucleic acid / nucleotide deaminase superfamily. Examples of such deaminases include cytidine deaminases that can replace cytosine or 5-methylcytosine with uracil or thymine, respectively; adenosine deaminases that can replace adenine with hypoxanthine; and guanosine deaminases that can replace guanine with xanthine. These deaminases can be used separately depending on the target base substitution.
[0179] The deaminases are not particularly limited in their origin, and examples include hagfish; mammals such as humans, monkeys, pigs, cattle, horses, rats, and mice. Examples of such deaminases include cytidine deaminases such as APOBEC (rAPOBEC1 from rat, hAPOBEC1, hAPOBEC2, hAPOBEC3 (hAPOBEC3A, 3B, 3C, 3D (3E), 3F, 3G, 3H), and hAPOBEC4 from humans); Anc689, which represents the ancestral amino acid sequence of APOBEC; AID (Activation-induced cytidine deaminase (AICDA)) from mammals (e.g., humans, pigs, cattle, horses, and monkeys); and PmCDA1 (Petromyzon marinus cytosine deaminase 1) from hagfish, which belongs to the AID family. Furthermore, examples of adenosine deaminases include TadA from Escherichia coli. Each of the above-mentioned deaminases includes variants thereof (e.g., TadA-8e, TadA7.10, and variants thereof). Among them, the deaminase is preferably at least one selected from APOBEC, PmCDA1, Anc689, and TadA (including variants thereof).
[0180] Furthermore, when the second fusion protein of the present invention comprises two or more deaminases, their combination can be appropriately selected depending on the target base substitution. Examples include combinations of the same or different cytidine deaminases, combinations of the same or different adenosine deaminases, combinations of the same or different guanosine deaminases, combinations of a cytidine deaminase and an adenosine deaminase, combinations of a cytidine deaminase and a guanosine deaminase, and combinations of an adenosine deaminase and a guanosine deaminase. Among these, the combination of deaminases is preferably a combination of TadA and PmCDA1.
[0181] The nucleotide sequences and amino acid sequences of these deaminases are known and can be obtained from public databases (Genbank, etc.). Furthermore, as long as the deaminases have deamination activity (deaminase activity), appropriate modifications (e.g., introduction of substitutions, deletions, insertions, and / or additions of amino acid residues) can be made based on the nucleotide sequences and amino acid sequences of these known deaminases.
[0182] [Second connector]
[0183] In the second fusion protein involved in the present invention, the nucleic acid base transferase and the second DNA binding domain can be directly bound or can be bound via a linker. In the case where a linker is present, it will be referred to as the "second linker" according to the situation in this specification. As a second linker, as long as the effect of the present invention (binding ability to the second DNA binding domain recognition sequence, nucleic acid base conversion activity, nickase activity of the first fusion protein, etc.) is not inhibited, there is no particular limitation on its chain length or type. As an example of such a second linker, its preferred embodiment is also included, and the linker 1 described in the specification of International Publication No. 2022 / 050377 is listed. For example, in the case where the second DNA binding domain is TALE, in the second fusion protein, the chain length of the second linker can be adjusted so that the distance between the TALE repeat domain and the aforementioned nucleic acid base transferase is 1 to 630 amino acid residues, with the distance also including the chain length of the N-terminal domain or C-terminal domain (more preferably the C-terminal domain) of TALE. In this case, the distance between the TALE repeat domain and the aforementioned nucleobase transferase (the number of amino acid residues from the amino acid residue adjacent to the C-terminal side of the C-terminus of the TALE repeat domain (or the amino acid residue adjacent to the N-terminal side of the N-terminus) as the first residue to the residue adjacent to the N-terminal side of the N-terminus of the nucleobase transferase (or the residue adjacent to the C-terminal side of the C-terminus)) is more preferably 10 to 500 amino acid residues, further preferably 25 to 300 amino acid residues, and further more preferably 52 to 174 amino acid residues.
[0184] In addition, as the second fusion protein, it is possible to include two or more (e.g., two) nucleic acid base transfer enzymes. In this case, the nucleic acid base transfer enzymes may be only one or a combination of two or more. For example, in the case where the second fusion protein of the present invention includes two nucleic acid base transfer enzymes (a first nucleic acid base transfer enzyme, a second nucleic acid base transfer enzyme), they may be combined in the order of a first nucleic acid base transfer enzyme, a second DNA binding domain, and a second nucleic acid base transfer enzyme. Among them, as the second fusion protein of the present invention, it is preferred to include one nucleic acid base transfer enzyme, preferably from the N-terminal side, in the order of the second DNA binding domain → (second linker as needed) → nucleic acid base transfer enzyme, or from the N-terminal side, in the order of nucleic acid base transfer enzyme → (second linker as needed) → second DNA binding domain, more preferably from the N-terminal side, in the order of the second DNA binding domain → (second linker as needed) → nucleic acid base transfer enzyme.
[0185] Furthermore, as the second fusion protein of the present invention, it is preferable to include a base removal repair inhibitor bound via a linker at the C-terminal side, i.e., the C-terminal end of the aforementioned nucleic acid base transfer enzyme (the aforementioned nucleic acid base transfer enzyme is located on the C-terminal side of the second DNA binding domain) or the C-terminal end of the second DNA binding domain (the aforementioned nucleic acid base transfer enzyme is located only on the N-terminal side of the second DNA binding domain). As examples of such linkers and base removal repair inhibitors, including preferred embodiments thereof, the linker 2 and base removal repair inhibitor described in the specification of International Publication No. 2022 / 050377 are listed.
[0186] In the second fusion protein, the binding of the nucleobase transferase to the second DNA binding domain can be performed at the nucleic acid level and the amino acid level, and the method is the same as the binding of the nickase to the first DNA binding domain in the first fusion protein described above.
[0187] In addition, the second fusion protein can be artificially synthesized based on the amino acid sequence information. Furthermore, as the first fusion protein of the present invention, epitope tags (e.g., Flag tags, HA tags, etc.) for purification and detection, and various transfer signals (e.g., nuclear transfer signals, chromatin transfer signals, etc.) can be added.
[0188] (Third fusion protein)
[0189] The third fusion protein of the present invention is a fusion protein comprising a third DNA-binding domain and a transcriptional regulator. The third DNA-binding domain of the present invention also includes preferred embodiments thereof, such as those described for the first DNA-binding domain. When the first DNA-binding domain is a TALE, the third DNA-binding domain is preferably at least one selected from a TALE and a zinc finger array, more preferably a TALE.
[0190] [Transcription regulator]
[0191] In the present invention, a transcriptional regulatory factor refers to a factor that binds to an enhancer on DNA and regulates its transcription, and more preferably a transcriptional activator that promotes transcription. Examples of such transcriptional regulatory factors include VPR, Rta, p65, Hsf1, TCF4, MEF2A, MEF2C, MEF2D, p53, E2F1, VP16, VP64, VP128, and VP160, and may be a single species or a combination of two or more. The nucleotide and amino acid sequences of these transcriptional regulatory factors are known and can be obtained from public databases (such as Genbank).
[0192] [Third connector]
[0193] In the third fusion protein involved in the present invention, the transcription regulatory factor and the third DNA binding domain can be directly bound or bound via a linker. In the case where a linker is present, it will be referred to as the "third linker" according to the circumstances in this specification. As the third linker, as long as it does not inhibit the effects of the present invention (binding ability to the third DNA binding domain recognition sequence, nucleic acid base conversion activity, nickase activity of the first fusion protein, etc.), there are no particular restrictions on its chain length and type. For example, in the case where the third DNA binding domain is TALE, in the third fusion protein, the chain length of the third linker can be adjusted so that the distance between the TALE repeat domain and the aforementioned transcription regulatory factor is 1 to 630 amino acid residues, with the distance also including the chain length of the N-terminal domain or C-terminal domain (more preferably the C-terminal domain) of TALE. In this case, the distance between the TALE repeat domain and the aforementioned transcription regulatory factor (the number of amino acid residues from the amino acid residue adjacent to the C-terminal side of the C-terminus of the TALE repeat domain (or the amino acid residue adjacent to the N-terminal side of the N-terminus) as the first residue to the residue adjacent to the N-terminal side of the N-terminus of the transcription regulatory factor (or the residue adjacent to the C-terminal side of the C-terminus)) is generally 10 to 500 amino acid residues, preferably 25 to 300 amino acid residues, and further preferably 52 to 174 amino acid residues. There is no particular limitation on the type of the third linker involved in the present invention. For example, the one listed as the second linker can be appropriately applied.
[0194] In the third fusion protein involved in the present invention, the transcriptional regulatory factor and the third DNA binding domain can be connected via a third linker as needed, either on the N-terminal side or on the C-terminal side. In addition, in the third fusion protein, the binding of the transcriptional regulatory factor to the third DNA binding domain can be carried out at the nucleic acid level and the amino acid level, as these methods are the same as the binding of the nickase and the first DNA binding domain in the first fusion protein mentioned above. Further, as the third fusion protein involved in the present invention, epitope tags (e.g., Flag tags, HA tags, etc.) for purification and detection, various transfer signals (e.g., nuclear transfer signals, chromatin transfer signals, etc.) can be added.
[0195] In the DNA editing system of the present invention, the first fusion protein, the second fusion protein, and the third fusion protein can each independently be in the form of a protein, a polynucleotide (DNA, RNA) encoding the protein, or a vector (expression vector) for expressing the protein. In addition, the DNA editing system of the present invention can be a composition comprising two or more of the first fusion protein, the second fusion protein, and, if necessary, the third fusion protein, and can be a composition or a kit comprising the aforementioned combination.
[0196] <Targeted DNA Editing Method>
[0197] The target DNA editing method of the present invention comprises the following steps:
[0198] The DNA editing system of the present invention is brought into contact with a target DNA, and the base at the target site of the target DNA is edited by the activity of the nucleobase transferase, wherein:
[0199] The distance between the first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain and the second DNA binding domain recognition sequence or its complementary sequence recognized by the second DNA binding domain is 8 to 48 bases.
[0200] In addition, as the target DNA editing method of the present invention, the following method is more preferred, wherein:
[0201] The first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain is present on the 5' side or 3' side of the aforementioned target site via a first spacer of 4 to 16 bases, and
[0202] The second DNA-binding domain recognition sequence or its complementary sequence recognized by the second DNA-binding domain is present on the opposite side of the target site from the first DNA-binding domain recognition sequence or its complementary sequence via a second spacer of 3 to 31 bases.
[0203] Furthermore, in the case where the DNA editing system of the present invention further comprises a third fusion protein, as the target DNA editing method of the present invention, it is preferred that the third DNA binding domain recognition sequence recognized by the third DNA binding domain is present on the 5' side or 3' side opposite to the target site of the first DNA binding domain recognition sequence via a third spacer of 1 to 500 bases.
[0204] (Target DNA)
[0205] In the present invention, the DNA containing the target site that becomes the object of target DNA editing is referred to as "target DNA". The target DNA involved in the present invention is double-stranded DNA. In order to show the correspondence with the above-mentioned DNA editing system, conveniently, at least one chain is a structure (structure 1) that contains the first DNA binding domain recognition sequence or its complementary sequence, the first spacer (referred to as "spacer 1" as the case may be), the aforementioned target site, the second spacer (referred to as "spacer 2" as the case may be), and the second DNA binding domain recognition sequence or its complementary sequence in order from the 5' side, or at least one chain is a structure (structure 2) that contains the second DNA binding domain recognition sequence or its complementary sequence, spacer 2, the aforementioned target site, spacer 1, and the first DNA binding domain recognition sequence or its complementary sequence in order from the 5' side. In addition, in the case where the DNA editing system of the present invention further includes a third fusion protein, it is a structure (structure 3) that contains the third DNA binding domain recognition sequence on the 5' side or 3' side opposite to the target site of the first DNA binding domain recognition sequence via a third spacer (referred to as "spacer 3" as the case may be).
[0206] "Complementary sequence comprising a DNA-binding domain recognition sequence" on the 5' side or 3' side, or on the side opposite to the target site, means that the complementary chain of the chain comprises a DNA-binding domain recognition sequence. That is, the first DNA-binding domain recognition sequence and the second DNA-binding domain recognition sequence, and the third DNA-binding domain recognition sequence as needed, can be independently set on the same chain as the aforementioned target site, or on its complementary chain. In addition, the first DNA-binding domain recognition sequence and the second DNA-binding domain recognition sequence can be set on the same chain, or on chains opposite to each other. However, the first DNA-binding domain recognition sequence and the third DNA-binding domain recognition sequence are preferably on the same chain.
[0207] The target site of the present invention refers to a single base that is a target for editing by the aforementioned nucleic acid base transfer enzyme (preferably deamination by a deaminase). However, this does not exclude the possibility that bases on both sides of the single base (particularly bases that are targets for deamination by a deaminase, preferably bases within the total spacer length described below, more preferably 1 to 10 bases each on the 5' and 3' sides of the target site, further preferably 1 to 5 bases, and even more preferably 1 to 2 bases) are further edited (deaminized in the case of a deaminase).
[0208] The first DNA binding domain recognition sequence, the second DNA binding domain recognition sequence, and the third DNA binding domain recognition sequence are sequences recognized by the first DNA binding domain, the second DNA binding domain, and the third DNA binding domain, respectively. As the number of bases of each DNA binding domain recognition sequence, in the case where each of the aforementioned DNA binding domains is a TALE, each independently, preferably 10 to 30 bases, preferably 13 to 25 bases, and more preferably 15 to 22 bases. In addition, in the case where each of the aforementioned DNA binding domains is a zinc finger array, each independently, preferably 6 to 45 bases, preferably 9 to 24 bases, and more preferably 12 to 18 bases.
[0209] The first DNA binding domain recognition sequence is preferably selected in a manner such that a spacer 1 with a chain length of 4 to 16 bases is present on the 5' side or 3' side of the aforementioned target site. By having the chain length of the spacer 1 within this range, there is a tendency to be able to replace one base of the aforementioned target site more specifically and efficiently. The chain length of the aforementioned spacer 1 is the length from the base adjacent to the 5' side or 3' side of the base of the target site as the first base to the base adjacent to the 3' side or 5' side of the first DNA binding domain recognition sequence or its complementary sequence. As such a chain length of the spacer 1, 5 to 14 bases are more preferred, 6 to 13 bases are further preferred, and 8 to 11 bases are further preferred.
[0210] The second DNA binding domain recognition sequence is preferably selected in such a way that a spacer 2 with a chain length of 3 to 31 bases is present on the opposite side of the aforementioned target site to the first DNA binding domain recognition sequence. By having the chain length of the spacer 2 within this range, there is a tendency to be able to replace one base of the aforementioned target site more specifically and efficiently. The chain length of the aforementioned spacer 2 is the length from the base adjacent to the 5' side or 3' side of the base of the target site as the first base to the base adjacent to the 3' side or 5' side of the second DNA binding domain recognition sequence or its complementary sequence. As such a chain length of the spacer 2, 5 to 25 bases or 7 to 31 bases are more preferred, 7 to 19 bases are further preferred, and 8 to 15 bases are further preferred.
[0211] The chain length of spacer 1 and the chain length of spacer 2 can be arbitrarily combined as long as they meet the following conditions for the total length of the spacers. For example, the chain length of spacer 1 / the chain length of spacer 2 are preferably a combination of 4 to 16 bases / 3 to 31 bases, 5 to 14 bases / 5 to 25 bases, 6 to 13 bases / 7 to 19 bases, and 8 to 11 bases / 8 to 15 bases.
[0212] In addition, the distance between the first DNA-binding domain recognition sequence and the second DNA-binding domain recognition sequence (the length from the base adjacent to the 5' side or 3' side of the first DNA-binding domain recognition sequence or its complementary sequence as the first base to the base adjacent to the 3' side or 5' side of the second DNA-binding domain recognition sequence or its complementary sequence, including the target site. In this specification, it is referred to as the "total spacer length" depending on the situation) is 8 to 48 bases from the perspective of being able to specifically and efficiently replace one base in the aforementioned target site. As the total spacer length, 11 to 40 bases or 12 to 48 bases are preferred, 14 to 27 bases are more preferred, and 17 to 24 bases are even more preferred. In addition, depending on the type of nucleic acid base transferase contained in the second fusion protein, for example, in the case of TadA (including variants), the total length of the aforementioned spacer is more preferably 17 to 19 bases and 21 to 22 bases, and further preferably 17 to 18 bases and 21 bases.
[0213] The third DNA binding domain recognition sequence is preferably selected in such a way that it exists via a spacer 3 of 1 to 500 bases on the 5' side or 3' side opposite to the target site of the first DNA binding domain recognition sequence. By having the chain length of the spacer 3 within this range, there is a tendency to further promote the substitution of the bases at the aforementioned target site. The chain length of the aforementioned spacer 3 is the length from the base adjacent to the 5' side or 3' side of the first DNA binding domain recognition sequence as the first base to the base adjacent to the 3' side or 5' side of the third DNA binding domain recognition sequence or its complementary sequence. As such a chain length of the spacer 3, 5 to 300 bases are more preferred, and 15 to 56 bases are further preferred.
[0214] There is no particular limitation on the nucleotide sequence of such a target DNA. By designing the first fusion protein and the second fusion protein, and each DNA binding domain in the third fusion protein as needed, such that the total length of the first DNA binding domain recognition sequence or its complementary sequence, the second DNA binding domain recognition sequence or its complementary sequence, and the spacer length, and the spacer 1, spacer 2, the third DNA binding domain recognition sequence or its complementary sequence, and the spacer 3 satisfy the above conditions as needed, and selecting the nucleic acid base transferase corresponding to the target base editing, it can be used as a target of the DNA editing method of the present invention.
[0215] The target DNA involved in the present invention can be a DNA present in the cell (endogenous DNA) or a DNA present outside the cell, depending on the purpose. The DNA present in the cell can be endogenous DNA or exogenous DNA. As endogenous DNA, genomic DNA and chloroplast DNA in the nucleus are listed, and as exogenous DNA, for example, DNA introduced into the cell is listed. In the case where the target DNA involved in the present invention is a DNA present in the cell (endogenous DNA), it is necessary to use nuclear DNA, chloroplast DNA, or exogenous DNA, except for mitochondrial DNA, as the DNA, preferably, nuclear DNA or exogenous DNA is more preferred. As the DNA present outside the cell, it can be DNA of cell origin or DNA amplified and synthesized outside the cell.
[0216] (Targeted DNA Editing Method)
[0217] In the target DNA editing method of the present invention, the DNA editing system is brought into contact with the target DNA, and the bases at the target site of the target DNA are edited by the activity of the nucleobase transferase.
[0218] When the aforementioned DNA editing system, i.e., the first fusion protein and the second fusion protein, is brought into contact with the target DNA, each DNA binding domain of the first fusion protein and the second fusion protein recognizes and binds to the corresponding DNA binding domain recognition sequence on the target DNA, thereby inducing the nickase linked to the first fusion protein and the base transferase linked to the second DNA binding domain to the target DNA. Thus, at the target site, the nickase cuts one strand of the double-stranded DNA and introduces a nick. At this time, for example, in the case where the structure of the first fusion protein is a structure that binds in the order of "TALE→ND1n→ND1 linker→ND1" from the N-terminal side, the nick can be selectively introduced into the strand with the first DNA binding domain recognition sequence (the first DNA binding domain recognition strand). As a result of the introduction of the nick, through a new mechanism that is different from the previous mechanism of using the CRISPR-Cas system that relies on the R-loop, the target base substitution (e.g., deamination of the base) can be effectively caused by the base transferase in the vicinity. It is speculated that the novel mechanism described above is due to the relaxation of the higher-order structure of the double-stranded DNA caused by the introduction of the nick, resulting in the generation of partially single-stranded DNA regions. Furthermore, when the third fusion protein is brought into further contact with the target DNA, its DNA-binding domain recognizes and binds to the corresponding DNA-binding domain recognition sequence on the target DNA, inducing the transcriptional regulatory factor associated with the third fusion protein to the target DNA and promoting the aforementioned target base substitution. This is speculated to be due to the further relaxation of the higher-order structure of the double-stranded DNA caused by the transcriptional regulatory factor.
[0219] In the target DNA editing method of the present invention, one embodiment of the positional relationship between the first fusion protein, the second fusion protein, and the target DNA is shown in FIG. , for example, using TALE as each DNA binding domain and deaminase as a nucleobase transferase. Figure 16 、 Figure 27 Concept map.
[0220] As one embodiment of the target DNA editing method of the present invention, Figure 16 When the embodiment is given as an example, for example, the following embodiment is listed: the target DNA is in the order of 5' side, TALE recognition sequence (second DNA binding domain recognition sequence), spacer 2, target site ( Figure 16 (A) C), spacer 1, complementary sequence of TALE recognition sequence (complementary sequence of the first DNA binding domain recognition sequence) (first embodiment: Figure 16 (A)), or from the 5' side in order, the complementary sequence of the TALE recognition sequence (the complementary sequence of the second DNA binding domain recognition sequence), spacer 2, target site ( Figure 16 (B) A), spacer 1, complementary sequence of TALE recognition sequence (complementary sequence of the first DNA binding domain recognition sequence) (second embodiment: Figure 16 (B)). Thus, when a nickase is induced near the target site via the first DNA-binding domain, and the structure of the first fusion protein is a structure in which binding occurs from the N-terminus in the order of "TALE → ND1n → ND1 linker → ND1," a nick is introduced into the strand having the first DNA-binding domain recognition sequence (the first DNA-binding domain recognition strand). Simultaneously, since a deaminase is also induced near the target site via the second DNA-binding domain, base substitution at the target site (C or A) can be effectively induced by the deaminase activity on the strand opposite to the first DNA-binding domain recognition strand where the nick was introduced.
[0221] The method for editing the target DNA of the present invention is not limited to this. For example, the target DNA may be a complementary sequence of the first DNA binding domain recognition sequence, spacer 1, target site, spacer 2, and the second DNA binding domain recognition sequence in order from the 5' side (third embodiment), or a complementary sequence of the first DNA binding domain recognition sequence, spacer 1, target site, spacer 2, and the second DNA binding domain recognition sequence in order from the 5' side (fourth embodiment), etc.
[0222] As a result, a single base substitution (e.g., C→U) is caused at the target site. In addition, various mutations can be introduced, for example, by mismatching of double-stranded DNA within the cell, such that the base on the opposite strand of the substituted strand forms a pair with the substituted base (e.g., G→A), or is replaced by another base during repair (e.g., U→A, G), or by generating a deletion or insertion of one base or more than a dozen bases.
[0223] Therefore, the DNA editing involved in the present invention includes deletion of one or more bases in the target site converted by the aforementioned nucleobase transferase and its vicinity, substitution with one or more other bases, or insertion of one or more bases, or a combination of these mutations.
[0224] The target DNA editing method of the present invention can be carried out in a cell or in a cell-free system. As the "intracellular" where the target DNA editing method of the present invention is used, it can be in a eukaryotic cell or in a prokaryotic cell, preferably in a eukaryotic cell. As the aforementioned eukaryotic cells, for example, animal cells (cells of mammals, fish, birds, reptiles, amphibians, insects, etc.), plant cells, algae cells, yeast are listed. As the aforementioned prokaryotic cells, for example, Escherichia coli, Salmonella, Bacillus subtilis, lactic acid bacteria, and hyperthermophilic bacteria are listed.
[0225] “Animal cells” include, for example, cells constituting individual animals, cells constituting organs / tissues removed from animals, and cultured cells derived from animal tissues. Specifically, for example, reproductive cells such as oocytes and sperm are listed; embryonic cells of embryos at various stages (e.g., 1-cell embryos, 2-cell embryos, 4-cell embryos, 8-cell embryos, 16-cell embryos, morula embryos, etc.); stem cells such as induced pluripotent stem (iPS) cells and embryonic stem (ES) cells; somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells, brain cells, and kidney cells. As oocytes used in the preparation of genome-edited animals, oocytes before and after fertilization can be used, preferably oocytes after fertilization, i.e., fertilized eggs. Particularly preferably, the fertilized egg is that of a pronuclear embryo. Oocytes can be thawed and applied as those stored frozen.
[0226] "Plant cells" include, for example, individual cells constituting a plant, cells constituting organs and tissues isolated from a plant, and cultured cells derived from plant tissues. Examples of plant organs and tissues include leaves, stems, stem tips (growing points), roots, tubers, and callus tissue.
[0227] Furthermore, the "cell-free system" used in the target DNA editing method of the present invention refers to a system without living cells (the aforementioned eukaryotic cells or prokaryotic cells). The cell-free system involved in the present invention is not particularly limited as long as it is a system in which the aforementioned DNA editing system and the aforementioned target DNA can come into contact. Examples include within a buffer solution; within a cell lysate of the aforementioned eukaryotic or prokaryotic cells; or within a cell extract.
[0228] There is no particular limitation on the method for contacting the aforementioned DNA editing system with the aforementioned target DNA. In cells, for example, the following methods for preparing cells in which the target DNA has been edited are listed, and methods for introducing or expressing the aforementioned DNA editing system into cells containing the aforementioned target DNA are included. In a cell-free system, for example, a solution of the target DNA is mixed with a solution of the aforementioned DNA editing system. There is no particular limitation on the solvents for these solutions, and for example, buffers such as phosphate buffer, Tris buffer, Good buffer, and borate buffer are preferred.
[0229] <Method for preparing cells with edited target DNA>
[0230] The method for preparing cells in which target DNA has been edited of the present invention comprises the following steps:
[0231] The DNA editing system of the present invention is introduced into a cell or expressed in a cell and brought into contact with a target DNA in the cell other than mitochondrial DNA, and the bases at the target site of the target DNA are edited by the activity of the nucleobase transferase, wherein:
[0232] The distance between the first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain and the second DNA binding domain recognition sequence or its complementary sequence recognized by the second DNA binding domain is 8 to 48 bases.
[0233] Furthermore, as a method for preparing a target DNA-edited cell of the present invention, the following method is more preferred, wherein:
[0234] The first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain is present on the 5' side or 3' side of the aforementioned target site via a first spacer of 4 to 16 bases, and
[0235] The second DNA-binding domain recognition sequence or its complementary sequence recognized by the second DNA-binding domain is present on the opposite side of the target site from the first DNA-binding domain recognition sequence or its complementary sequence via a second spacer of 3 to 31 bases.
[0236] Furthermore, in the case where the DNA editing system of the present invention further includes a third fusion protein, as a method for preparing cells in which the target DNA is edited according to the present invention, the third DNA binding domain recognition sequence recognized by the third DNA binding domain is preferably present on the side opposite to the target site of the first DNA binding domain recognition sequence via a third spacer of 1 to 500 bases.
[0237] In the methods of preparing target DNA-edited cells of the present invention (hereinafter, simply referred to as "preparation methods"), the aforementioned DNA editing system and target DNA are as described above for the DNA editing system and target DNA editing methods of the present invention. The target DNA in the preparation methods of the present invention is intracellular DNA other than mitochondrial DNA, more preferably nuclear genomic DNA. Depending on the purpose of editing the genomic DNA, a first fusion protein and a second fusion protein can be designed, and further, a third fusion protein can be designed as needed.
[0238] In addition, in the preparation method of the present invention, the aforementioned DNA editing system, i.e., the first fusion protein, the second fusion protein, and the third fusion protein as needed, are brought into contact with the target DNA in the cell by the following: by introducing the aforementioned fusion proteins into the cell in the form of proteins, introducing the aforementioned fusion proteins into the cell in the form of polynucleotides, and / or introducing the aforementioned fusion proteins into the cell in the form of expression vectors, so that they are expressed in the cell. Therefore, as the aforementioned DNA editing system, the aforementioned fusion proteins can be introduced into the cell in the form of proteins, introduced into the cell in the form of RNA or DNA (polynucleotide) encoding the protein so that it is expressed in the cell, or introduced into the cell in the form of a vector (expression vector) expressing the protein so that it is expressed in the cell.
[0239] When each of the aforementioned fusion proteins is introduced into cells in the form of an expression vector and expressed in the cells, for example, vectors expressing each of the aforementioned fusion proteins may be introduced into cells separately, or vectors expressing these proteins in combination may be introduced into cells.
[0240] Furthermore, when each of the aforementioned fusion proteins is introduced into a cell in the form of an expression vector for intracellular expression, the polynucleotides encoding each of the aforementioned fusion proteins can be independently and appropriately codon-optimized according to the cell into which they are introduced. Furthermore, the expression vector preferably comprises a promoter and / or other regulatory sequences operably linked to the polynucleotide to be expressed. Furthermore, the expression vector preferably does not integrate into the host genome and is capable of stably expressing the encoded protein. Such expression vectors can be prepared according to appropriate conventional methods.
[0241] As a method for introducing the aforementioned fusion proteins, polynucleotides encoding the proteins, and vectors expressing the proteins into cells, known methods for introducing proteins, DNA, or RNA fragments into cells can be appropriately adopted depending on the type of cells. Examples of such methods include electroporation, microinjection, particle gun, calcium phosphate method, polyethyleneimine (PEI) method, liposome method (lipofection method), DEAE-dextran method, cationic lipid-mediated transfection, virus (adenovirus, lentivirus, adeno-associated virus, baculovirus, etc.), Agrobacterium method, lithium acetate method, spheroplast method, heat shock method (calcium chloride method, rubidium chloride method), etc. Such methods are described in many standard laboratory manuals such as "Davis et al., Basic Methods in Molecular Biology, New York: Elsevier, 1986".
[0242] When each of the aforementioned fusion proteins is introduced into a cell or expressed in a cell, each fusion protein comes into contact with the target DNA in the cell, and through the target DNA editing described in the target DNA editing method of the present invention, the target base is replaced at the target site. As a result, a cell in which the target DNA has been edited can be obtained.
[0243] In addition, the present invention provides a method for preparing a non-human individual comprising cells in which the target DNA has been edited. The method includes the step of preparing a non-human individual from cells obtained by the above-mentioned preparation method. As the aforementioned non-human individual, for example, non-human animals and plants are listed. As the aforementioned non-human animals, mammals (mice, rats, guinea pigs, hamsters, rabbits, monkeys, pigs, cattle, goats, sheep, etc.), fish, birds, reptiles, amphibians, and insects are listed. In the case of preparing model animals, mammals are preferably rodents such as mice, rats, guinea pigs, and hamsters, and mice are particularly preferred. As the aforementioned plants, for example, cereals, oil crops, forage crops, fruits, and vegetables are listed. As specific crops, for example, rice, corn, bananas, peanuts, sunflowers, tomatoes, rapeseed, tobacco, wheat, barley, potatoes, soybeans, cotton, and carnations can be exemplified.
[0244] As a method for preparing a non-human individual from cells that have undergone the aforementioned target DNA editing, a well-known method can be used. In the case of preparing a non-human individual from cells in animals, germ cells or pluripotent stem cells are usually used. For example, the aforementioned DNA editing system is microinjected into oocytes, and the obtained oocytes are transplanted into the uterus of a female non-human mammal in a pseudo-pregnancy state, and offspring can be obtained thereafter. In addition, in plants, it has been known since ancient times that their somatic cells have differentiation totipotency. For example, by microinjecting the aforementioned DNA editing system into plant cells and regenerating a plant body from the obtained plant cells, a plant body with the desired DNA edited can be obtained. In addition, from the obtained non-human individual, the desired DNA-edited offspring and clones can be obtained.
[0245] Confirmation of the presence or absence of target DNA editing and determination of genotype can be performed based on conventionally known methods, for example, PCR, sequencing, Southern blotting, etc. can be utilized.
[0246] <Reagent Kit>
[0247] The kit of the present invention is a kit for use in the target DNA editing method of the present invention, the method for preparing cells in which the target DNA is edited, or the method for preparing a non-human individual of the present invention, comprising:
[0248] At least one selected from the group consisting of a first fusion protein, a first fusion protein expression vector, and a polynucleotide encoding the first fusion protein, and
[0249] At least one selected from the group consisting of a second fusion protein, a second fusion protein expression vector, and a polynucleotide encoding the second fusion protein.
[0250] The kit of the present invention preferably further comprises at least one selected from the group consisting of a third fusion protein, a third fusion protein expression vector, and a polynucleotide encoding the third fusion protein.
[0251] The first fusion protein, the second fusion protein, and the third fusion protein are as described above in the DNA editing system of the present invention. They can be in the form of a protein, a polynucleotide encoding the protein, or a vector (expression vector) that expresses the protein.
[0252] In the case of an expression vector, the fusion protein can be designed by the user in a manner corresponding to the target site of the target DNA.
[0253] The first fusion protein expression vector may be at least one selected from the following:
[0254] (i) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, and a first DNA binding domain, and
[0255] (ii) a vector comprising polynucleotides encoding ND1, ND1n, and an ND1 linker, and an insertion site for a polynucleotide encoding a first DNA binding domain,
[0256] The second fusion protein expression vector may be at least one selected from the following:
[0257] (iii) a vector comprising a polynucleotide encoding a nucleobase transferase and a second DNA binding domain, and
[0258] (iv) a vector comprising a polynucleotide encoding a nucleic acid base transferase and an insertion site for a polynucleotide encoding a second DNA binding domain,
[0259] The third fusion protein expression vector may be at least one selected from the following:
[0260] (vii) a vector comprising a polynucleotide encoding a transcriptional regulator and a third DNA binding domain, and
[0261] (viii) A vector comprising a polynucleotide encoding a transcriptional regulatory factor and an insertion site for a polynucleotide encoding a third DNA-binding domain.
[0262] In the vectors of (i) to (iv), (vii), and (viii) above, the order of each structure (domain, linker, insertion site, etc.) can be appropriately adjusted according to the target first fusion protein, second fusion protein, and third fusion protein to be expressed. In addition, for example, in the case where the first DNA binding domain, the second DNA binding domain, and the third DNA binding domain are TALEs, the polynucleotide inserted into the insertion site in (ii), (iv), and (viii) may be only the TALE repeat sequence. In this case, the vectors of (ii), (iv), and (viii) may contain a polynucleotide encoding the N-terminal domain of TALE and a polynucleotide encoding the C-terminal domain as needed. Furthermore, the first fusion protein expression vector, the second fusion protein expression vector, and the third fusion protein expression vector as needed may be one vector. In addition, each vector of (i) to (iv), (vii), and (viii) preferably contains an expression unit that can express each polynucleotide.
[0263] The kit of the present invention may further include one or more additional reagents. Examples of such additional reagents include, but are not limited to, dilution buffer, reconstitution solution, wash buffer, nucleic acid introduction reagent, protein introduction reagent, and control reagent (e.g., a control deaminase). Furthermore, the kit may further include instructions for practicing the method of the present invention.
[0264] The various components included in the kit of the present invention may be housed in separate containers or in the same container. Each component may be housed in a container in an amount suitable for each use, or in a single container for multiple uses. Each component may be housed in a container in a dry form or dissolved in an appropriate solvent (including a buffer, stabilizer, preservative, preservative, etc.).
[0265] Other DNA editing methods
[0266] In the present invention, the first fusion protein of the present invention may be a method that further comprises the aforementioned nucleobase transferase, that is, a fusion protein comprising a first DNA binding domain, a nickase of the present invention, and a nucleobase transferase. Thus, as a method for editing target DNA, a method comprising the following steps can be provided: contacting the fusion protein with the target DNA, and editing the bases at the target site of the target DNA through the aforementioned nucleobase transferase activity; as a method for preparing a cell in which the target DNA has been edited, a method comprising the following steps can be provided: introducing the fusion protein into a cell or expressing it in a cell and contacting it with a target DNA in the cell other than mitochondrial DNA, and editing the bases at the target site of the target DNA through the aforementioned nucleobase transferase activity.
[0267] In the fusion protein of this situation, as the order of the first DNA binding domain, nickase, and nucleic acid base transferase, nickase and nucleic acid base transferase can be combined at any one of the N-terminal side and the C-terminal side of the first DNA binding domain. In this case, either nickase or nucleic acid base transferase can be on the first DNA binding domain side. In addition, as the aforementioned order, the first DNA binding domain can be sandwiched so that nickase and nucleic acid base transferase are combined. In this case, either nickase or nucleic acid base transferase can be on the N-terminal side or the C-terminal side. In addition, nucleic acid base transferase and the first DNA binding domain can be directly combined or can be combined via a joint. As the aforementioned joint of this situation, its preferred embodiment is also included, and the one recorded as the above-mentioned second joint is enumerated. In the case of comprising the first DNA binding domain, nickase, nucleic acid base transferase, and each joint, as these joints (first joint, second joint) and their binding methods, its preferred embodiment is also included, as described above.
[0268] In addition, in this case, the method of bringing the fusion protein into contact with the target DNA, the method of introducing it into the cell, or the method of expressing it in the cell also includes preferred embodiments thereof, which are the same as the above-mentioned method of bringing the DNA editing system into contact with the target DNA, the method of introducing it into the cell, or the method of expressing it in the cell.
[0269] Therefore, the present invention provides a kit for use in a target DNA editing method and a method for preparing a target DNA-edited cell, comprising at least one selected from the group consisting of the aforementioned fusion protein, fusion protein expression vector, and a polynucleotide encoding the fusion protein, wherein:
[0270] The aforementioned fusion protein expression vector is at least one selected from the following: (v) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, a nucleic acid base transferase, and a first DNA binding domain, and (vi) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, a nucleic acid base transferase, and an insertion site for a polynucleotide encoding the first DNA binding domain.
[0271] The kit preferably further comprises at least one selected from the group consisting of a third fusion protein, a third fusion protein expression vector, and a polynucleotide encoding the third fusion protein.
[0272] Example
[0273] Hereinafter, the present invention will be described in more detail based on test examples, but the present invention is not limited to the following test examples.
[0274] [Test Example 1]
[0275] 1. Method
[0276] (1) Preparation of TALE vector set and TALE expression plasmid
[0277] The sequence and structure of the TALE vector set were constructed according to this experimental example. The basic construction method followed the literature (Sakuma et al., (2013) Scientific Reports, 3, pp. 1-8). Specifically, the following sequences were artificially synthesized: a modular sequence (non-repeat-variable di-residue: non-RVD) encoding four variable residues consisting of two amino acids (RVD: HD, NG, NI, NN) was modified, and a restriction enzyme BsAI recognition sequence was added to both ends (a total of 16 sequences: 1HD to 4HD, 1NG to 4NG, 1NI to 4NI, and 1NN to 4NN). These sequences were inserted into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan) to create a modular plasmid set (a total of 16 sequences: pEX1HD to pEX4HD, pEX1NG to pEX4NG, pEX1NI to pEX4NI, and pEX1NN to pEX4NN). Subsequently, the nucleotide sequences FUS2_axx (7 in total: xx = 1a, 2a, 2b, 3a, 3b, 4a, 4b) and FUS2_b(1-4) (4 in total) constituting the array plasmids were prepared by artificial DNA synthesis. These sequences were inserted into pCR8 / GW / TOPO (Thermo Fisher Scientific, Waltham, MA, USA) to prepare pCR8_FUS2_axx and pCR8_FUS_b(1-4). These sequences were then used as capture vectors in the initial assembly step (step 1) of the Platinum Gate system described in the aforementioned literature by Sakuma et al., and array plasmids linked to TALE repeats were prepared according to the method described in the same literature.
[0278] Then, a target vector was prepared by inserting a sequence encoding the N-terminal domain of TALE prepared by artificial DNA synthesis, a shortened module sequence corresponding to the final module of DNA binding repeats, and a sequence encoding the C-terminal domain of TALE following the 3' end into pcDNA3.1s in which the drug resistance gene expression unit was removed from pcDNA3.1(+) (Thermo Fisher Scientific). As the C-terminal domain of TALE, three types were used: one having 63 amino acid residues in the C-terminal domain (63), one having 47 amino acid residues in the C-terminal domain (47), and one having WT C-terminal domain (amino acid residue number: 180). For the sequence encoding "63", artificial DNA synthesis was performed with reference to the sequence of the C-terminal domain contained in Addgene's pTALEN_v2 (ID: 32189~32192) (nucleotide sequence number: 66 and amino acid sequence number: 67), for the sequence encoding "47", artificial DNA synthesis was performed with reference to the sequence of the C-terminal domain contained in Addgene's ptCMV-153 / 47-VR-NG (Addgene ID: 50704) (nucleotide sequence number: 70, amino acid sequence number: 71), and for the sequence encoding "WT", artificial DNA synthesis was performed with reference to the amino acid sequence of the C-terminal domain contained in Addgene's pTALETF_v2 (ID: 32185~32188). In addition, as the N-terminal domain of TALE, the sequence of the N-terminal domain corresponding to the C-terminal domains "63" and "WT" contained in artificial DNA synthesis reference 136: Addgene's ptCMV-136 / 63-VR-HD (ID: 50699) (nucleotide sequence number: 64, amino acid sequence number: 65) and the sequence of the N-terminal domain corresponding to the C-terminal domain "47" contained in artificial DNA synthesis reference 153: Addgene's ptCMV-153 / 47-VR-HD (ID: 50703) (nucleotide sequence number: 68, amino acid sequence number: 69) were used. In addition, using the array plasmid and destination vector prepared above, three TALE (TALE63, TALE47, TALEWT) expression plasmids (TALE-136 / 63, TALE-153 / 47, TALE-136 / WT) were prepared by the Golden Gate method according to the method described in the literature of Sakuma et al.
[0279] (2) Preparation of single-chain FokI (scFokI), single-chain ND1 (scND1), and single-chain ND2 (scND2) expression plasmids
[0280] Artificial DNA synthesis was carried out by linking a nucleotide sequence encoding the HTS95 amino acid sequence (Sun, N., & Zhao, H. (2014) Molecular BioSystems, 10(3), p. 446-453, HTS95 linker: amino acid residues 197 to 291 of the amino acid sequence of sequence number 61) and a sequence (FokI-95-FokI, nucleotide sequence number: 60, amino acid sequence number: 61) that linked two partial genes encoding the nuclease domain of the FokI gene (bases 1 to 589 of the nucleotide sequence of sequence number 60, amino acid residues 1 to 196 of the amino acid sequence of sequence number 61: simply referred to as "FokI" in the experimental examples). In addition, a nucleotide sequence (FokI-60-FokI, nucleotide sequence number: 62, amino acid sequence number: 63) encoding an amino acid sequence substituted from the amino acid sequence of the HTS95 linker to the amino acid sequence of the GGGGS×12 linker (60 amino acid residues in total) was also artificially synthesized. Regarding the ND1 and ND2 genes, artificial DNA was synthesized: ND1-60-ND1 (nucleotide sequence number: 72, amino acid sequence number: 73), in which a linker sequence encoding a GGGGS×12 linker was inserted between a portion of the gene encoding the nuclease domain 1 (ND1) (bases 1 to 585 of the nucleotide sequence of SEQ ID NO: 72, amino acid residues 1 to 195 of the amino acid sequence of SEQ ID NO: 73 (amino acid sequence of SEQ ID NO: 98)); and ND2-95-ND2 (nucleotide sequence number: 86, amino acid sequence number: 87), in which a linker sequence encoding an HTS95 linker was inserted between a portion of the gene encoding the nuclease domain 2 (ND2) (bases 1 to 573 of the nucleotide sequence of SEQ ID NO: 86, amino acid residues 1 to 191 of the amino acid sequence of SEQ ID NO: 87). Furthermore, during the synthesis of the artificial DNA, the restriction enzyme recognition sequences AleI / XmaI were added upstream and PshAI / SacII downstream of the sequences encoding the GGGGS×12 linker or HTS95 linker, respectively, as common sequences. These synthetic sequences were inserted into pEX-A2J2 (Eurofins Genomics, Tokyo, Japan).
[0281] Subsequently, ND1-60-ND1 / pEX-A2J2 and ND2-95-ND2 / pEX-A2J2 were digested with restriction enzymes AleI and PshAI to isolate a 190 bp DNA fragment (60 aa) encoding a GGGGS×12 linker, a 295 bp DNA fragment (95 aa) encoding an HTS95 linker, and fragments of the remaining vector components, respectively. The 60-amino acid residue DNA fragments were ligated using T4 DNA ligase and then digested with restriction enzymes XmaI and SacII. The resulting DNA fragments were separated by agarose gel electrophoresis to separate and purify a 380 bp DNA fragment (120 aa: GGGGS×24 linker) and a 570 bp DNA fragment (180 aa: GGGGS×36 linker). The above-mentioned 95aa, 120aa, and 180aa DNA fragments were respectively inserted into the fragments of the vector portion obtained by cutting ND1-60-ND1 / pEX-A2J2 with the restriction enzymes AleI and PshAI. In addition, the above-mentioned 60aa, 120aa, and 180aa DNA fragments were respectively inserted into the fragments of the vector portion obtained by cutting ND2-95-ND2 / pEX-A2J2 with the restriction enzymes AleI and PshAI. Plasmids ND1-95-ND1 (nucleotide sequence number: 78, amino acid sequence number: 79) / pEX-A2J2 and ND1-1 20-ND1 (nucleotide sequence number: 74, amino acid sequence number: 75) / pEX-A2J2, ND1-180-ND1 (nucleotide sequence number: 76, amino acid sequence number: 77) / pEX-A2J2, ND2-60-ND2 (nucleotide sequence number: 80, amino acid sequence number: 81) / pEX-A2J2, ND2-120-ND2 (nucleotide sequence number: 82, amino acid sequence number: 83) / pEX-A2J2, and ND2-180-ND2 (nucleotide sequence number: 84, amino acid sequence number: 85) / pEX-A2J2.
[0282] Subsequently, the TALE63-scFokI expression plasmid, TALE63-scND1 expression plasmid, and TALE63-scND2 expression plasmid were prepared by the following method, in which the above-mentioned scFokI (FokI-y-FokI, y: 95, 60), scND1 (ND1-y-ND1, y: 95, 60, 120, 180), and scND2 (ND2-y-ND2, y: 95, 60, 120, 180) sequences were respectively linked to the downstream of the TALE sequence of the TALE expression plasmid (TALE-136 / 63) prepared in (1). That is, first, the above-mentioned plasmids FokI-95-FokI / pEX-A2J2, FokI-60-FokI / pEX-A2J2, ND1-95-ND1 / pEX-A2J2, ND1-60-ND1 / pEX-A2J2, ND1-120-ND1 / pEX-A2J2, ND1-180-ND1 / pEX-A2J2, ND2-95-ND2 / pEX-A2J2, ND2-60-ND2 / pEX-A2J2, ND2-120-ND2 / pEX-A2J2, and ND2-180-ND2 / pEX-A2J2 were used as templates to amplify the sequences of scFokI, scND1, and scND2, respectively, by PCR. They were inserted downstream of the sequence of TALE63 (N-terminal domain: 136 amino acid residues / TALE repeat domain / C-terminal domain: 63 amino acid residues) of the platinum TALE structure of the TALE expression plasmid by the In-Fusion method (TaKaRa Bio Inc, Shiga, Japan), and TALE63-scFokI expression plasmid, TALE63-scND1 expression plasmid, and TALE63-scND2 expression plasmid were prepared respectively. In addition, the sequence of the TALE in the case of inserting the scFokI sequence further added two amino acids "LK" on the C-terminal side of the amino acid sequence of the C-terminal domain. In addition, similarly, for each sequence of scND1, a TALE47-scND1 expression plasmid was prepared by linking the sequence of TALE47 downstream of the TALE expression plasmid (TALE-153 / 47) prepared in (1). Furthermore, similarly, expression plasmids for TALE63-ND1mono, TALE63-ND2mono, TALE47-ND1mono, and TALE47-ND2mono were prepared by ligating the artificially synthesized DNA sequences of ND1 and ND2 downstream of the TALE sequence of the TALE expression plasmid (TALE-136 / 63 or TALE-153 / 47) prepared in (1). The templates and primers used in these preparations are shown in Table 1 below.
[0283] In each expression plasmid provided for SSA testing, a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence on Rosa26 (Rosa26-L, R), the TALE recognition sequence on APC (APC-L, R), or the TALE recognition sequence on the HPRT1 gene (HPRT1-L, R) shown in Table 8 below was inserted.
[0284]
[0285] (3) Preparation of reporter plasmid for single-strand annealing (SSA) assay
[0286] For the SSA test reporter plasmid, first, based on the sequence information described in the reference (Mashiko et al., (2013) Scientific Reports, 3, 3355, DOI: 10.1038 / srep03355), PCR was performed using artificially synthesized EGFP cDNA as a template to obtain two amplification products, one at the N-terminus and the other at the C-terminus. Subsequently, PCR was performed again using primers that added a linker sequence between the N-terminus and the C-terminus to obtain the amplified products. These amplified products were then inserted into the BamHI / EcoRV restriction enzyme recognition sequence of pcDNA3.1s using the In-Fusion method to create a plasmid (pcEGxxFP) expressing a sequence with a linker sequence inserted between the N-terminus and the C-terminus of the EGFP sequence (referred to as "EGxxFP"). The templates and primers used in these preparations are shown in Table 2 below.
[0287]
[0288] As target genes used in the SSA test, APC (human adenomatous polyposis coli) was isolated from the human genome by PCR, and Rosa26 and HPRT1 were isolated from the hamster genome. These were then inserted into the restriction enzyme BamHI / EcoRI recognition sequences within the linker inserted between the N-terminus and C-terminus of EGFP within pcEGxxFP by in-fusion to prepare reporter plasmids for each SSA test. Sequences containing the nucleotide sequences of the target genes APC, Rosa26, and HPRT1 used in the SSA test evaluation (EGxAPCxFP, EGxRosa26xFP, and EGxHPRT1xFP) are shown in order as SEQ ID NOs: 94-96.
[0289] (4) Preparation of plasmids for TALE C-terminal length studies
[0290] The chain length between scND1 and the TALE repeat domain of TALE was studied by varying the C-terminal length using the C-terminal end of TALE47 of the TALE47-scND1 expression plasmid prepared in (2) as a starting point. In the case of longer lengths, PCR was performed using TALE-136 / WT prepared in (1) as a template, and the amplified product was linked between the C-terminal domain of TALE47 of the TALE47-scND1 expression plasmid prepared in (2) and scND1. 9 TALE47-z-scND1 expression plasmids having chain lengths ranging from 63 amino acid residues to 180 amino acid residues (z: 63, 75, 90, 105, 120, 135, 150, 165, 180) were prepared. Furthermore, in the case of a shorter sequence, PCR was performed using the TALE47-scND1 expression plasmid prepared in (2) as a template to prepare TALE47-z-scND1 expression plasmids having four different C-terminal domains (z: 0, 12, 24, 36) with chain lengths ranging from 0 to 36 amino acid residues. The templates and primers used in these preparations are shown in Table 3 below.
[0291]
[0292] (5) Preparation of flexible joints
[0293] Three types of flexible linkers were prepared: a GSS linker, a SAGG linker, and a GGGGS linker, allowing for insertion into the target sequence with adjustable repeat counts. First, sequences containing a restriction enzyme recognition sequence of XmaI at the 5' end and a BamHI or MroI recognition sequence at the 3' end were synthesized using forward and reverse oligonucleotides: a GSS × 4 adapter, a SAGG × 3 adapter, and a GGGGS × 3 adapter. The oligonucleotides used in these preparations are shown in Table 4 below.
[0294]
[0295] After phosphorylation and annealing at the 5' end of each oligonucleotide, it was inserted into a plasmid containing one each of XmaI, BamHI, and MroI (for example, since the 60-amino acid linker sequence of TALE63-ND1-60-ND1 contains one each of XmaI, BamHI, and MroI cleavage sequences, TALE63-ND1-60-ND1 was used as a dummy vector in this experimental example). These vectors were GSS linker vectors, SAGG linker vectors, and GGGGS linker vectors, respectively.
[0296] In addition, as flexible linker repeat sequences to be inserted into each linker vector, sequences containing the restriction enzyme BglII recognition sequence at the 5' end and BamHI recognition sequence at the 3' end, or the recognition sequence of XmaI recognition sequence at the 5' end and MroI recognition sequence at the 3' end were synthesized using respective forward and reverse oligonucleotides: GSS×4, GSS×7, SAGG×3, SAGG×4, GGGGS×3, and GGGGS×4. The oligonucleotides used in these preparations are also shown in Table 4 above. After phosphorylation and annealing at the 5' end of each oligonucleotide, a ligation reaction was performed, followed by double cleavage with the restriction enzymes BglII and BamHI, and XmaI and MroI, and fragments with the desired number of repeats were recovered by agarose gel electrophoresis. For example, in the case of GSS×4 (number of repeats: 4), repeat sequences such as GSS×8, GSS×12, and GSS×16 are obtained. By mixing with GSS×7 and performing a ligation reaction, DNA fragments with various numbers of repeats such as GSS×11, GSS×14, and GSS×15 can be obtained.
[0297] By inserting DNA fragments having each target repeat number (final target repeat number minus the number of repeats inserted into the linker vector) into the BamHI or MorI site of the GSS linker vector, SAGG linker vector, or GGGGS linker vector, polynucleotides encoding linkers having the final target repeat number were obtained. These polynucleotides were isolated by double cleavage with the restriction enzymes XmaI and BamHI, and XmaI and MorI.
[0298] Next, in order to insert a DNA fragment with the final target number of repeats in place of the HTS95 linker (95 amino acid residues) in the ND1-95-ND1 sequence, two plasmids, TALE24-scND1 (bam) and TALE24-scND1 (mro), were prepared, each with a recognition sequence for the restriction enzyme BamHI or MorI inserted before ND1 on the C-terminal side (on the N-terminal side) in the TALE24-scND1 expression plasmid (scND1: ND1-95-ND1) prepared in (4). The templates and primers used in the preparation of these plasmids are shown in Table 5 below.
[0299]
[0300] Then, each plasmid was double-cut with restriction enzymes XmaI and BamHI, and XmaI and MorI, and then dephosphorylated at the 5' end by alkaline phosphatase. Thereafter, a ligation reaction was carried out together with the DNA fragment of the final target repeat number prepared above, and finally, TALE24-ND1-GSS×32 linker-ND1 expression plasmid, TALE24-ND1-SAGG×24 linker-ND1 expression plasmid, and TALE24-ND1-GGGGS×19 linker-ND1 expression plasmid were obtained. In each plasmid, the sequence of the TALE repeat domain is a sequence corresponding to the nucleotide sequence of the TALE recognition sequence on APC (APC-L, R), the TALE recognition sequence on Rosa26 (Rosa26-L, R), and the TALE recognition sequence on HPRT1 (HPRT1-L, R) shown in Table 8 below. The nucleotide sequences and amino acid sequences of ND1-GSS×32 linker-ND1, ND1-SAGG×24 linker-ND1, and ND1-GGGGS×19 linker-ND1 are shown in SEQ ID NOs: 88 to 93 in order.
[0301] (6) Preparation of TALE-scND1 mutant (scND1n)
[0302] To introduce a mutation into one of the ND1s of scND1, the N-terminal ND1 and the C-terminal ND1 of the two were first placed on separate plasmids. Specifically, the TALE47-12-ND1-95-ND1 (hereinafter referred to as "TALE12-ND1-95-ND1") expression plasmid prepared in (4) above was used as a template, and PCR was performed using the primer set shown in Table 6 below. The PCR fragments were ligated by the In-Fusion method to prepare a TALE12-ND1 / pcDNA3.1s plasmid containing up to "TALE12-ND1" and a 95-ND1 / pcDNA3.1s plasmid containing the remaining "95-ND1". Then, using TALE12-ND1 / pcDNA3.1s as a template, a site-specific mutagenesis method using PCR was performed using the primer sets N-D450N, N-D450A, and N-D467A shown in Table 7 below, respectively, to introduce three mutations: D450N, D450A, or D467A mutations into ND1 on the N-terminal side. Furthermore, using 95-ND1 / pcDNA3.1s as a template, a site-specific mutagenesis method using PCR was performed using the primer sets C-D450N, C-D450A, and C-D467 shown in Table 7 below, respectively, to introduce three mutations: D450N, D450A, or D467A mutations into ND1 on the C-terminal side.
[0303] (7) Preparation of TALE-scND1n expression plasmid
[0304] Two plasmids each having ND1 at the N-terminal side and ND1 at the C-terminal side before introduction of the mutation and the six plasmids prepared in (6) above were double-cleaved with restriction enzymes XmaI and ScaI, and a total of eight plasmids were prepared by appropriately combining the N-terminal side and the C-terminal side of TALE12-ND1-95-ND1 for ligation reaction to prepare TALE12-ND1(D450N)-95-ND1 and TALE12-ND1(D450A) in which ND1 at the N-terminal side was introduced with the mutation. The expression plasmids for TALE12-scND1n (Example) containing six scND1n mutants were constructed, including three types of TALE12-ND1 (D450N), TALE12-ND1-95-ND1 (D450A), and TALE12-ND1-95-ND1 (D467A), each containing a C-terminal ND1 mutation. "D450" corresponds to the aspartic acid at position 66 of the amino acid sequence of SEQ ID NO: 98, and "D467" corresponds to the aspartic acid at position 83 of the amino acid sequence of SEQ ID NO: 98.
[0305]
[0306]
[0307] In each TALE12-scND1n expression plasmid used for SSA testing, a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence on Rosa26 (Rosa26-L, R), the TALE recognition sequence on APC (APC-L, R), or the TALE recognition sequence on the HPRT1 gene (HPRT1-L, R) was inserted. Furthermore, in the TALE12-scND1n expression plasmid used for T7E1 testing, a TALE repeat domain corresponding to the TALE recognition sequence on APC (off4-L) was inserted. Furthermore, in each TALE12-scND1n expression plasmid for the reporter test, a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence on Rosa26 (Rosa26-L) and the recognition sequence on the reporter (Rosa26-LM4.2~LM.2+4) was inserted, and in each TALE12-scND1n expression plasmid for the endogenous DNA editing test, a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence on APC (APC-L2-1~3) was inserted. In addition, in this specification, when recorded as "TALEz(ww)-", "z" represents the chain length of the C-terminal domain of TALE, and "ww" represents the corresponding TALE recognition sequence. Each TALE recognition sequence is shown in Table 8 below.
[0308]
[0309] (8) Preparation of Cas9 and its mutant expression plasmids
[0310] First, the region from the U6 promoter to the BGH poly A addition sequence of pX330 (Addgene, Cambridge, MA; Plasmid 42230) was artificially synthesized and inserted between the EcoRV and BamHI recognition sites of pBlueScriptII (SK+) (Stratagene, La Jolla, CA, USA) via in-fusion to create the Cas9 expression plasmid pX330_BS. Next, the Cas9 gene in pX330_BS was subjected to site-specific mutagenesis using PCR to generate expression plasmids for various mutants: the nickase-type nCas9 (D10A) expression plasmid, the nCas9 (H840A) expression plasmid, and the dCas9 (D10A + H840A) expression plasmid, which lacks nicking activity.
[0311] (9) Preparation of guide RNA expression plasmid
[0312] First, the pX330_BS prepared in (8) above was treated with restriction enzymes XbaI and NotI to remove the Cas9 gene. After blunt-end reaction, a plasmid (pX330_BS-ΔCas9) was prepared by self-ligation. Then, oligonucleotides designed in a manner corresponding to the target sequence of the guide RNA shown in Table 9 below (in Table 9, the complementary sequence (other than the underlined sequence) of the target sequence (on the antisense strand) of the guide RNA) were annealed and inserted into pX330_BS-ΔCas9 by BpiI treatment and ligation to prepare plasmids expressing each guide RNA (gRNA-A to D) for each target gene of Rosa26 and HPRT1, and plasmids expressing each guide RNA (gRNA-A to D and gRNA-off4-L3 to L5) for the target gene APC. Table 9 below shows the complementary sequence (other than the underlined sequence) of the target sequence of the guide RNA and the PAM sequence (underlined) of Cas9.
[0313]
[0314] (10) Preparation of other plasmids used in this experimental example
[0315] TALE-deaminases: TALE47-12-AID, ABE8e-32-TALE47 expression plasmids, and their negative control plasmids; AncBE4max and ABE8e expression plasmids were prepared in the same manner as the TALE47-12-AID expression plasmid, ABE8e-32-47 expression plasmid, and their negative control plasmids; and AncBE4max expression plasmid and ABE8e expression plasmid described in International Publication No. 2022 / 050377. "AID" denotes PmCDA1 deaminase, "ABE8e" denotes TadA-8e deaminase (a variant of TadA), "CBE" below corresponds to TALE47-12-AID with C / G converted to T / A, and "ABE" corresponds to ABE8e-32-TALE47 with A / T converted to G / C.
[0316] Into each TALE-deaminase expression plasmid supplied for reporter assays and endogenous DNA editing assays, a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence on APC (APC-R) was inserted.
[0317] (11) Quantification of DNA double-strand cleavage activity by SSA assay
[0318] (i)TALE-scFokI, TALE-scND1, TALE-scND2
[0319] The SSA test for TALE-scFokI, TALE-scND1, and TALE-scND2 was performed as follows. Specifically, HEK293T cells grown in DMEM medium containing 10% FBS were cultured for 1×10 4 One day before transfection, 1000 cells were seeded into each well of a 96-well plate. For each of the TALE-scFokI expression plasmid, TALE-scND1 expression plasmid, and TALE-scND2 expression plasmid prepared in (2) above, 33 ng of either a plasmid containing a TALE repeat domain corresponding to Rosa26-L, APC-L, or HPRT1-L (TALE-L) or a plasmid containing a TALE repeat domain corresponding to Rosa26-R, APC-R, or HPRT1-R (TALE-R) were added, along with 33 ng of the SSA assay reporter plasmid prepared in (3) above. These plasmids were then introduced into HEK293T cells using Lipofectamine 3000 (Thermo Fisher Scientific). To reduce the number of plasmids to be introduced, pBluescript II (SK+) (Stratagene, La Jolla, CA, USA) was used to supplement the plasmids, resulting in a total of 100 ng of the introduced plasmids. After 48 hours of culture after transfection, the fluorescence per unit area of each well caused by EGFP protein was measured using a microplate reader. To reduce variations caused by cell localization within the wells, fluorescence was measured at 16 points in each well, and the average value was used as the EGFP fluorescence intensity (RFI) to evaluate double-strand cleavage activity. The average and standard deviation of the values from three independent wells were plotted graphically.
[0320] (ii)TALE-scND1n
[0321] The SSA assay using TALE-scND1n was performed as follows. Specifically, double-strand cleavage activity was evaluated by fluorescence measurement in the same manner as in (i) above, except that 2 ng of any one of the TALE12-scND1n expression plasmids prepared in (7) above, 16.5 ng of the guide RNA expression plasmid prepared in (9) above, 33 ng of the nCas9 (D10A) expression plasmid prepared in (8) above, and 33 ng of the SSA assay reporter plasmid prepared in (3) above were used.
[0322] (12) Preparation of reporter plasmid for base substitution activity detection
[0323] As a reporter, pNLF1-C [CMV Hygro] (manufactured by PROMEGA, WI, USA) was used to change the methionine codons (ATG) at positions 71 and 107 to alanine codons (ATA) by site-specific mutagenesis to create pNLF-M / A. Subsequently, oligonucleotides designed to correspond to the sequences shown in Table 10 below were annealed to pNLF-M / A and inserted into pNLF-M / A treated with the restriction enzymes NheI and XhoI. Reporter plasmids were prepared with each sequence (sequence shown in Table 10) inserted for studying the length of Spacer 1 (the distance between the target base (target site, hereinafter referred to as the same) and the TALE recognition sequence of TALE-scND1n or its complementary sequence). In addition, oligonucleotides designed to correspond to the sequences shown in Table 11 below were annealed and inserted into pNLF-M / A treated with restriction enzymes NheI and XhoI to prepare reporter plasmids of various sequences (sequences shown in Table 11) for studying the length of insertion spacer 2 (the distance between the target base and the TALE recognition sequence of TALE-deaminase or its complementary sequence). The insertion sequence includes the TALE recognition sequence of TALE-deaminase (underline (*1) in Tables 10 and 11) or its complementary sequence (underline (*3) in Table 11), a codon containing the target base (underline (*2) in Tables 10 and 11: 5'-ACG-3'), a complementary sequence of the TALE recognition sequence of TALE-scND1n (underline (*3) in Table 10, underline (*4) in Table 11) and a spacer sequence therebetween. In the study of the length of spacer 1, the length of spacer 1 was 4 to 16 bases, and in the study of the length of spacer 2, the length of spacer 2 was 7 to 16 bases.
[0324]
[0325]
[0326] A conceptual diagram showing the structure of the completed reporter plasmid is shown in Figure 17 In this reporter plasmid, the target codon is located adjacent to the start codon of NanoLuc luciferase ( Figure 17 (A)) or stop codon ( Figure 17 (B)) corresponding position, based on this reporter plasmid, the target codon ACG is converted to ATG by TALE-deaminase binding to the TALE recognition sequence ( Figure 17 (A)), or the stop codon (TAG) is changed to a tryptophan codon (TGG) ( Figure 17(B)) NanoLuc luciferase was expressed. Guide RNAs for AncBE4max (gRNA-CBE-L4M2, gRNA-ABE-L4M2) as positive controls were prepared according to (9) above. The complementary sequences of the target sequences of each guide RNA (except underlined) and the PAM sequence of Cas9 (underlined) are also shown in Table 9 above.
[0327] (13) Transfection into HEK cells
[0328] (i) Reporter test
[0329] The reporter test using each reporter plasmid prepared in (12) as the target DNA was performed as follows. Specifically, 20 ng of any one of the reporter plasmids prepared in (12); 75 ng of any one of the TALE-deaminase expression plasmids and negative control plasmids prepared in (10); 75 ng of any one of the TALE12-scND1n expression plasmids prepared in (7); 20 ng of the guide RNA expression plasmid prepared in (9); and 5 ng of pGL4.54 as a reference plasmid were combined and introduced into each well using Lipofectamine LTX (manufactured by Thermo Fisher Scientific) at a concentration of 5 × 10 4 The cells were cultured in HEK293T cells for 24 hours. The reference plasmid was a plasmid expressing firefly luciferase (Fluc).
[0330] In addition, as a positive control, 20 ng of any one of the reporter plasmids prepared in (12) above; 75 ng of any one of the AncBE4max expression plasmid and ABE8e expression plasmid prepared in (10) above; 20 ng of the guide RNA expression plasmid prepared in (9) above; and 5 ng of the reference plasmid pGL4.54 were similarly introduced into HEK293T cells and cultured for 24 hours.
[0331] (ii) Endogenous (intranuclear) DNA editing testing
[0332] When endogenous (nuclear) DNA was used as the target DNA, 60 ng of the TALE-deaminase expression plasmid prepared in (10) and 30 ng of the TALE12-scND1n expression plasmid prepared in (7) were combined and introduced into each well using Lipofectamine LTX (manufactured by Thermo Fisher Scientific) at 3 × 10 4In the case of T7E1 testing, 75 ng of the TALE12-scND1n expression plasmid prepared in (7), 75 ng of either the Cas9 expression plasmid pX330_BS or its mutant expression plasmid prepared in (8), and 25 ng of the guide RNA expression plasmid prepared in (9) were added to HEK293T cells and cultured for 48 hours.
[0333] (14) Determination of NanoLuc luciferase activity
[0334] After culturing for 24 hours after transfection in (i) above (13), the medium was removed, the cells were washed with PBS(-), and the cells were lysed with Passive Lysis Buffer (manufactured by PROMEGA) to prepare a cell lysate. The cell lysate was diluted 100-fold with DMEM medium, and the NanoLuc luciferase activity and firefly luciferase activity were measured using a Nano-Glo Dual-Luciferase Reporter Assay System (manufactured by PROMEGA) and a TriStar S LB942 microplate reader (manufactured by Berthold Technologies).
[0335] As a negative control for each score, the NanoLuc luciferase activity score was measured in cells transfected with a negative control plasmid for TALE-deaminases. Furthermore, for the positive control (cells transfected with an AncBE4max expression plasmid or an ABE8e expression plasmid), the activity score was measured in cells transfected with the pX330_BS-ΔCas9 plasmid (cells not expressing the guide RNA) as a negative control (AncBE4max-NC or ABE8e-NC).
[0336] The activity was quantified as follows. First, the activity score of each NanoLuc luciferase was normalized using the firefly luciferase activity score of the reference plasmid. Then, using the normalized activity score, the activity score of the cells into which the TALE-deaminase negative control plasmid was introduced was subtracted from the activity score of the cells into which the TALE-deaminase expression plasmid was introduced to obtain the activity value of the TALE-deaminase activity. In addition, AncBE4max-NC or ABE8e-NC was subtracted from the activity score of the cells into which the AncBE4max expression plasmid or ABE8e expression plasmid was introduced to obtain the AncBE4max activity value or ABE8e activity value. Furthermore, in order to compare the TALE-deaminase activity between each experiment, the TALE-deaminase activity value was displayed as the relative activity when the AncBE4max activity value or ABE8e activity value was 1. The experiment was repeated at least 3 times, and the average value was assigned a standard deviation and graphed.
[0337] (15) Analysis of base editing activity of endogenous (nuclear) DNA
[0338] 48 hours after transfection in (ii) of the above (13), the culture medium was removed, and after washing with PBS (-), 50 μL of DNAzol (Molecular Research Center, Inc.) was added to dissolve the cells to prepare a cell lysate. Using the above cell lysate as a template, the APC-F2 and APC-R primers listed in Table 12 below were applied to amplify the region of the target site containing the above endogenous (nuclear) DNA using PrimeSTAR Max (manufactured by Takara Bio Inc.). After purification of the obtained PCR product, sequencing was commissioned to Fasmac Co., Ltd. EditR (Kluesner et al., The CRISPR J 1, 239-250 (2018)) was applied to analyze the sequence data, and the ratio (%) of T in each C within the target range on the target site was calculated.
[0339]
[0340] (16) T7E1 test
[0341] After transfection in step (ii) of (13), PCR was performed using the method described in (15) above, using the off4-F and off4-R primers described in Table 12. The resulting PCR products were subjected to the T7E1 assay according to the manufacturer's protocol.
[0342] <2. Results>
[0343] (1) Confirmation of double-strand cleavage activity by SSA test
[0344] (i) Confirmation of TALE-scFokI activity
[0345] First, the activity of the reporter TALE-scFokI was confirmed as a monomeric nuclease lacking base recognition activity. The TALE used was a TALE with the TALE repeat domains of the TALEN Right TALE (APC-R: nucleotide sequence of SEQ ID NO: 116 / complementary strand of nucleotides 119 to 136 of SEQ ID NO: 94) and Left TALE (APC-L: nucleotide sequence of SEQ ID NO: 117 / nucleotides 85 to 102 of SEQ ID NO: 94), which targets the TALE recognition sequence on the APC gene, whose intracellular activity has been confirmed in a non-patent literature (Sakuma et al., Sci. Rep., 3: 3379 (2013)). As a positive control, TALENs having TALE repeat domains corresponding to APC-R and APC-L (TALEN (APC-R) + TALEN (APC-L)) were placed and evaluated by the SSA test method in HEK293T cells ((11)(i) of [Test Example 1] <1. Method>). As a result, regarding the TALE-scFokI expression plasmid (scFokI: FokI-95-FokI) prepared in (2) of [Test Example 1] <1. Method>, there was no significant difference in double-stranded cleavage activity between the one having only the TALE repeat domain corresponding to APC-R (TALE (APC-R)-FokI-95-FokI) and the one having only the TALE repeat domain corresponding to APC-L (TALE (APC-L)-FokI-95-FokI) and the case where only the SSA test reporter plasmid was used (blank control (mock)). Figure 1 ).
[0346] Therefore, in anticipation of improved activity, the TALE-scFokI expression plasmid (scFokI: FokI-60-FokI) in which the linker sequence between FokI was changed to a GGGGS×12 linker in the same manner as in (5) of [Test Example 1]<1. Method> was tested in the same manner as in the case where the TALE repeat domain corresponding to APC-R (TALE(APC-R)-FokI-60-FokI) and the TALE repeat domain corresponding to APC-L (TALE(APC-L)-FokI-60-FokI) were only included. However, as in the above, no double-strand cleavage activity was confirmed ( Figure 1 ). This suggests that the cleavage activity of scFokI is very low, making it difficult to effectively utilize for genome editing in human cells.
[0347] (ii) Confirmation of double-strand cleavage activity of TALE-scND1 and TALE-scND2
[0348] Then, following the structure of scFokI, the same study was conducted using ND1 and ND2, which were discovered as replacement factors for the FokI nuclease domain. The TALE63-ND1 mono expression plasmid (TALE63 (APC-R / L)-ND1 mono expression plasmid) having a TALE repeat domain corresponding to APC-R or APC-L prepared in (2) of [Test Example 1] <1. Method> and their combination, the TALE63-scND1 expression plasmid (linker: 60, 95, 120, 180 aa) having a TALE repeat domain corresponding to APC-L, and the TALE63-ND1 mono expression plasmid (linker: 60, 95, 120, 180 aa) having a TALE repeat domain corresponding to APC-R or A The TALE63-ND2mono expression plasmid (TALE63 (APC-R / L)-ND2mono expression plasmid) with a TALE repeat domain corresponding to PC-L and their combination, and the TALE63-scND2 expression plasmid (linkers: 60, 95, 120, 180 aa) with a TALE repeat domain corresponding to APC-L were evaluated using the SSA test in HEK293T cells ((11)(i) of [Test Example 1] <1. Method>). The conceptual diagram of the structure of each protein (TALE-ND1mono, TALE-scND1, TALE-ND2mono, TALE-scND2) expressed by these expression plasmids is shown in FIG. Figure 2 A Flag tag (3×FLAG: not shown) and a nuclear translocation signal NLS were added to the N-terminus of each protein.
[0349] The results of the test showed that the double-stranded cleavage activity of the combination of TALE63-ND1mono expression plasmid (TALE63(APC-R / L)-ND1mono expression plasmid) with TALE repeat domains corresponding to APC-R or APC-L (TALE(APC-R)ND1mono + TALE(APC-L)ND1mono) and the combination of the same TALE63-ND2mono expression plasmid (TALE(APC-R)ND2mono + TALE(APC-L)ND2mono) was significantly reduced, but surprisingly, in the case of linking and single-stranding (in the case of scND1), both showed excellent activity ( Figure 3 、 Figure 4 On the other hand, in ND2, when linked and single-stranded (in the case of scND2), no double-stranded cleavage activity was shown ( Figure 3Furthermore, when the TALE63-scND1 expression plasmid was changed to the TALE47-scND1 expression plasmid, high activity was shown in TALE47-ND1-95-ND1 ( Figure 4 ).
[0350] about Figure 4 The result of scND1 is the case where the TALE repeat domain corresponds to APC-R, but the other five cases, specifically, the case where the TALE repeat domain corresponds to APC-L; the case where the TALE repeat domain corresponds to the Right TALE (HPRT1-R: nucleotide sequence of sequence number 132 / the complementary chain of the nucleotide sequence of sequence number: 96) and the Left TALE (HPRT1-L: nucleotide sequence of sequence number 133 / the complementary chain of the nucleotide sequence of sequence number: 96) of TALEN, which is the TALE recognition sequence on HPRT1; the case where the Left TALE (Rosa26-L: nucleotide sequence of sequence number 121 / the complementary chain of the nucleotide sequence of sequence number: 65) and the Right TALE (Rosa26-L: nucleotide sequence of sequence number 121 / the complementary chain of the nucleotide sequence of sequence number: 65) of TALEN (TALE nucleases), which is the TALE recognition sequence on hamster Rosa26. Similarly, in the case of TALE (Rosa26-R: nucleotide sequence of sequence number 122 / nucleotide sequence of sequence number 65, bases 130 to 147) (Sato et al., (2015) Stem Cell Reports, 14; 5(1), p. 75-82), TALE47-ND1-95-ND1 showed higher double-strand cleavage activity ( Figure 5A 、 Figure 5B ) These results indicate that, in scND1, the linker between ND1s is preferably the HTS95 linker found in the HTS95 amino acid sequence.
[0351] (2) Study of TALE-scND1 structure
[0352] Then, the effect of the sequence length on the activity of the C-terminal side of TALE was studied. First, the C-terminal side sequence in scND1 prepared in (4) of [Test Example 1] <1. Method> was extended (z: 63, 75, 90, 105, 120, 135, 150, 165, 180 amino acid residues), and when the SSA test ((11) (i) of [Test Example 1] <1. Method>) was performed using TALE47-z (Rosa26-L) -scND1 or TALE47-z (Rosa26-R) scND1 having a TALE repeat domain corresponding to Rosa26-L or Rosa26-R, no change in activity corresponding to the number of amino acid residues on the C-terminal side was observed ( Figure 6 On the other hand, when the C-terminal side sequence in scND1 prepared in (4) of [Test Example 1]<1. Method> was shortened (z: 36, 24, 12, 0 amino acid residues), and the SSA test ((11)(i) of [Test Example 1]<1. Method>) was performed using TALE47-z(Rosa26-L / R)-scND1, TALE47-z(APC-L / R)-scND1, or TALE47-z(HPRT1-L / R)-scND1 having a TALE repeat domain corresponding to Rosa26-L / R, APC-L / R, or HPRT1-L / R) was performed, especially when the C-terminal side sequence was more than 12 amino acid residues, it was stable and highly active in any case ( Figure 7A 、 Figure 7B ).
[0353] In the above (1)(ii), the result obtained was that the linker length between ND1 in scND1 was the most suitable at 95 amino acid residues, but this linker is a sequence unique to the HTS95 (95aa) linker. In contrast, the other linkers of 60, 120, and 180 amino acid residues are composed of a repeating sequence of the GGGGS sequence. Therefore, it is unclear whether the chain length of 95 amino acid residues is the most suitable as a linker or the sequence of the HTS95 linker is the most suitable. Therefore, the expression plasmids prepared in (5) of [Test Example 1] <1. Method> and inserted with sequences encoding three types of linkers (GSS×32 linker, SAGG×24 linker, GGGGS×19 linker) were used to compare the activities using the SSA test ((11)(i) of [Test Example 1] <1. Method>). The conceptual diagram of the structure of these expression plasmids is shown in FIG. Figure 8 Flag tags (3×FLAG: not shown) and nuclear translocation signal NLS were added to the N-terminal side of each protein. As a result, relatively good activity was observed as a whole when the HTS95 linker was used ( Figure 9 (A) to (C)), when APC was used as the target gene, the same activity was observed for all linkers ( Figure 9 (B)).
[0354] (3) Nicking of TALE-scND1
[0355] (i) SSA test
[0356] Nickase activity was reported in TALENs by introducing D450N, D450A, or D467A mutations into the FokI nuclease domain of either TALEN-L or TALEN-R (Wu et al., (2014) Biochemical and Biophysical Research Communications, 446, 261, DOI: 10.1016 / j.bbrc.2014.02.099). Since these amino acid residues D450 and D467 are also conserved in ND1, nickase activity was investigated by introducing mutations corresponding to D450N, D450A, or D467A (mutations D66N, D66A, or D83A in SEQ ID NO: 98) into the N-terminal or C-terminal ND1 of TALE-scND1. First, the TALE12-scND1n expression plasmids prepared in (7) of [Test Example 1] <1. Method>, which have a TALE repeat domain corresponding to Rosa26-L, were used: three types of TALE12(Rosa26-L)-ND1(D450N)-95-ND1, TALE12(Rosa26-L)-ND1(D450A)-95-ND1, and TALE12(Rosa26-L)-ND1(D467A)-95-ND1, in which ND1 at the N-terminal side was introduced with a mutation; and TALE12(Rosa26-L)-ND1-95-ND1(D450N) in which ND1 at the C-terminal side was introduced with a mutation. , 3 types of expression plasmids, TALE12(Rosa26-L)-ND1-95-ND1(D450A), TALE12(Rosa26-L)-ND1-95-ND1(D467A), a total of 6 expression plasmids, and gRNA-A or gRNA-B expression plasmids designed on the target gene Rosa26 (prepared in (9) of [Test Example 1]<1. Method>), and nCas9(D10A) expression plasmid (prepared in (8) of [Test Example 1]<1. Method>), and Rosa26-L was used as the TALE recognition sequence of TALE12-scND1n to implement SSA test ([Test Example 1]<1. Method>(11)(ii)). Figure 10 (A) is a conceptual diagram showing the positional relationship between the PAM sequence on the target gene (Rosa26), the cleavage point by nCas9 (D10A), the TALE (Rosa26-L), and the guide RNA.
[0357] The test results showed that the three types of ND1 introduced mutations on the N-terminal side confirmed strong double-stranded cleavage activity when acting with gRNA-A. On the other hand, the double-stranded cleavage activity was weak when combined with gRNA-B. When combined with gRNA-A, all three types of TALE12 (Rosa26-L) -ND1 (mutant) -95-ND1 preferentially added nicks to the TALE recognition chain. On the other hand, it was shown that the three types of ND1 introduced mutations on the C-terminal side confirmed equivalent double-stranded cleavage activity when combined with either gRNA-A or gRNA-B. All three types of TALE12 (Rosa26-L) -ND1-95-ND1 (mutant) added nicks to either the TALE recognition chain or the non-recognition chain ( Figure 11A ).
[0358] In addition, TALE12-scND1n expression plasmids with TALE repeat domains corresponding to Rosa26-R were used: 3 types of TALE12(Rosa26-R)-ND1(D450N)-95-ND1, TALE12(Rosa26-R)-ND1(D450A)-95-ND1, and TALE12(Rosa26-R)-ND1(D467A)-95-ND1 with mutations introduced into ND1 at the N-terminal side, and ND1 with mutations introduced into ND1 at the C-terminal side. Three types of expression plasmids, namely TALE12(Rosa26-R)-ND1-95-ND1(D450N), TALE12(Rosa26-R)-ND1-95-ND1(D450A), and TALE12(Rosa26-R)-ND1-95-ND1(D467A), for a total of six expression plasmids, were used to perform SSA test using Rosa26-R as the TALE recognition sequence of TALE12-scND1n ((11)(ii) of [Test Example 1]<1. Method>). Figure 10 (A) is a conceptual diagram showing the positional relationship between the PAM sequence on the target gene (Rosa26), the cleavage point by nCas9 (D10A), the TALE (Rosa26-R), and the guide RNA.
[0359] As a result of the test, the three ND1 mutations introduced at the N-terminal side all showed relatively strong double-stranded cleavage activity when combined with gRNA-D. On the other hand, in combination with gRNA-C, the double-stranded cleavage activity of TALE12(Rosa26-R)-ND1(D450N)-95-ND1 was weak, but relatively sufficient activity was confirmed in the other two, indicating that an incision was added to one of the TALE recognition chain or non-recognition chain. In addition, the three ND1 mutations introduced at the C-terminal side all showed strong double-stranded cleavage activity when combined with gRNA-C, while the double-stranded cleavage activity was weak when combined with gRNA-D, indicating that all three TALE12(Rosa26-R)-ND1-95-ND1 (mutants) preferentially added incisions to the TALE non-recognition chain ( Figure 11B ).
[0360] These results confirm that TALE12-ND1(D450N)-95-ND1 preferentially nicks the TALE recognition strand in both Rosa26-L and Rosa-R. "ND1(D450N)-95-ND1" will be used as "scND1n" from now on. However, nickase activity was confirmed in N-terminal ND1 mutants other than D450N and all C-terminal ND1 mutants, which does not negate their applicability.
[0361] Furthermore, regarding the TALE recognition sequences on other target genes: APC-L, APC-R, HPRT1-L, HPRT1-R, each TALE12-scND1n was prepared, and the gRNA-A, B, C, D designed on the target gene APC ( Figure 10 (B)) or gRNA-A, B, C, D designed on the target gene HPRT1 ( Figure 10 (C)), and nCas9 (D10A), and SSA test was performed using these target genes as targets ((11)(ii) of [Test Example 1] <1. Method>). Figure 10 (B) and (C) are conceptual diagrams showing the positional relationship between the PAM sequence on the target gene (B: APC, C: HPRT1), the cleavage point by nCas9 (D10A), and the TALE (B: APC-L / R, C: HPRT-1-L / R) and guide RNA.
[0362] As a result of the test, particularly strong double-strand cleavage activity was confirmed when TALE12(APC-L)-scND1n was combined with RNA-A, TALE12(APC-R)-scND1n was combined with gRNA-D, TALE12(HPRT1-L)-scND1n was combined with gRNA-A, and TALE12(HPRT1-R)-scND1n was combined with gRNA-D. Figure 12 、 Figure 13 On the other hand, when TALE12(APC-L)-scND1n was combined with gRNA-B, TALE12(APC-R)-scND1n was combined with gRNA-C, TALE12(HPRT1-L)-scND1n was combined with gRNA-B, and TALE12(HPRT1-R)-scND1n was combined with gRNA-C, the double-strand cleavage activity was weak ( Figure 12 、 Figure 13 ), showing that these TALE-scND1ns preferentially incorporate nicks specifically in the TALE recognition strand.
[0363] (ii) T7E1 test
[0364] Then, the T7E1 test was also performed in the genomic target to confirm that TALE-scND1n preferentially nicked the TALE recognition chain. For TALE12(off4-L)-scND1n prepared in (7) of [Test Example 1] <1. Method>, guide RNAs were designed at three locations ( Figure 14 (A), gRNA-off4-L3~5). It was investigated whether each guide RNA interacted with each Cas9 prepared in (8) of [Test Example 1]<1. Method> to exert its function. As a result, cleavage bands were confirmed in all guide RNAs, indicating that the guide RNA exerted its function ( Figure 14 (B)). Then, TALE12(off4-L)-scND1n, and any one of nCas9(D10A), nCas9(H840A), and dCas9(D10A+H840A) prepared in (8) of [Test Example 1]<1. Method> were allowed to act with each guide RNA to study which DNA chain TALE12(off4-L)-scND1n cuts. The positional relationship between the PAM sequence, the cleavage point by nCas9, the guide RNA, and the TALE of TALE12(off4-L)-scND1n is shown in Figure 14(A). nCas9 (D10A) adds a nick to the same strand as the guide RNA target sequence, and nCas9 (H840A) adds a nick to the complementary strand of the guide RNA target sequence. For gRNA-off4-L3, a cleavage band was confirmed when it acted with nCas9 (H840A). On the other hand, for gRNA-off4-L4 and L5, a cleavage band was confirmed when it acted with nCas9 (D10A) ( Figure 14 (C) Based on the above confirmation, TALE-scND1n also preferentially adds nicks to the TALE recognition strand in genomic targets.
[0365] (4) Confirmation of base substitution activity by TALE-scND1n and TALE-deaminase
[0366] In this test example, TALE-AID was used as a cytidine deaminase-type TALE-deaminase, and ABE8e-TALE was used as an adenosine deaminase-type TALE-deaminase. Figure 15 Conceptual diagram showing the structures of TALE-AID (A) and ABE8e-TALE (B).
[0367] In addition, respectively, Figure 16 (A) shows a conceptual diagram of base editing through collaboration between TALE-scND1n and TALE-AID. Figure 16 (B) shows a conceptual diagram of base editing by collaboration of TALE-scND1n and ABE8e-TALE. In these collaborative systems, each TALE (TALE-L, TALE-R) is designed to bind to the TALE recognition sequence sandwiching the same target base near the target base ((A): cytosine, (B): adenine) on the target DNA. In the case of applying TALE-AID ( Figure 16 (A)), when the TALE-scND1n bound to the TALE recognition sequence of the target DNA adds a nick to the DNA chain near the target base (TALE recognition chain), the DNA chain containing the target base becomes one chain, and the target base C is deaminated to U by the TALE-AID bound to the TALE recognition sequence of the target DNA, and is ultimately expected to be replaced by T. Similarly, in the case of ABE8e-TALE, a substitution from A to G of the target base is expected ( Figure 16 (B)).
[0368] Various reporters prepared in (12) of [Test Example 1] <1. Method>, with different lengths of spacer 2 (the distance between the target base and the TALE recognition sequence of TALE-deaminase or its complementary sequence) and spacer 1 (the distance between the target base and the TALE recognition sequence of TALE-scND1n or its complementary sequence) ( Figure 17 ), and the base substitution activity was detected. First, the reporter test of (13) (i) and (14) of [Test Example 1] <1. Method> was used to study whether base editing occurs through the collaboration of TALE-scND1n and TALE-deaminase. As reporters, Rosa26-L2~L5, R1 (Table 10) were used, as TALE-scND1n expression plasmids, the TALE12 (Rosa26-L) -scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method> was used, and as TALE-deaminase expression plasmids, the TALE-AID (TALE47-12-AID) expression plasmid prepared in (10) of [Test Example 1] <1. Method> was used ( Figure 15 (A) ), the study was conducted with a spacer 2 length of 13 bases and a spacer 1 length of 4, 7, 10, 13, and 16 bases.
[0369] The results of the test showed that even in the presence of TALE-AID (47-12-AID), no base substitution activity was detected in the absence of TALE-scND1n (w / o scND1n) or in the absence of TALE-scND1n (scND1n-NC). In addition, no base substitution activity was detected in the presence of TALE-scND1n (w / o 47-12-AID) alone. On the other hand, when both TALE-AID and TALE-scND1n bound near the target base and the length of spacer 1 was 7 to 10 bases, base substitution activity was detected ( Figure 18 These results suggest that base editing occurs through the collaboration of TALE-scND1n and TALE-deaminase.
[0370] (5) Study on the length of spacer 1 and spacer 2
[0371] Since base editing was shown to occur through cooperation between TALE-scND1n and TALE-deaminase, studies were performed to optimize the length of spacer 1 and spacer 2.
[0372] (TALE-AID)
[0373] As a reporter, CBE-L4M2-7bp~16bp (Table 11) was used, as a TALE-scND1n expression plasmid, the TALE12(Rosa26-LM.2)-scND1n expression plasmid prepared in (7) of [Test Example 1]<1. Method> was used, and as a TALE-deaminase expression plasmid, the TALE-AID(TALE47-12-AID) expression plasmid prepared in (10) of [Test Example 1]<1. Method> was used. Figure 15 (A)), the length of spacer 1 is 10 bases, and the length of spacer 2 is 7 to 16 bases. The reporter test of (13) (i) and (14) of [Test Example 1] <1. Method> was used for the study. As a result of the test, activity was detected when the length of spacer 2 was 8 to 15 bases, and particularly high activity was shown when the length was 10 bases and 11 bases ( Figure 19 ).
[0374] Then, in the case where the length of spacer 2, which showed relatively high activity, was 9 to 13 bases, the length of spacer 1 was also studied. As a TALE-scND1n expression plasmid, the TALE12 (Rosa26-LM.3.2 to LM.2+4)-scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method> was applied. As a result, particularly high activity was confirmed when the length of spacer 2 was 10 to 11 bases and the length of spacer 1 was 9 to 11 bases ( Figure 20 The activity is particularly high when the distance between the two TALE recognition sequences is 20 to 23 bases, and the maximum activity is shown when the distance is 22 bases.
[0375] (ABE8e-TALE)
[0376] As a reporter, ABE-L4M2-7bp~16bp (Table 11) was used, as a TALE-scND1n expression plasmid, the TALE12(Rosa26-LM1.2)-scND1n expression plasmid prepared in (7) of [Test Example 1]<1. Method> was used, and as a TALE-deaminase expression plasmid, the ABE8e-TALE(ABE8e-32-TALE47) expression plasmid prepared in (10) of [Test Example 1]<1. Method> was used. Figure 15 (B)), the length of spacer 1 was 9 bases, and the length of spacer 2 was 7 to 16 bases. The reporter test of (13)(i) and (14) of [Test Example 1] <1. Method> was used to study the activity. As a result of the test, activity was detected at all lengths of spacer 2, and particularly high activity was shown at 11 bases and 12 bases ( Figure 21 ).
[0377] Then, when the length of spacer 2 was 9 to 14 bases, the length of spacer 1 was also studied. As a TALE-scND1n expression plasmid, the TALE12 (Rosa26-LM.4.2 to LM.2+4)-scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method> was used. As a result, particularly high activity was confirmed when the length of spacer 2 was 10 to 12 bases and the length of spacer 1 was 9 to 11 bases ( Figure 22 The activity was particularly high when the distance between the two TALE recognition sequences was 20 to 23 bases, and the maximum activity was shown at 21 and 22 bases.
[0378] (6) Base editing of endogenous (nuclear) DNA by TALE-scND1n and TALE-deaminase
[0379] Since the above-mentioned reporter test showed that the cooperative system of TALE-scND1n and TALE-deaminase has base editing activity, the endogenous DNA editing test of (13)(ii) and (15) of [Test Example 1] <1. Method> was used to study whether base editing can also be performed in endogenous (nuclear) DNA. First, regarding TALE-AID, as a TALE-scND1n expression plasmid, the TALE12-scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method> was used, and TALE had three different lengths of TALE repeat domains corresponding to the TALE recognition sequence on APC: APC-L2-1~3 ( Figure 23 (A)). In addition, Figure 23 The sequence of the target site (APC) described in (A) is shown in SEQ ID NO: 182. In addition, as a TALE-deaminase expression plasmid, the TALE-AID (TALE47-12-AID) prepared in (10) of [Test Example 1] <1. Method> was used, and the TALE was a TALE having a TALE repeat domain corresponding to the TALE recognition sequence on APC: APC-R ( Figure 23 (A)). Thus, the results of the three types of distances between the two TALE recognition sequences, 23, 24, and 25 bases, showed that all of them were within a wide range of C (C3 to C19: Figure 23 When the distance between the two TALE recognition sequences was 23 bases, a particularly high activity was confirmed. The ratio of T in each C was about 40% in C11 to C13 ( Figure 23 (B)).
[0380] [Test Example 2]
[0381] 1. Method
[0382] (1) Preparation of an expression plasmid in which ABE8e (V108W) is located at the C-terminus of TALE
[0383] In order to study the improvement of the activity of ABE8e-TALE (ABE8e-32-TALE47) prepared from (10) of [Test Example 1] <1. Method>, four types of TALE-ABE8e (TALE16-XTEN-ABE8e, TALE47-XTEN-ABE8e, TALE47-ABE8e, TALE47-12-ABE8e) expression plasmids were prepared in which ABE8e (V108W) was configured on the C-terminal side of TALE. Specifically, the ABE8e-TALE (ABE8e-32-TALE47) expression plasmid prepared in (10) of [Test Example 1] <1. Method> was used as a template, and the primer set described in Table 13 below (combination of ABE_F_infXTEN and ABE_R_infCterm) was used to amplify the gene encoding ABE8e (V108W, nucleotide sequence number: 193, amino acid sequence number: 194). Sequences encoding the respective TALEs (TALE-153 / 16 (nucleotide sequence number: 195, amino acid sequence number: 196), TALE-153 / 47) were amplified by PCR using the TALE-153 / 47 expression plasmid prepared in (1) of [Test Example 1] <1. Method> as a template and the primer sets described in Table 13 below. Furthermore, an oligonucleotide (nucleotide sequence number: 197) encoding a 2aa+XTEN linker (amino acid sequence number: 198) that adds two amino acids to the XTEN linker was synthesized and annealed to its complementary sequence. These nucleotide fragments were combined and ligated using the In-Fusion method to create three TALE-ABE8e expression plasmids (TALE16-XTEN-ABE8e, TALE47-XTEN-ABE8e, and TALE47-ABE8e).
[0384] In addition, by the In-Fusion method, the fragment (ABE8e (V108W)) amplified by PCR using the above-mentioned ABE8e-TALE expression plasmid as a template and the primer set described in the following Table 13 (a combination of ABE_F and ABE_R_infCterm) was linked to the fragment amplified by PCR using the TALE-AID (TALE47-12-AID) expression plasmid prepared in (10) of [Test Example 1] <1. Method> as a template and the primer set described in the following Table 13, to prepare a TALE-ABE8e (TALE47-12-ABE8e) expression plasmid.
[0385] The templates and primers used in the preparation of these plasmids are shown in the following Table 13. For evaluation by reporter testing, a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence (APC-R) shown in Table 8 above was inserted into each of these expression plasmids.
[0386]
[0387] (2) Preparation of TALE-BspD6I for evaluation of incision activity
[0388] A construct for evaluating the nicking activity of BspD6I (Nt.BspD6I (C) described in Document 1) was prepared. Here, in order to carry out activity comparison, BspD6I was configured on the C-terminal side of TALE in the same manner as TALE12-scND1n prepared in (7) of [Test Example 1] <1. Method>. The nucleotide sequence encoding the nicking activity domain of the BspD6I gene (nucleotide sequence number: 199, amino acid sequence number: 200) was fully synthesized and amplified by PCR using the primer set described in Table 14 below. In addition, TALE-153 / 47 prepared in (1) of [Test Example 1] <1. Method> was used as a template and amplified by PCR using the primer set described in Table 14 below. These nucleotide fragments were linked by the In-Fusion method to prepare a TALE12-BspD6I expression plasmid. The conceptual diagram of the structure of the protein (TALE12-BspD6I) expressed in the expression plasmid is shown in Figure 24 .
[0389] The templates and primers used in the preparation of these plasmids are shown in Table 14 below. The TALE12-BspD6I expression plasmid used in the nicking activity evaluation contained a TALE repeat domain corresponding to the nucleotide sequence of the TALE recognition sequence (Rosa26-L, R) on Rosa26 shown in Table 8 above. The nicking activity of BspD6I in the nucleus was evaluated according to the SSA test in (11) (ii) of [Test Example 1] <1. Method>.
[0390]
[0391] (3) Preparation of TALE-BspD6I for evaluation of base substitution activity
[0392] To evaluate base substitution activity, a fusion protein of a TALE having the amino acid sequence described in Document 1 and BspD6I was used. The amino acid sequence of this TALE differs partially from the C-terminal domain sequence of TALE-136 / 63 prepared in (1) of [Test Example 1] <1. Method>, but it can be said that they are subspecies derived from the same biological species. A mutation was introduced into the TALE-136 / 63 expression plasmid prepared in (1) of [Test Example 1] <1. Method>, and the C-terminal domain was reduced to 41 amino acids. A TALE-136 / 41N (nucleotide sequence number: 205, amino acid sequence number: 206) expression plasmid was prepared, and BspD6I prepared in (2) of [Test Example 2] <1. Method> was added to its C-terminus. Specifically, the TALE12-BspD6I expression plasmid prepared in (2) of [Test Example 2] <1. Method> was used as a template and amplified by PCR using the primer set listed in Table 15 below. Furthermore, the TALE-136 / 63 expression plasmid prepared in (1) of [Test Example 1] <1. Method> was used as a template and the primer sets listed in Table 15 below were used to amplify the N-terminal, internal, and C-terminal sides of the TALE by PCR. These nucleotide fragments were combined and ligated by the In-Fusion method to prepare the TALE41N-BspD6I expression plasmid.
[0393] In the analysis of mitochondrial targets and the analysis of nuclear genomic targets, the above-mentioned TALE41N-BspD6I expression plasmid, the TALE12-scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method>, the TALE47-12-ABE8e expression plasmid prepared in (1) of [Test Example 2] <1. Method>, and the TALE47-12-AID expression plasmid prepared in (10) of [Test Example 1] <1. Method> were used. In the analysis of mitochondrial targets, TALE repeat domains corresponding to the nucleotide sequences of the TALE recognition sequences listed in Table 16 below were inserted into these plasmids. Similarly, in the analysis of nuclear genomic targets, TALE repeat domains corresponding to the nucleotide sequences of the TALE recognition sequences listed in Table 17 below were inserted.
[0394] Furthermore, regarding the aforementioned TALE41N-BspD6I expression plasmid and TALE12-scND1n expression plasmid, a mitochondrial expression type mitoTALE41N-BspD6I expression plasmid and a mitoTALE12-scND1n expression plasmid were prepared, in which SOD-MTS+3×FLAG was added instead of 3×FLAG+NLS added to the N-terminus. Furthermore, regarding the TALE47-12-ABE8e expression plasmid prepared in (1) of [Test Example 2] <1. Method>, a mitochondrial expression type mitoTALE47-12-ABE8e expression plasmid was prepared, in which COX8A-MTS+3×HA was added instead of 3×FLAG+NLS added to the N-terminus. Specifically, an artificially synthesized SOD-MTS+3×FLAG sequence (nucleotide sequence number: 220, amino acid sequence number: 221) was used as a template, and PCR amplification was performed using the primer sets listed in Table 15 below. In addition, the COX8A-MTS + 3 × HA sequence (nucleotide sequence number: 222, amino acid sequence number: 223) was artificially synthesized as a template and amplified by PCR using the primer set described in Table 15 below. Furthermore, the above-mentioned TALE41N-BspD6I expression plasmid was used as a template and amplified by PCR using the primer set described below. In addition, the TALE12-scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method> and the TALE47-12-ABE8e expression plasmid prepared in (1) of [Test Example 2] <1. Method> were used as templates, respectively, and amplified by PCR using the primer set described in Table 15 below. These nucleotide fragment combinations were linked by the In-Fusion method to prepare mitoTALE41N-BspD6I expression plasmid, mitoTALE12-scND1n expression plasmid, and mitoTALE47-12-ABE8e expression plasmid, respectively.
[0395] The templates and primers used in the preparation of these plasmids are shown in the following Table 15. Into these mitochondrial expression plasmids, TALE repeat domains corresponding to the nucleotide sequences of the TALE recognition sequences on the respective mitochondrial targets shown in the following Table 16 were inserted.
[0396]
[0397]
[0398]
[0399] (4) Cloning of mitochondrial target sequences
[0400] Using HEK293T cell genomic DNA as a template, PCR was performed using the primer sets listed in Table 18 to generate amplified DNA fragments: MT-ND1 (nucleotide sequence number: 246) and MT-ND4site2 (nucleotide sequence number: 247). These fragments were cloned into the EcoRV recognition site of pBlueScriptII (SK+) (Stratagene, USA). The plasmids were designated pBS / MT-ND1 and pBS / MT-ND4site2, respectively. Both MT-ND1 and MT-ND4site2 are mitochondrial-specific sequences (mitochondrial targets).
[0401]
[0402] (5) Confirmation of base substitution activity on mitochondrial targets
[0403] HEK293T cells grown in DMEM medium containing 10% FBS were cultured at 1 × 10 cells / mL on the day before transfection. 4 Each well of a 96-well plate was inoculated with 100 ng of the mitoTALE41N-BspD6I expression plasmid or 25 ng of the mitoTALE12-scND1n expression plasmid prepared in (3) of [Test Example 2] <1. Method>, which contains a TALE repeat domain sequence for each TALE recognition sequence, and a combination of 25 ng of the mitoTALE47-12-ABE8e expression plasmid were introduced into HEK293T cells using Lipofectamine 3000 (manufactured by Thermo Fisher Scientific).
[0404] In addition, a combination of 25 ng of a TALE41N-BspD6I expression plasmid or a TALE12-scND1n expression plasmid that is not a mitochondrial expression type and a TALE47-12-ABE8e expression plasmid, which has a TALE repeat domain sequence inserted into the TALE recognition sequence, and 10 ng of plasmids (pBS / MT-ND1, pBS / MT-ND4site2) for cloning each mitochondrial target sequence prepared in (4) of [Test Example 2] <1. Method> was used to make a total of 100 ng through pcNDA3.1s and introduced into HEK293T cells.
[0405] Cells were collected 48 hours after each introduction, and genomic DNA (including mitochondrial DNA and plasmid DNA) was purified using PureLink Genomic DNA Kits (manufactured by Thermo Fisher Scientific). Using 10 ng of each genomic DNA as a template, KOD-ONE (manufactured by TOYOBO) was used to amplify the target on the plasmid (pBS) and the target in the mitochondria (MT-ND1 (nucleotide sequence number of the PCR product: 248) and MT-ND4site2 (nucleotide sequence number of the PCR product: 249)) by PCR using the EditR amplification primer set described in Table 19 below. The obtained PCR product was purified using NucleoSpin Gel and PCR Clean-up (manufactured by MACHERY-NAGEL), and 10 ng of it was used as a template. The sequencing primers described in Table 19 below were used to sequence MT-ND1 on the plasmid and MT-ND4site2 in the mitochondria, respectively, using BigDye Terminator V3.1 (manufactured by Thermo Fisher Scientific). The obtained sequence data were analyzed using EditR on a 3730×l DNA Analyzer (manufactured by Thermo Fisher Scientific), and the ratio (%) of G in each A between each TALE recognition sequence was calculated.
[0406]
[0407] (6) Confirmation of base substitution activity on nuclear genomic targets
[0408] As in (5) of [Test Example 2] <1. Method>, 25 ng of the TALE41N-BspD6I expression plasmid or TALE12-scND1n expression plasmid prepared in (3) of [Test Example 2] <1. Method> and inserted with the TALE repeat domain sequence for each TALE recognition sequence described in Table 17, and 50 ng of the TALE47-12-ABE8e expression plasmid or the TALE47-12-AID expression plasmid prepared in (10) of [Test Example 1] <1. Method> were diluted to a total of 100 ng through pBS and introduced into HEK293T cells. The recovered and purified genomic DNA was used as a template, and the four targets (ABE-site9, BCL11A, VEGFA-3, HEK-1) on the nuclear genome were amplified by PCR using the EditR amplification primer set described in the following Table 20. Sequencing was performed using the sequencing primers listed in Table 20 below, and EditR analysis was performed to calculate the ratio (%) of T in each C or the ratio (%) of G in each A among the TALE recognition sequences.
[0409]
[0410] (7) Preparation of TALE-deaminase expression plasmid for base substitution of endogenous (nuclear) DNA
[0411] Separately, a TALE47-12-AID expression plasmid for base substitution of endogenous DNA was prepared in the same manner as (10) in <1. Method> of [Test Example 1] above, and a TALE47-12-ABE8e expression plasmid for base substitution of endogenous DNA was prepared in the same manner as (1) in <1. Method> of [Test Example 2] above. TALE repeat domains corresponding to the nucleotide sequences of the respective TALE recognition sequences listed in Table 21 below were inserted into these expression plasmids.
[0412]
[0413] (8) Preparation of TALE-scND1n expression plasmid for base substitution of endogenous (nuclear) DNA
[0414] Plasmids expressing TALE12-scND1n for base substitution of endogenous DNA were prepared in the same manner as in (7) of [Test Example 1] <1. Method>. TALE repeat domains corresponding to the nucleotide sequences of the TALE recognition sequences listed in Table 22 below were inserted into these expression plasmids.
[0415]
[0416] (9) Preparation of TALE-VPR
[0417] An artificially synthesized nucleotide sequence encoding the VPR protein (nucleotide sequence number: 312, amino acid sequence number: 313) was used as a template and amplified by PCR using the primer set listed in Table 23 below. Furthermore, the TALE-153 / 47 expression plasmid prepared in (1) of [Test Example 1] <1. Method> was used as a template and amplified by PCR using the primer set listed in Table 23 below. These nucleotide fragments were ligated by the In-Fusion method to prepare a TALE47-VPR (TALE47-12-VPR) expression plasmid.
[0418] The templates and primers used in the preparation of the above plasmids are shown in Table 23 below. In order to confirm the effect of promoting the base substitution activity of TALE-VPR, a TALE repeat domain corresponding to the nucleotide sequence of the three TALE recognition sequences listed in Table 24 below, which was set further outside the TALE recognition sequence bound by TALE12-scND1n in the APC locus, was inserted into the TALE47-VPR expression plasmid.
[0419]
[0420]
[0421] (10) Preparation of reporter plasmid for detecting base substitution activity of TALE-ABE8e
[0422] Oligonucleotides designed to correspond to the sequences shown in Table 25 below were annealed and inserted into pNLF-M / A treated with the restriction enzymes NheI and XhoI to prepare reporter plasmids containing various sequences (sequences shown in Table 25) for studying the length of spacer 2 (the distance between the target base and the TALE recognition sequence of the TALE-deaminase). The inserted sequence consisted of the TALE recognition sequence of the TALE-deaminase (underlined (*1) in Table 25), a codon containing the target base (underlined (*2) in Table 25: 5'-TAG-3'), a complementary sequence to the TALE recognition sequence of the TALE-scND1n (underlined (*3) in Table 25), and a spacer sequence therebetween. The length of spacer 2 ranged from 7 to 16 bases. In this reporter plasmid, the target codon is located at a position corresponding to the stop codon upstream of the NanoLuc luciferase insertion. According to this reporter plasmid, when the target codon TAG is converted to TGG (the A at the target base (target site) is replaced by G) by the TALE-deaminase that binds to the TALE recognition sequence, NanoLuc luciferase is expressed.
[0423]
[0424] (11) Transfection of HEK cells and analysis of base editing activity
[0425] (i) Reporter test
[0426] The reporter test was performed in the same manner as (13)(i) and (14) of the above-mentioned [Test Example 1] <1. Method>, except that each reporter plasmid prepared in (11) of the above-mentioned [Test Example 2] <1. Method> was used as the reporter plasmid, and TALE-ABE8e prepared in (1) of the above-mentioned [Test Example 2] <1. Method> was used as the TALE-deaminase expression plasmid.
[0427] (ii) Base editing testing of endogenous (nuclear) DNA
[0428] In the case where endogenous (nuclear) DNA is used as the target DNA, 60 ng of the TALE47-12-AID expression plasmid prepared in (7) of the above-mentioned [Test Example 2] <1. Method>; 7.5 ng or 1.5 ng (VEGFA3 site) of the TALE12-scND1n expression plasmid prepared in (8) of the above-mentioned [Test Example 2] <1. Method>; 60 ng of the TALE47-12-ABE8e expression plasmid prepared in (7) of the above-mentioned [Test Example 2] <1. Method>; and 30 ng of the TALE12-scND1n expression plasmid prepared in (8) of the above-mentioned [Test Example 2] <1. Method> are combined, and transfection of HEK293T cells and preparation of cell lysate are performed in the same manner as in (13)(ii) and (15) of the above-mentioned [Test Example 1] <1. Method>. Using the cell lysate as a template, the primers listed in Tables 26 and 12 below were used to amplify the region encompassing the target site of the endogenous (nuclear) DNA using KODFXNeo (manufactured by TOYOBO). The resulting PCR product was purified and sequenced, and the sequence data was analyzed using EditR to calculate the ratio (%) of T per C or G per A within the target region of the target site.
[0429]
[0430] <2. Results>
[0431] (1) Evaluation of the nickase activity of BspD6I in the nucleus
[0432] The same method as in (11)(ii) of [Test Example 1] <1. Method> was used to evaluate the combination of the TALE12(Rosa26-L)-scND1n (scND1: ND1(D450N)-95-ND1, the same below) expression plasmid prepared in (7) of [Test Example 1] <1. Method> and the TALE12(Rosa26-R)-scND1n expression plasmid prepared in (2) of [Test Example 2] <1. Method>. Double nicking activity of a combination of a 2(Rosa26-L)-BspD6I expression plasmid and a TALE12(Rosa26-R)-BspD6I expression plasmid, or a combination of the same TALE12(Rosa26-L, R)-BspD6I expression plasmid and the nCas9(D10A) expression plasmid prepared in (8)(9) of [Test Example 1]<1. Method> and the guide RNA (gRNA-A, B, D) designed on the target gene Rosa26.
[0433] The graph showing the EGFP fluorescence intensity in each combination is shown in Figure 25 .like Figure 25 As shown, TALE12-scND1n alone does not exhibit double nicking activity, but by combining nCas9 with the corresponding guide RNA, or by combining the TALE repeat domain with two TALE recognition sequence binding domains (L and R) sandwiching the target base, it exhibits activity comparable to that of the CRISPR-nCas9 system. On the other hand, no such activity was confirmed with any of the TALE12-BspD6I combinations, and the nicking activity of BspD6I in the nucleus could not be detected.
[0434] (2) Evaluation of base substitution activity of BspD6I / scND1n on mitochondrial targets MT-ND1 and MT-ND4site2
[0435] The ratio (%) of G in each A when each TALE-nickase (TALE41N-BspD6I or TALE12-scND1n) was introduced for each target in the mitochondria (mitochondorial) and each target on the introduced plasmid (plasmid) using the method (5) of [Test Example 2] <1. Method> is shown in Table 27 below. In this specification, "mock" refers to each ratio in cells introduced with 100 ng of pBS alone for mitochondrial targets and to each ratio in cells introduced with 10 ng of pBS / MT-ND1 or pBS / MT-ND4site2 and 90 ng of pcNDA3.1s for targets on plasmids. In addition, in MT-ND1, the distance between TALE recognition sequences when TALE41N-BspD6I is applied is 16 bases as described in Document 1, and A2 to A14 are within this range. The distance between TALE recognition sequences when TALE12-scND1n is applied is 22 bases, and A2 to A19 are within this range. In addition, in MT-ND4site2, the distance between TALE recognition sequences when TALE41N-BspD6I is applied is 15 bases as described in Document 1, and A4 to A13 are within this range. The distance between TALE recognition sequences when TALE12-scND1n is applied is 22 bases, and A4 to A21 are within this range.
[0436]
[0437] As shown in Table 27, BspD6I exhibited base substitution activity for targets within mitochondria, but no significant activity was detected with scND1n. Conversely, scND1n exhibited base substitution activity for targets on plasmids, but no significant activity was detected with BspD6I. Thus, for the same mitochondrial target, BspD6I and scND1n exhibited significant differences in base substitution activity depending on target localization, confirming the presence of location-dependent nicking activity. Furthermore, this result does not contradict the results of (1) in [Test Example 2] <2. Results>.
[0438] (3) Evaluation of base substitution activity of BspD6I / scND1n on four nuclear genomic targets
[0439] According to the method (6) of [Test Example 2] <1. Method>, the base substitution activity of TALE47-12-ABE8e was evaluated when each TALE-nickase (TALE12-scND1n or TALE41N-BspD6I) was used to cooperate with the nuclear genomic targets: ABE-site-9 and BCL11A. The ratio (%) of G in each A is shown in Table 28 below. In addition, the base substitution activity of TALE47-12-AID was evaluated when each TALE-nickase (TALE12-scND1n or TALE41N-BspD6I) was used to cooperate with the nuclear genomic targets: VEGFA-3 and HEK-1. The ratio (%) of T in each C is shown in Table 29 below. In addition, when TALE41N-BspD6I is applied, the distance between the TALE recognition sequences in ABE-site-9, VEGFA-3, and HEK-1 is 15 bases, and the distance between the TALE recognition sequences in BCL11A is 16 bases, and A1 to A15 or C1 to C13 are within this range. When TALE12-scND1n is applied, the distance between the TALE recognition sequences in ABE-site-9 and BCL11A is 22 bases, and the distance between the TALE recognition sequences in VEGFA-3 and HEK-1 is 20 bases, and A1 to A22 or C1 to C17 are within this range.
[0440]
[0441]
[0442] As shown in Tables 28 and 29, scND1n showed base substitution activity against all four nuclear genomic targets, but BspD6I showed no significant base substitution activity against any of them. In this experiment, BspD6I was found to have no nicking activity at 4 / 4 of the nuclear genomic targets.
[0443] (4) Comparison of base editing activity of various TALE-ABE8e structures
[0444] [Test Example 2] A conceptual diagram of the structure of each TALE-ABE8e (TALE47-12-ABE8e, TALE16-XTEN-ABE8e, TALE47-XTEN-ABE8e, TALE47-ABE8e) expressed in the expression plasmid prepared in (1) of <1. Method> is shown in FIG. Figure 26 In addition, a conceptual diagram of base editing through the collaboration of TALE-scND1n and TALE-ABE8e is shown in Figure 27 .
[0445] As the reporter plasmid, ABE-C-L4M2-8bp~16bp (for TALE-ABE8e, Table 25) prepared in (10) of [Test Example 2] <1. Method> or ABE-L4M2-8bp~16bp (for ABE8e-TALE, Table 11) prepared in (12) of [Test Example 1] <1. Method> was used as the TALE-scND1n expression plasmid, and TALE12(Rosa26-LM.1.2)-scND1n expression plasmid prepared in (7) of [Test Example 1] <1. Method> was used as the TALE-deaminase expression plasmid. The enzyme expression plasmid was prepared using the TALE47-12-ABE8e expression plasmid, TALE16-XTEN-ABE8e expression plasmid, TALE47-XTEN-ABE8e expression plasmid, TALE47-ABE8e expression plasmid prepared in (1) of [Test Example 2] <1. Method>, or the ABE8e-32-TALE47 expression plasmid prepared in (10) of [Test Example 1] <1. Method>, with the length of spacer 1 being 9 bases and the length of spacer 2 being 8 to 16 bases, and the reporter test of (11) (i) of [Test Example 2] <1. Method> was performed. The results are shown in Figure 28 .
[0446] like Figure 28 As shown, all cases showed base substitution activity, but among the various TALE-ABE8e structures, TALE47-12-ABE8e had the highest base substitution activity. Furthermore, compared to ABE8e-32-TALE47, in which the deaminase was fused to the N-terminus, TALE47-12-ABE8e, which had the deaminase fused to the C-terminus, showed a tendency to have higher activity. Therefore, in subsequent experiments, TALE47-12-ABE8e was used as the TALE-ABE8e.
[0447] (5) Study on the length of spacer 1 and spacer 2
[0448] Subsequently, the optimization of the length of spacer 1 and the length of spacer 2 was studied. As a reporter plasmid, ABE-C-L4M2-7bp~16bp (for TALE-ABE8e, Table 25) prepared in (10) of [Test Example 2]<1. Method> was used, as a TALE-scND1n expression plasmid, TALE12(Rosa26-LM.4.2~LM.2+3)-scND1n expression plasmid prepared in (7) of [Test Example 1]<1. Method> was used, as a TALE-deaminase expression plasmid, TALE47-12-ABE8e expression plasmid prepared in (1) of [Test Example 2]<1. Method> was used, the length of spacer 1 was 6 to 13 bases, and the length of spacer 2 was 7 to 16 bases, and the reporter test of (11)(i) of [Test Example 2]<1. Method> was performed. The results are shown in Figure 29 .
[0449] like Figure 29 As shown, particularly high base substitution activity was observed when the length of spacer 1 was 8 to 11 bases and when the length of spacer 2 was 9 to 13 bases. Base substitution activity was high when the distance between the two TALE recognition sequences was 19 to 23 bases, and particularly high when the distances were 20 and 21 bases.
[0450] (6) Study on the optimal distance between TALEs in endogenous (nuclear) DNA
[0451] The results of each reporter test showed that when the distance between TALE recognition sequences in TALE-AID is 20 to 23 bases ( Figure 20 ), when the distance between TALE recognition sequences in TALE-ABE8e is 19 to 23 bases ( Figure 29 ), and the base substitution activity is particularly high, and thus, the distance between the most suitable TALE recognition sequences in the endogenous (nuclear) DNA was studied according to the method of (11) (ii) endogenous DNA editing test of [Test Example 2] <1. Method>. That is, as a TALE-AID expression plasmid, TALE47-12-AID prepared in (7) of [Test Example 2] <1. Method> was applied, as a TALE-ABE8e expression plasmid, TALE47-12-ABE8e prepared in (7) of [Test Example 2] <1. Method> was applied, and as a TALE-scND1n expression plasmid, TALE12-scND1n expression plasmid prepared in (8) of [Test Example 2] <1. Method> was applied. In Figure 30 (A) to (D) show the positional relationship between the TALE recognition sequence of TALE-deaminase and the TALE recognition sequence of TALE-scND1n in each endogenous (nuclear) target site (HEK1, VEGFA3, ABE-site9, BCL11A). Figure 30 The sense strands of the nucleotide sequences described in (A) to (D) are shown in nucleotide sequence numbers 360 to 363. The distance between TALE recognition sequences is 16 to 24 bases ( Figure 30 (A), (B)), or 17 to 22 bases ( Figure 30 (C), (D)). The TALE recognition sequence of each TALE-deaminase is the sequence described in Table 21 above, and the TALE recognition sequence of each TALE-scND1n is the sequence described in Table 22 above. The ratio (%) of T in each C of the target base between the two TALE recognition sequences or the ratio (%) of G in each A of the target base is shown in Figures 31-35 .
[0452] In TALE-AID, when the distance between TALEs in the HEK1 site is 19 bases ( Figure 31 ), VEGFA3 site and BCL11A site 20 bases ( Figure 32 、 Figure 33 ), respectively showing the maximum base substitution activity. In addition, in TALE-ABE8e, when the distance between TALEs in ABE-site9 is 21 bases ( Figure 34 ), 20 bases in the BCL11A site ( Figure 35 ), showing the maximum base substitution activity, respectively.
[0453] (7) Verification of the versatility of base editing through collaboration between TALE-scND1n and TALE-deaminase
[0454] Based on the results of (6) of the above-mentioned [Test Example 1]<2. Results>, a base editing experiment was conducted in other endogenous (intranuclear) target sites through the collaboration of TALE-scND1n and TALE-deaminase to study the versatility of the present invention. As a TALE-AID expression plasmid, the TALE47-12-AID prepared in (7) of [Test Example 2]<1. Method> was applied; as a TALE-ABE8e expression plasmid, the TALE47-12-ABE8e prepared in (7) of [Test Example 2]<1. Method> was applied; as a TALE-scND1n expression plasmid, the TALE12-scND1n expression plasmid prepared in (8) of [Test Example 2]<1. Method> was applied. Figure 36 (A) to (G) show the positional relationship between the TALE recognition sequence of TALE-deaminase and the TALE recognition sequence of TALE-scND1n in each endogenous (nuclear) target site (AAVS1-2, APC, MALTA1, RPCI, ABE-site5, HEK2, DYRK1A). Figure 36The sense strands of the nucleotide sequences described in (A) to (G) are shown in nucleotide sequence numbers 364 to 370. The distance between TALE recognition sequences is 20 bases in TALE-AID ( Figure 36 (A) to (D)), 20 bases or 21 bases in TALE-ABE8e ( Figure 36 (E) to (G)). The TALE recognition sequence added to the TALE of each TALE-deaminase is the sequence described in Table 21 above, and the TALE recognition sequence added to the TALE of each TALE-scND1n is the sequence described in Table 22 above. The results of TALE-AID (the ratio (%) of T in each C of the target base between the two TALE recognition sequences) are shown in Table 21. Figure 37 The results in TALE-ABE8e (the ratio (%) of G in each A of the target base between the two TALE recognition sequences) are shown in Figure 38 .
[0455] like Figure 37 and Figure 38 As shown, sufficient base substitution activity was observed at multiple target sites for both TALE-AID and TALE-ABE8e, confirming that the DNA editing system of the present invention comprising TALE-scND1n and TALE-deaminase is a technology with high versatility for nuclear DNA.
[0456] (8) Promotion of base editing activity by TALE-VPR through collaboration between TALE-scND1n and TALE-deaminase
[0457] In the DNA editing system of the present invention comprising TALE-scND1n and TALE-deaminase, the effect of promoting base editing activity by TALE-VPR was studied. The APC site was selected as the endogenous (intranuclear) target site, and the TALE of TALE-VPR was designed to bind upstream of TALE-scND1n with a distance of 15 to 56 bases between the TALE recognition sequences. Figure 39The positional relationship of the TALE recognition sequences of the TALE-deaminase prepared in (7) of [Test Example 2] <1. Method>, the TALE-scND1n prepared in (7) of [Test Example 1] <1. Method>, and the TALE recognition sequences of the TALE-VPR prepared in (9) of [Test Example 2] <1. Method> at the APC site is shown. The TALE of the TALE-deaminase has a TALE repeat domain corresponding to the TALE recognition sequence at the APC site: APC-R3 (Table 21), and the TALE of the TALE-scND1n has a TALE repeat domain corresponding to the TALE recognition sequence at the APC site: APC-L2-3 (Table 8, Table 22). The TALE of the TALE-VPR has a TALE repeat domain corresponding to the TALE recognition sequences at the APC site: APC-TA-L1 to L3 (Table 24). Separately, the results of TALE-AID obtained in the method of (11)(ii) endogenous DNA editing test of [Test Example 2] <1. Method> (ratio (%) of T in each C of the target base between the two TALE recognition sequences) are shown in Figure 40 The results in TALE-ABE8e (the ratio (%) of G in each A of the target base between the two TALE recognition sequences) are shown in Figure 41 .
[0458] like Figure 40 and Figure 41 As shown, in all TALE-deaminases, the base substitution activity was enhanced by TALE-VPR. In particular, in TALE-ABE8e, which has relatively low base substitution activity, the activity promotion effect was significant. Based on these results, the DNA editing system of the present invention confirmed the promotion effect of TALE-VPR base editing activity.
[0459] Industrial applicability
[0460] As described above, the nickase of the present invention has nickase activity with 1 molecule, for example, by becoming a fusion protein with a DNA binding domain, site-specific nickase activity is exerted. Therefore, the fusion protein can replace the nicking function of nickase type Cas9, by cutting the target DNA (target DNA in the cell other than the mitochondrial DNA) only by a single-stranded protein, and the DNA editing by various nucleic acid editing enzymes in the target site of the periphery becomes possible. For example, by combining TALE-deaminase as a nucleic acid editing enzyme, a base substitution technology on the genome only by protein can be provided, and by applying various deaminases, the base editing of the target site can be realized Cas9-independently. Thus, according to the present invention, it is possible to provide a method that does not rely on CRISPR-Cas systems such as Cas9, can specifically and effectively edit the target DNA in the cell other than the mitochondrial DNA by nucleic acid base transferase only by protein, and a nickase useful thereto. Thus, the present invention, as an excellent genome editing tool, is expected to be applied in a wide range of industrial fields such as medical treatment, agriculture, and industry.
Claims
1. A nickase comprising two domains, ND1 and ND1n, bound via an ND1 linker, wherein: ND1 is at least one polypeptide selected from the following (a) to (c): (a) a polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 98, (b) a polypeptide comprising the following amino acid sequence: one or more of the amino acid sequences described in SEQ ID NO: 98 are substituted, deleted, inserted, and / or added, and the amino acids corresponding to positions 66 and 83 of the amino acid sequence described in SEQ ID NO: 98 are aspartic acid, (c) a polypeptide comprising an amino acid sequence having 80% or more homology to the amino acid sequence of SEQ ID NO: 98, wherein the amino acids corresponding to amino acids 66 and 83 of the amino acid sequence of SEQ ID NO: 98 are aspartic acid, Furthermore, ND1n is at least one polypeptide selected from the following (an) to (cn): (an) a polypeptide comprising the following amino acid sequence: aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted with any other amino acid, (bn) A polypeptide comprising the following amino acid sequence: one or more of the amino acid sequences described in SEQ ID NO: 98 are substituted, deleted, inserted, and / or added, and the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence described in SEQ ID NO: 98 is substituted with any amino acid other than aspartic acid; (cn) A polypeptide comprising the following amino acid sequence: having more than 80% homology with the amino acid sequence recorded in sequence number 98, and the amino acid corresponding to aspartic acid at position 66 and / or aspartic acid at position 83 of the amino acid sequence recorded in sequence number 98 is a substituted amino acid substituted by any amino acid other than aspartic acid.
2. The nicking enzyme according to claim 1, wherein The length of the ND1 linker is 30 to 300 amino acid residues. A fusion protein comprising a first DNA binding domain and the nicking enzyme of claim 1 .
4. A DNA editing system comprising: The first fusion protein as the fusion protein of claim 3, and A second fusion protein comprises a second DNA binding domain and a nucleobase transferase.
5. The DNA editing system of claim 4, further comprising: A third fusion protein comprises a third DNA binding domain and a transcriptional regulator.
6. A method for editing target DNA, comprising the steps of: The DNA editing system according to claim 4 is brought into contact with a target DNA, and the base at the target site of the target DNA is edited by the activity of the nucleic acid base transferase, wherein: The distance between the first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain and the second DNA binding domain recognition sequence or its complementary sequence recognized by the second DNA binding domain is 8 to 48 bases.
7. A method for preparing a cell in which target DNA has been edited, comprising the steps of: The DNA editing system according to claim 4 is introduced into a cell or expressed in a cell and brought into contact with a target DNA in the cell other than mitochondrial DNA, and the bases at the target site of the target DNA are edited by the activity of the nucleic acid base transferase, wherein: The distance between the first DNA binding domain recognition sequence or its complementary sequence recognized by the first DNA binding domain and the second DNA binding domain recognition sequence or its complementary sequence recognized by the second DNA binding domain is 8 to 48 bases.
8. The method according to claim 6 or 7, wherein The first DNA binding domain recognition sequence or its complementary sequence is present on the 5' side or 3' side of the target site via a first spacer of 4 to 16 bases, and The second DNA-binding domain recognition sequence or its complementary sequence is present on the opposite side of the target site to the first DNA-binding domain recognition sequence or its complementary sequence via a second spacer of 3 to 31 bases.
9. A kit for use in the method according to claim 6 or 7, comprising: At least one selected from the group consisting of a first fusion protein, a first fusion protein expression vector, and a polynucleotide encoding the first fusion protein, and at least one selected from the group consisting of a second fusion protein, a second fusion protein expression vector, and a polynucleotide encoding the second fusion protein, wherein: The first fusion protein expression vector is at least one selected from the group consisting of: (i) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, and a first DNA binding domain, and (ii) a vector comprising a polynucleotide encoding ND1, ND1n, and an ND1 linker and an insertion site for a polynucleotide encoding the first DNA binding domain, The second fusion protein expression vector is at least one selected from the following: (iii) a vector comprising a polynucleotide encoding a nucleic acid base transferase and a second DNA binding domain, and (iv) a vector comprising an insertion site for a polynucleotide encoding a nucleic acid base transferase and a polynucleotide encoding a second DNA binding domain.
10. The kit according to claim 9, further comprising at least one selected from the group consisting of: a third fusion protein comprising a third DNA binding domain and a transcriptional regulatory factor, a third fusion protein expression vector, and a polynucleotide encoding the third fusion protein, wherein: The third fusion protein expression vector is at least one selected from the following: (vii) a vector comprising a polynucleotide encoding a transcriptional regulatory factor and a third DNA binding domain, and (viii) a vector comprising an insertion site for a polynucleotide encoding a transcriptional regulatory factor and a polynucleotide encoding a third DNA binding domain. The fusion protein of claim 3 , further comprising a nucleobase transferase.
12. A method for editing target DNA, comprising the steps of: The fusion protein according to claim 11 is brought into contact with a target DNA, and the bases at the target site of the target DNA are edited by the activity of the nucleobase transferase.
13. A method for preparing a cell in which target DNA has been edited, comprising the steps of: The fusion protein according to claim 11 is introduced into a cell or expressed in a cell and brought into contact with a target DNA in the cell other than mitochondrial DNA, and the bases of the target site of the target DNA are edited by the activity of the nucleobase transferase.
14. A kit for use in the method according to claim 12 or claim 13, comprising at least one selected from the group consisting of the fusion protein, a fusion protein expression vector, and a polynucleotide encoding the fusion protein, wherein: The fusion protein expression vector is at least one selected from the following: (v) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, a nucleic acid base transferase, and a first DNA binding domain, and (vi) a vector comprising a polynucleotide encoding ND1, ND1n, an ND1 linker, a nucleic acid base transferase, and an insertion site for a polynucleotide encoding the first DNA binding domain.
Citation Information
Patent Citations
Polypeptide containing DNA-binding domain
JP2015033365A
Novel nuclease domain and uses thereof
WO2020045281A1
Separator for electrochemical element and electrochemical element using same
WO2020050377A1
Method for editing target DNA, method for producing cell having edited target DNA, and DNA edition system for use in said methods
WO2022050377A1