Targeted deaminonases and base editing using them
Non-toxic cytosine and adenine deaminases in fusion proteins enable targeted base editing in plant and animal cells, particularly mitochondria and chloroplasts, addressing the limitations of existing gene editing technologies and enhancing crop trait improvement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INST FOR BASIC SCI
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-10
AI Technical Summary
Existing gene editing technologies, such as CRISPR-Cas9, are not suitable for editing DNA sequences in plant organelles like mitochondria and chloroplasts due to difficulties in delivering guide RNA and expressing proteins, which are essential for studying gene function and improving crop traits.
Development of non-toxic, full-length cytosine and adenine deaminases in the form of fusion proteins with DNA-binding proteins, enabling targeted base editing in plant and animal cells, including mitochondrial and chloroplast genomes, by using split deaminases that only activate near the target DNA.
Achieves precise and efficient base editing in plant and animal cells, including organelles, without inducing DNA double-strand breaks, allowing for the study of gene function and improvement of crop traits.
Smart Images

Figure 2026062789000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to isolated forms of cytosine or adenine deaminoase or its variants, non-toxic full-length cytosine deaminoase or its variants, fusion proteins containing the same, base editing compositions, and methods for editing bases using the same. [Background technology]
[0002] A fusion protein, consisting of a DNA-binding protein and a deaminosease, enables targeted nucleotide substitution or base editing in a gene without generating DNA double-strand breaks (DSBs), editing point mutations that induce genetic defects, or single nucleotide transpositions in a targeted manner to introduce a desired single nucleotide mutation into prokaryotic cells, human cells, and other eukaryotic cells.
[0003] Unlike nucleases such as CRISPR-Cas9, which induce small insertions or deletions (indels) at target sites, this method alters a single base within a window of several nucleotides at the target site. Therefore, it can be used to edit point mutations that induce genetic disorders or to generate single nucleotide polymorphisms (SNPs) in cultured cells, animals, and plants.
[0004] Examples of fusion proteins in which DNA-binding proteins and deaminoses are linked include: 1) Base editors (BEs) containing catalytically-deficient Cas9 (dCas9) or D10A Cas9 nickase (nCas9) derived from S. pyogenes and rAPOBEC1, a rat cytosine deaminose; 2) Target-AID containing dCas9 or nCas9 and PmCDA1, an activation-induced cytidine deaminase (AID) ortholog from sea lamprey, or human AID; and 3) CRISPR-X, which contains dCas9 and sgRNAs linked to an MS2 RNA hairpin to recruit hyperactivated AID variants fused to an MS2-binding protein.
[0005] Thus, programmed gene editing tools such as ZFNs (zinc finger nucleases), TALENs (transcription activator-like effector nucleases), CRISPR (clustered regularly interspaced short palindromic repeat) systems, and base editing techniques for CRISPR-linked protein 9 (Cas9) mutants and base-deaminoenzyme proteins with low nucleic acid degradation efficiency have been developed in plant genetic research and crop trait improvement through alteration of base sequences. However, such tools are not suitable for editing the DNA sequences of plant organelles, including mitochondria and chloroplasts, mainly because it is difficult to transmit guide RNA to the organelles or to simultaneously express two compounds in the organelles. Plant organelles encode numerous essential genes necessary for photosynthesis and respiration. Methods and tools for editing the genes of such organelles are essential for studying the function of these genes and for improving crop productivity and traits. For example, targeted mutations in the mitochondrial atp6 gene can cause male infertility, a trait useful for reproduction, and specific point mutations in the 16S rRNA gene in the chloroplast genome can cause antibiotic resistance.
[0006] Bacterial toxin DddA tox This is the enzymatically active portion of a bacterial toxin derived from Burkholderia cenocepacia, which can deaminate cytosine in double helix DNA. An example of a deamination enzyme is DddA. tox Because it is toxic to cells, DddA is used to avoid toxicity in host cells. tox It is used by splitting it into two inactivated splits, but each half can be used as a functional DdCBE pair by linking it to a DNA-binding protein designed to bind to DNA.
[0007] In principle, this deamination enzymatic reaction is activated only when two inactive halves are positioned close to the target DNA by a DNA-binding protein. Therefore, cytosine-to-thymine base editing occurs midway between the binding sites of the two DNA-binding proteins. The two inactive halves, bound to a TALE (transcription activator-like effector) array using a DNA-binding protein, become functional when bound near the target DNA by the TALE. The cytosine-to-thymine base conversion is induced in a region of 14-18 bases between the two TALE binding sites. DddA is one such divided halves. tox This imposes many constraints on the experiment.
[0008] Overall length DddA tox Because of its toxicity, cloning using common E. coli is not possible; therefore, cloning is performed using E. coli that also express the immune gene that inhibits toxicity.
[0009] On the other hand, mitochondrial DNA plays a crucial role in cellular respiration, which occurs through the mitochondrial oxidative phosphorylation (OXPHOS) mechanism. Since the OXPHOS mechanism is essential for survival, mutations in mitochondrial DNA can cause serious dysfunction in various organs and muscles, such as tissues that require a lot of energy. In human mitochondrial diseases, normal mitochondrial DNA and mitochondrial DNA with single-nucleotide mutations coexist, resulting in a numerical heteroplasmy of mitochondrial DNA. The balance between mutations and normal mitochondrial DNA determines the development of clinically symptomatic mitochondrial diseases. In vitro and in vivo, programmable nucleohydrolases have been used in a way that cleaves mutated mitochondrial DNA while leaving normal mitochondrial DNA intact. However, such nucleohydrolases cannot introduce or reverse specific mutations in mitochondria because double-strand breaks in DNA are not efficiently repaired in mitochondria through non-homologous end joining or homologous recombination, as they are in the nucleus.
[0010] Mitochondrial base editing can be used to create models for a variety of diseases that were previously inaccessible, or to develop therapeutic agents to treat them. From this perspective, the need for developing highly efficient mitochondrial base editing enzymes is increasing.
[0011] Against this technical background, the present invention was completed by reducing non-selective base editing by substituting residues in deaminationases, or by making them non-toxic through novel full-length deaminationases, and confirming that they can be used as the target CBE (cytosine base editor) or ABE (adenine base editor) and can edit DNA. [Overview of the project] [Problems that the invention aims to solve]
[0012] An object of the present invention is to provide a DNA-binding protein, a cytosine or adenine deaminase or a variant thereof in an isolated form, a non-toxic full-length cytosine deaminase, or a fusion protein containing a variant thereof.
[0013] An object of the present invention is to provide a nucleic acid encoding the fusion protein.
[0014] An object of the present invention is to provide a composition for base editing containing the fusion protein or nucleic acid.
[0015] An object of the present invention is to provide a base editing method including a step of treating a cell with the composition.
[0016] To achieve the above object, the present invention provides a fusion protein comprising (i) a DNA-binding protein; and (ii) a first and a second split derived from a cytosine deaminase or a variant thereof, wherein the first and the second split are each in a form that binds to the DNA-binding protein.
[0017] The present invention provides a fusion protein comprising (i) a DNA-binding protein; and (ii) a non-toxic full-length cytosine deaminase derived from a cytosine deaminase or a variant thereof.
[0018] The present invention provides a fusion protein comprising (i) a DNA-binding protein; (ii) a cytosine deaminase or a variant thereof; and (iii) an adenine deaminase, wherein the cytosine deaminase or a variant thereof comprises (a) a non-toxic full-length cytosine deaminase or (b) a first and a second split derived from a cytosine deaminase or a variant thereof, and the first and the second split are each in a form that binds to the DNA-binding protein.
[0019] The present invention provides a nucleic acid encoding the fusion protein. The present invention provides a base editing composition comprising the aforementioned fusion protein or nucleic acid. The present invention provides a base editing composition for eukaryotic cells, comprising the aforementioned fusion protein or nucleic acid.
[0020] The present invention also provides a composition for base editing of plant cells, comprising the fusion protein or nucleic acid; and an NLS (nuclear localization signal) peptide or a nucleic acid encoding it.
[0021] The present invention also provides a composition for base editing of plant cells, comprising the fusion protein or nucleic acid; and a chloroplast transit peptide or a nucleic acid encoding it.
[0022] The present invention also provides a composition for base editing of plant cells, comprising the fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding it.
[0023] In some cases, the present invention also provides a composition for base editing of plant cells, further comprising a nuclear export signaling protein or a nucleic acid encoding the same.
[0024] The present invention also provides a method for base editing of plant cells, comprising the step of treating plant cells with the composition.
[0025] The present invention also provides a method for base editing of plant cells, comprising the steps of treating plant cells with the fusion protein or nucleic acid and an NLS (nuclear localization signal) peptide or nucleic acid encoding it.
[0026] The present invention also provides a method for base editing of plant cells, comprising the steps of: treating plant cells with the fusion protein or nucleic acid; and a chloroplast transit peptide or nucleic acid encoding it.
[0027] The present invention also provides a method for base editing of plant cells, comprising the steps of: treating plant cells with the fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or nucleic acid encoding it.
[0028] The present invention also provides a composition for base editing of animal cells, comprising the fusion protein or nucleic acid; and an NLS (nuclear localization signal) peptide or a nucleic acid encoding it.
[0029] The present invention also provides a composition for base editing of animal cells, comprising the fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or nucleic acid encoding it.
[0030] The present invention also provides a composition for base editing of animal cells, which optionally further comprises a nuclear export signaling protein or a nucleic acid encoding it.
[0031] The present invention also provides a method for base editing of animal cells, comprising the step of treating animal cells with the composition.
[0032] The present invention also provides a method for base editing of animal cells, comprising the steps of treating animal cells with the fusion protein or nucleic acid and an NLS (nuclear localization signal) peptide or a nucleic acid encoding it.
[0033] The present invention also provides a method for base editing of animal cells, comprising the steps of treating animal cells with the fusion protein or nucleic acid and a mitochondrial targeting signal (MTS) or nucleic acid encoding it.
[0034] The present invention also provides a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and the cytosine deaminose of the fusion protein or a variant thereof is derived from bacteria and is specific to double-stranded DNA.
[0035] The present invention also provides a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, the cytosine deaminose of the fusion protein or a variant thereof is derived from bacteria and is specific to double-stranded DNA, a DNA-binding protein is bound to the N-terminus of the cytosine deaminose or its variant, and a DNA-binding protein is bound to the C-terminus of the adenine deaminose of the fusion protein.
[0036] The present invention also provides a composition for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, the cytosine deaminoase of the fusion protein or a variant thereof is a full-length non-toxic cytosine deaminoase, the cytosine deaminoase of the fusion protein or a variant thereof is derived from bacteria and is specific to double-stranded DNA.
[0037] The present invention also provides a composition for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, the cytosine deaminoase of the fusion protein or its variant is a split cytosine deaminoase comprising a first split and a second split, and the cytosine deaminoase of the fusion protein or its variant is derived from bacteria and is specific to double-stranded DNA.
[0038] The present invention also provides a method for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and the cytosine deaminosease or a variant of the fusion protein is of bacterial origin and is specific to double-stranded DNA.
[0039] The present invention also relates to a method for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease.
[0040] The cytosine deaminoase or its variant of the aforementioned fusion protein is derived from bacteria and is specific to double-stranded DNA.
[0041] The present invention provides a method in which a DNA-binding protein is bound to the N-terminus of the cytosine deaminose or its variant, and a DNA-binding protein is bound to the C-terminus of the adenine deaminose of the fusion protein.
[0042] The present invention also provides a method for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and the cytosine deaminose of the fusion protein or a variant thereof is of bacterial origin and is specific to double-stranded DNA. [Brief explanation of the drawing]
[0043] [Figure 1]This figure shows the results of ZFD optimization using the pTarget plasmid. a. ZFD structure. Half of the Split-DddAtox fuses to the C-terminus of the ZFP (type C). b. Optimization of the ZFD platform using the pTarget library. The pTarget plasmid contains spacers ranging in size from 1 to 24 bp (shown in red) and ZFP DNA binding sites (shown in green). The ZFD structure includes AA linkers of varying lengths (shown in yellow and orange) and other DddAtox splitting sites and directions (shown in blue). cd. ZFD activity was measured at target sites in the pTarget library to investigate the effects of the variables described in (b). ZFD pairs with the same (c) or different (d) lengths of linkers were tested in the ZFDs on the left and right. Base edit frequencies were measured by targeted deep sequencing of the relevant regions of the pTarget plasmid. Data are shown as mean ± standard error of mean (sem) obtained from n=2 biologically independent samples. [Figure 2] This figure shows the results of examining the ZFD efficiency with diverse linkers in pTarget plasmids. a. Frequency of C / G to non-C / G editing in diverse pTarget spacers with lengths of 1-24 bps, as depicted in the heatmap. Diverse ZFD configurations were tested, including different types of linkers between the ZFP and the split DddAtox, and the sites where the DddAtox was split. b. Overall activity of each ZFD pair. In the nomenclature used on the x-axis, the ZFD on the left is shown below, and the ZFD on the right is shown above. c. Base editing efficiency by spacer length. "AA" indicates the number of amino acids in the linker. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. [Figure 3]This figure shows the results of verifying the efficiency of ZFDs with 24AA linkers and other linkers. The heatmap shows the effect of ZFD linker length on the editing efficiency from C / G to non-C / G. In the ZFD pairs, the left ZFD is fixed with 24AA linkers, and the right ZFD contains a variable-length linker, and vice versa. Error bars are the standard error of the mean (sem) of n=2 biologically independent samples. [Figure 4] This figure shows the results of confirming the activity of ZFDs targeting the nucleus in vivo. a. Structure of nuclear DNA-targeted ZFDs. Half of Split-DddAtox fuses to the C-terminus (C-type) or N-terminus (N-type) of the ZFP. ZFD pairs are designed in CC or NC configurations, consisting of a ZFD on the left side of the C-type and a ZFD on the right side of the C-type, or a ZFD on the left side of the N-type and a ZFD on the right side of the C-type, respectively. b. Frequency of ZFD-induced base editing at nuclear DNA target sites in HEK 293T cells. Data are shown mean ± sem for n=3 biologically independent samples. cf. ZFD-induced base editing efficiency at each base position within the spacer of NUMBL(c), INPP5D-2(d), TRAC-CC(e), and TRAC-NC(f) target sites in HEK 293T cells. Data are shown mean ± sem for n=3 biologically independent samples. g) ZFD-induced base edit frequency in K562 cells after delivery of ZFD protein or ZFD-encoding plasmid by electroporation or direct delivery. ZFD proteins with one or four NLS were tested, and left and right ZFDs were used equimorally. For electroporation, an Amaxa 4D-Nucleofector was used. For direct delivery, K562 cells were incubated with cell medium containing left and right ZFD proteins. Cells were treated once (1x) or twice (2x) in the same manner. For biologically independent n=2 samples, data are shown mean ± sem. [Figure 5] This is a schematic diagram showing the structure of a ZFD targeting nuclear DNA. There are four possible ZFD structures. While the NC and CN structures are structurally identical, the ZFD structures on the left and right differ in type. [Figure 6] This figure shows the results of examining the deletion rate of ZFDs targeting the in vivo nucleus. All ZFDs tested generated deletions at a frequency of less than 0.4%. Data are shown mean ± sem for n=3 biologically independent samples. [Figure 7] This figure shows the results of an in vitro activity test on recombinant ZFD proteins. a. Purification steps of ZFD pairs targeting TRAC sites. GST-tagged proteins were purified in E. coli cell lysates using glutathione Sepharose beads. The purification steps were monitored using polyacrylamide gel electrophoresis. The gel was stained with Coomassie blue. Lane 1, molecular weight marker. Lane 2, sample of cells in which protein expression was not induced by IPTG. Lane 3, sample of cells in which protein expression was induced by IPTG. Lane 4, soluble fraction after sonication. Lane 5, insoluble fraction after sonication. Lane 6, column pass-through fraction. Lane 6, wash fraction. Lane 7, elution fraction. The size of representative markers is shown on the left. Red boxes indicate ZFD proteins. b. ZFD binding sites on the left and right. Red arrows indicate sites where ZFD-induced deamination is possible. c. Summary of ZFD activity against PCR amplicons containing TRAC sites. The TRAC-NC ZFD pair first deaminates cytosine to produce uracil (shown in red). Then, the USER enzyme cleaves the uracil to create a gap (shown in red triangles). d. Untreated PCR amplicons (left) and PCR amplicons treated with the ZFD pair (right) analyzed by agarose gel electrophoresis. [Figure 8] This is a schematic diagram of various sequences (configurations) of mitochondrial DNA-targeted ZFDs. There are four possible mitoZFD configurations. The NLS of conventional ZFDs has been replaced by MTS and NES. Although the NC and CN configurations are structurally identical, the configuration types of the ZFDs on the left and right sides are different. [Figure 9]This figure shows the results of confirming the mitochondrial gene base editing efficiency of mitoZFD. a. Base editing frequency of mtDNA induced by mitoZFD and TALE-DdCBE in HEK 293T cells. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. bg. mitoZFD-induced base editing efficiency at each base position within the spacer of the ND2(b), ND4L(c), COX2(d), ND6(e), and ND1(f) target sites in HEK 293T cells, and TALE-DdCBE-induced base editing efficiency at the ND1(g) target site. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. Comparison of DNA and amino acid changes in the ND1 gene introduced by mitoZFD and TALE-DdCBE. The frequency (%) of sequencing reads for each mutant allele was measured by targeted deep sequencing. The spacer regions for the ZFD pair and the TALE-DdCBE pair are shown by the blue dotted lines. [Figure 10] This figure shows the results of confirming the base editing efficiency of single-cell-derived clone populations isolated from MT-ZFD-treated HEK293T cells. Single-cell-derived clones were obtained for allelic analysis. In each single-cell-derived clone, the C / G-to-non-C / G editing frequency was determined by targeted deep sequencing. a) Single-cell-derived clone from an HEK 293T cell population treated with ND1-target mitoZFD. b) Single-cell-derived clone from an HEK 293T cell population treated with ND2-target mitoZFD. c) Single-cell-derived clone from an untreated HEK 293T cell population. ZFP binding sites are shown in green. High editing frequencies in clones that underwent mitoZFD-induced editing are shown in red. [Figure 11]This figure shows the results of confirming the base editing efficiency of a single-cell-derived clone population isolated from HEK293T cells treated with MT-ZFD. Analysis of alleles from single-cell-derived clones exhibiting high-frequency base editing. The table shows the amino acids altered as a result of base editing at ND1. In the reference sequence at the top, red text indicates spacers. In the alleles, red text indicates changes in the amino acid sequence. (* indicates a stop codon.) [Figure 12a] and [Figure 12b] This figure shows the results of confirming the base editing efficiency of a single-cell-derived clone population isolated from HEK293T cells treated with MT-ZFD. Analysis of alleles from single-cell-derived clones showing high-frequency base editing. The table shows the amino acids altered as a result of base editing at ND2. In the reference sequence at the top, red text indicates spacers. In the alleles, red text indicates changes in the amino acid sequence. [Figure 13] This figure shows the results of confirming the base editing efficiency of the ZFD and TALE-DdCBE combination. a. DNA sequences of the binding regions of the mitoZFD and TALE-DdCBE pairs. Sites recognized by TALE-DdCBE are highlighted in green, and those recognized by mitoZFD are highlighted in blue. The upper sequence shows the mtDNA heavy strand, and the lower sequence shows the mtDNA light strand. b. Frequency of cytosine edited by ZFD, TALE-DdCBE, and ZFD / DdCBE mixed pairs. Data were obtained using targeted deep sequencing, and the data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. c. Heatmap of base editing activity at each base position. Red boxes indicate the spacer region of each configuration. Blue arrows indicate the mtDNA position. [Figure 14]This figure shows the results of examining editing efficiency by mRNA vs. plasmid at different ZFD concentrations. The target specificity of ND1-targeted mitoZFD across the entire mitochondrial genome varies depending on the concentration of the mRNA or plasmid encoding the ZFD. The on- and off-target base edit frequencies are determined by the sequencing of the entire mtDNA. The graph shows the results for HEK 293T cells transfected with the indicated concentrations of ND1-targeted mitoZFD-encoding plasmid or mRNA. Red arrows indicate target sites, red dots indicate base edit frequencies at target sites, and gray dots indicate SNPs that are also present in the control. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. [Figure 15] The results of examining editing efficiency by mRNA vs. plasmid at different ZFD concentrations are shown. a. The top figure shows ZFD binding at the ND1 site. ZFD binding sites are shown in green. Target cytosines between spacers are shown in red. On-target activity as defined by the overall mtDNA sequencing data in Figure 14. Activity decreases as the amount of plasmid or mRNA encoding the transfected mitoZFD decreases. b. Number of C / G sites edited at a frequency of >1% for each plasmid or mRNA amount. c. Mean C / G-to-T / A editing frequency of all C / Gs in the mitochondrial genome at each plasmid or mRNA concentration. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. [Figure 16]This report shows the results of editing efficiency by mRNA vs. plasmid at different ZFD concentrations. On and off-target base edit frequencies are determined by the sequencing of the entire mtDNA. The graph shows the results for HEK 293T cells transfected with the indicated concentrations of ND2-target mitoZFD-encoding plasmid or mRNA. Red arrows indicate target sites, red dots indicate base edit frequencies at target sites, and gray dots indicate SNPs present in both the target and control cells. Data are presented as mean ± standard error of the mean (sem) obtained from two biologically independent samples. [Figure 17] This figure shows the results of examining editing efficiency by mRNA vs. plasmid at different ZFD concentrations. a. The top figure shows ZFD binding at the ND2 site. ZFD binding sites are shown in green. Target cytosines between spacers are shown in red. On-target activity as defined by the overall mtDNA sequencing data in Figure 16. Activity decreases as the amount of plasmid or mRNA encoding the transfected mitoZFD decreases. b. Number of C / G sites edited at a frequency of >1% for each plasmid or mRNA amount. c. Mean C / G-to-T / A editing frequency for all C / Gs in the mitochondrial genome at each plasmid or mRNA concentration. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. [Figure 18]This figure shows the results of creating and verifying the editing efficiency of mitochondrial whole-grain sequencing / QQ mutants. a. The QQ mitoZFD mutant eliminates nonspecific DNA contacts by including an R(-5)Q mutation in each zinc finger of the ZFD. (If there is no R at position -5 of the zinc finger framework, a nearby K or R is converted to Q.) b. Sequence of the entire mtDNA of mitoZFD-treated cells. On-site and off-site editing frequencies are shown as red and black dots, respectively. Data are shown as mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. All C / G to T / A base edits with >1% efficiency are shown. cd. Editing efficiency and specificity vary depending on the dose of ZFD-encoding mRNA transmitted. c. Mean C / G-to-T / A editing frequency for all C / Gs in the mitochondrial genome. d. Number of edited C / Gs with a base editing frequency greater than 1%. [Figure 19] Golden Gate assembly system for plant-based base editors. Schematic diagram of Golden Gate assembly for cp-DdCBE and mt-DdCBE structures. For each position of the target sequence, a total of 424 sets (= 6 x 64 x 3 parts + 2 x 16 x 2 parts + 2 x 4 x 1 part) were selected from TALE subarray plasmids and mixed with the desired vector to generate plasmids encoding DdCBE targeting specific sequences. [Figure 20]Chloroplast and mitochondrial base editing in plants. a, b, c, d, Frequency and pattern of chloroplast base editing induced by cp-DdCBE in 16s rDNA (a, b) and psbA (c, d). Split DdCBE pairs G1333 and G1397 were transfected into lettuce and rapeseed protoplasts. e, Efficiency and pattern of mitochondrial base editing induced by mt-DdCBE in the fATP6 gene. Split DdCBE pairs G1333 and G1397 were transfected into lettuce and rapeseed protoplasts. a, c, e, TALE binding regions are shown in blue, and spacer cytosine is shown in orange. In all graphs, error bars indicate the mean ± standard deviation. Three independent biological copies. b, d, f, Translated nucleotides are shown in red. Edited allele % (mean ± standard deviation) was obtained from three independent experiments. [Figure 21] DNA editing of plant organs by DdCBE. a. Schematic diagram of plant organ mutagenesis. b. Efficiency of C·G to T·A conversion in cp-DdCBE transfected calluses cultured in the absence of spectinomycin, including a representative Sanger sequencing chromatogram. Converted nucleotides are shown in red to the left. Arrows indicate substituted nucleotides in the chromatogram. c. Summary of DdCBE-driven plant organ mutagenesis. Mutant calluses are shown to have a much higher editing frequency than simulated callus frequencies. d. Frequency of C to T conversion induced after transfection of lettuce protoplasts with mRNA encoding cp-DdCBE targeting 16srDNA. Error bars are mean ± sd. n=3 independent biological replicas. e. Editing frequency and pattern of spectomycin-resistant calluses at 2.5 months. f. Efficiency of C·G to T·A conversion in streptomycin-resistant plants transfected with DdCBE mRNA using a representative Sanger sequencing chromatogram. The arrows indicate substituted nucleotides in the chromatogram. Scale bar: 1 mm. [Figure 22]Chloroplast and mitochondrial base editing strategies. Both cp-DdCBE and mt-DdCBE preproteins contain either CTP (chloroplast transit peptide) or MTS (Mitochondrial targeting signal), and are therefore transported to chloroplasts and mitochondria after translation in plant cells. The preproteins pass through the outer and inner membranes of organelles, where CTP and MTS are cleaved by interstitial and mitochondrial-processed peptidases, respectively, before the cp-DdCBE and mt-DdCBE (mature proteins) form their final forms (conformation). [Figure 23] Time course of editing via DdCBE plasmid in lettuce protoplasts. Transfected protoplasts were collected at each time point, and editing efficiency was analyzed by targeted deep sequencing. Frequencies (mean ± standard deviation) were obtained from three independent experiments. [Figure 24] Base editing frequency of the psbB gene. After transfection of rapeseed protoplasts with plasmids encoding cp-DdCBE vs. Left-G1333-N+Right-G1333-C, which targets the chloroplast psbB gene, the base editing efficiency of the spacer region was analyzed by targeted deep sequencing. The TALE binding region, target cytosine, and converted nucleotides are shown in blue, orange, and red, respectively. Frequencies (mean ± standard deviation) were calculated for n=3 independent experiments. [Figure 25] Base editing efficiency of the mitochondrial RPS14 gene. After transfection of rapeseed protoplasts with a plasmid encoding mt-DdCBE vs. Left-G1333-N + Right-G1333-C, which targets the RPS14 gene, the conversion efficiency from C to T was analyzed by targeted deep sequencing. The TALE binding region, target cytosine, and converted nucleotides are shown in blue, orange, and red, respectively. Frequencies (mean ± standard deviation) were calculated for n=3 independent experiments. [Figure 26] Base editing efficiency targeting chloroplast genomes in callus. Frequency and pattern of base editing via DdCBE at target sites in 16srDNA and psbA in lettuce and rapeseed callus after 4 weeks of culture. Converted nucleotides in the spacer region are shown in red. [Figure 27] Base editing efficiency targeting the mitochondrial genome in callus. The frequency and pattern of DdCBE-mediated base editing at target sites of the ATP6 and RPS14 genes in rapeseed callus were confirmed by targeted deep sequencing. Converted nucleotides in the target spacer region are shown in red. [Figure 28] DNA-free base editing. Frequency and pattern of chloroplast base editing at target sites in 16srDNA after transfection of lettuce protoplasts with DdCBE mRNA. Target dip sequencing was performed after culturing protoplasts for 7 days. Converted nucleotides in the target spacer region are shown in red. [Figure 29] This figure shows the gel electrophoresis results indicating the absence of DdCBE mRNA or DNA sequences in the protoplast and callus (M is a marker). [Figure 30] Selection of 16srDNA mutations. Red arrows indicate streptomycin-resistant green callus. [Figure 31] No off-target mutations were found near the DdCBE target site in antibiotic-resistant callus or seedlings. (a), (b) Off-target activity was analyzed by targeted dip sequencing. The TALE binding site and spacer region are indicated by green and red underlines, respectively. (a) Spectinomycin-resistant callus cultured in lettuce protoplasts transfected with the DdCBE plasmid. (b) Shoots obtained from streptomycin-resistant seedlings. [Figure 32]Off-target activity analysis at the five sites with the highest homology to the on-target site. Five potential off-target sites of the 16S rRNA gene-specific DdCBE in the lettuce chloroplast genome were selected, including up to nine mismatches at the TALE binding site. TALE binding sequences and mismatched nucleotides are shown in blue and red, respectively. Off-target mutation frequencies were measured using targeted dip sequencing in protoplasts and drug-resistant callus or suits transfected with DdCBE plasmid or DdCBE mRNA. Frequencies (mean ± standard deviation) were obtained from three independent experiments. [Figure 33] Schematic diagrams of DdCBE assembly and mitochondrial DNA editing. a. Description of one-pot Golden-gate assembly for efficient DdCBE construction. A total of 424 sequences (64 triplicate recognition sequences × 6 + 16 duplicate recognition sequences × 2 + 4 single-recognition sequences × 2) and expression vectors were mixed to create the left and right modules for final plasmid construction. b. Schematic diagram of the interaction between the target gene ND5 in mouse mitochondrial DNA and DdCBE. TALE binding sites are shown in gray, and base editing ranges are shown in black. Each repeat variable diresidue module is shown in orange, blue, green, and yellow, which are adenine "NI", thymine "NG", guanine "NN", and cytosine "HD" for recognition, respectively. [Figure 34]Mouse mitochondrial ND5-point mutations resulting from DdCBE-mediated base editing. a. Target sequence and efficiency in NIH3T3 cells for cytosine-to-thymine base editing mediated by DdCBE deaminationase. In the target sequence, translation codons are underlined and editable sites are shown in red. Combinations for DdCBE transfection are denoted as left or right, -G1333 or -G1397, and -N or -C. The p-values for the C10 mutations Left-G1333-N+Right-G1333-C, left-G1333-C+right-G1333-N, left-G1397-N+right-G1397-C, and left-G1397-C+right-G1397-N were 0.0012, 0.0003, 0.0014, and 0.0009, respectively, and for the C13 mutations, they were 0.0116, 0.0076, 0.0030, and 0.0003, respectively (*p<0.05 and **p<0.01, using Student's two-tailed t test). b. Base editing efficiency in mouse blastocysts. Sequence analysis data were obtained from blastocysts developed from conjugates microinjected with left-G1397-N and right-G1397-C-DdCBE mRNA. c. Alignment diagram of neonatal mutant sequences. Targeted dip sequencing was performed on genetic DNA extracted from tissues obtained from the tail immediately after birth and from the toes at 7 and 14 days postnatally. Edited bases are shown in red. Edit frequency of mutant mitochondrial genomes is shown. d. Edit efficiency in various tissues of adult F0 mice (sipup-1). Sequencing data were obtained from each tissue at 50 days postnatally. In all graphs, dark gray bars and light gray bars represent the editing frequencies of m.C12539T (C10) and m.G12542A (C13) mutations, respectively. Error bars are the standard error of the mean (sem) for n=3 biologically independent samples. [Figure 35]Transmission of mutant mitochondrial DNA to germ cells. a. To observe the gonadal transmission of mtDNA mutations, female F0 (sipup-3) mice were crossed with wild-type C57BL6 / J males to obtain F1 offspring (101, 102), after which targeted deep sequencing was performed. Edited bases are shown in red. Shows the editing frequency of the mutant mitochondrial genome. b. Base editing efficiency in various tissues of F1 offspring (101) obtained using targeted deep sequencing of genomic DNA. Dark gray and light gray bars show the frequencies of m.C12539T (C10) and m.G12542A (C13) mutations, respectively. Error bars are the standard error of the mean (sem) of n=3 biologically independent samples. [Figure 36]Mouse mitochondrial ND5 G12918A mutations generated by DdCBE. a. DdCBE target for creating the m.G12918A point mutation that causes a D393N change in the ND5 protein. The target codon is underlined, and possible editing sites are shown in red. b. Efficiency of cytosine-thymine base editing using DdCBE in NIH3T3 cells. Combinations of transfected DdCBE pairs are shown. Error bars are sem. For n=3 biologically independent samples (ns are not meaningful, *p<0.05, **p<0.01, using Student's two-tailed t test). The P-values for the C6 mutations left -G1333-N + right -G1333-C, left -G1333-C + right -G1333-N, left -G1397-N + right -G1397-C, and left -G1397-C + right -G1397-N are 0.0052, 0.0099, 0.0027, and 0.0040, respectively. The P-value for ns is 0.4971. c, m. Base editing efficiency of point mutations in G12918A mouse blastocysts. Sequencing data were obtained from blastocysts developed after microinjection of mRNA encoding left -G1397-C and right-G1397-N-DdCBE into one-cell stage embryos. d, Mice with ND5 point mutation (F0). F0 offspring carrying the ND5 point mutation developed after microinjection of DdCBE mRNA. Alignment of mutant sequences identified in newborns. Edited bases are shown in red, and the right side shows the editing frequency of the mutant mitochondrial gene. [Figure 37]Mouse mitochondrial ND5 nonsense mutations produced by cytosine deaminationase-mediated base editing. a. DdCBE target sequences for generating m.C12336T nonsense mutation and m.G12341A silent mutation. The m.C12336T(C9) mutation produces a Q199 stop mutation in the ND5 protein, while m.G12341A(C14) causes a silent Q200Q mutation. Transcriptional triplets are underlined, and possible editing sites are shown in red. b. Efficiency of cytosine-thymine base editing for generating nonsense mutations in NIH3T3 cells. Shows transfected DdCBE pair combinations. Dark gray and light gray bars indicate the frequencies of m.C12336T(C9) and m.G12341A(C14) mutations, respectively. Error bars indicate sem. For n=3 biologically independent samples (ns not meaningful, *p<0.05, **p<0.01, using Student's two-tailed t test), the p-values for C9 mutations Left-G1333-N+Right-G1333-C, left-G1333-C+right-G1333-N, left-G1397-N+right-G1397-C, left-G1397-C+right-G1397-N were 0.0065, 0.1143, 0.0266, and 0.0037, respectively, and the p-values for C14 mutations were 0.0077, 0.0144, 0.0406, and 0.0214, respectively. c, editing efficiency in mouse blastocysts. Sequence analysis data were obtained from blastocysts developed after microinjection of mRNA encoding left-G1333-N and right-G1333-C-DdCBE into conjugates. Dark gray and light gray bars indicate the frequencies of C9 and C14 mutations, respectively. d. Neonatal mutant sequence alignment. Edited bases are shown in red, and the right side shows the editing frequency of the mutant mitochondrial genome. e. Sanger sequence analysis chromatograms of wild-type and edited mice. Red arrows indicate substituted nucleotides. [Figure 38]This is a schematic diagram of the Golden Gate cloning process for generating the DdCBE structure. All reactions occur simultaneously in a single tube. Arrows do not indicate a continuous reaction process. Using the BsaI enzyme, empty expression vectors and module vectors were cleaved to remove the linearized backbone and TALE module inserts, which contained compatible aggregated ends. Next, T4 DNA ligase was used to combine the backbone and the six module inserts to construct the final DdCBE structure. Eight DdCBE replicating backbone plasmids were used. For SOD2 MTS, these were Left-G1333-N, Left-G1333-C, Left-G1397-N, Left-G1397-C. For COX8A MTS, these were Right-G1333-N, Right-G1333-C, Right-G1397-N, Right-G1397-C. [Figure 39] ND5 mutant mice (F0). (a) ND5 silent mutant mouse, (b) ND5 G12918A mutant mouse, (c) ND5 nonsense mutant mouse generated after microinjection of DdCBEmRNA. [Figure 40] (a) Schematic diagram of the vector containing DdCBE-NES and the NES sequence. (b) Mouse m.G12918 ND5 gene and ND5-like sequence on chromosome 4 in the nucleus. Mitochondrial TrnA and chromosome 5 sequence in the nucleus. Mitochondrial Rnr2 and chromosome 6 sequence in the nucleus. (c) Editing efficiency of DdCBE and DdCBE-NES in the ND5 gene using NIH3T3 cell line. (d) Editing efficiency of DdCBE and DdCBE-NES in the TrnA gene using NIH3T3 cell line. (e) Editing efficiency of DdCBE and DdCBE-NES in the Rnr2 gene using NIH3T3 cell line. The orange graph shows the editing efficiency of DdCBE, and the gray graph shows the editing efficiency of DdCBE-NES. (f) DNA recognition sequence of the mitoTALEN TALE sequence. (g) DdCBE base editing efficiency in experimental groups treated with and without mitoTALEN. All graphs are for n=2, and the error bars represent the mean standard error. [Figure 41]Improved editing efficiency in mouse embryos and individuals using DdCBE-NES and mitoTALEN. (a) Base editing efficiency of multiple mitochondrial DNA targets (mtND5, mtTrnA, mtRNR2) in blastocysts using DdCBE-NES. (b) Comparison of m.G12918A base editing efficiency using DdCBE, DdCBE-NES, and mitoTALEN. (c) Comparison of m.G12918A base editing efficiency in individual mice. All graphs are n>=3, and error bars are the mean standard error. (ns = statistically insignificant, *p<0.05, **p<0.01, ***p<0.001 obtained using Student's two-tailed t test.) [Figure 42] (a) Schematic diagram of DdCBE protein modification. (b, c) Crystal structure of DddAtox deaminationase. Residues at the binding surface are shown as bars. (b) G1397-N and G1397-C splits are shown in purple and light blue, respectively. (b) G1333-N and G1333-C splits are shown in orange and green, respectively. (d, e) G1397-N and G1397-C (d) and G1333-N and G1333-C (e) during binding surface mutations are shown in red. [Figure 43] (a) Graph of base editing efficiency of G1397-binding surface mutations. Editing range and target cytosine are shown at the top. Mutants and wild-type / TALE-free DddAtox proteins were cotransfected as shown on the side. For Left-DdCBE, the TALE-free DddAtox protein is G1397-N, and for Right-DdCBE, the TALE-free DddAtox protein is G1397-C. (b) Heatmap showing target cytosine-thymine (guanine-adenine) base editing efficiency for DdCBE and mutants. [Figure 44](a) Graph of base editing efficiency of G1333 binding surface mutations. Editing range and target cytosine are shown at the top. Mutant and wild-type / TALE-free DddAtox proteins were cotransfected as shown on the side. For Left-DdCBE, the TALE-free DddAtox protein is G1333-N, and for Right-DdCBE, the TALE-free DddAtox protein is G1333-C. (b) Heatmap showing target cytosine-thymine (guanine-adenine) base editing efficiency of DdCBE and mutant. [Figure 45] This figure shows the results of a comparison of the amino acid sequences of the wild type and the novel full-length DddA. [Figure 46] This figure shows the morphology of the full-length DddA molecule transmitted to animal or plant cells. [Figure 47] This figure shows the results of confirming the activity of TC Motif in substituting cytosine with thymine at the human cell genome context ROR1 site (a), HEK3 site (b), and TYRO3 site (c). [Figure 48] This diagram illustrates the advantages of the overall length DddA. [Figure 49] This figure shows the results of measuring the activity of full-length DddA in human cell genome contexts: TRAC site 1 (a), TRAC site 2 (b), FANCF (c), and HBB (d). [Figure 50] This figure shows the results of measuring DddA activity in the human cell genome contexts TYRO3 (a), ROR1 (b), HEK3 (c), EMX1 site 2 (d), TRAC site 1 (e), and HBB (f) using DddA-dCas9(D10A, H840A)-UGI. [Figure 51]Base editing efficiency of full-length DddAtox in HEK293T cells. (a) Schematic diagram of a structure-based screening method for full-length DddAtox. Red alanine indicates that a positively charged amino acid residue has been replaced with alanine. (b) E. coli transformants of DddA mutants with alanine substitution are shown, with E1347A used as a control group and an active site mutant. (c) Frequency of DddA AAAAA and CBE editing and insertion / deletion at the TYRO3 site. (d) Allele frequency at the TYRO3 site: C to T substitutions are shown in red. Protospacers are shown in blue, and protospacer adjacent motifs (PAMs) are shown in orange. [Figure 52] Non-toxic DddA GSVG. (a) Schematic diagram of error-prone PCR-based non-toxic full-length DddAtox variant screening. Edit frequencies of (b) and alleles fused to the N-terminus and C-terminus of Cas9, nCas9(D10A), nCas9(H840A), and dCas9(D10A, H840A). Protospacers are shown in blue, and protospacer-adjacent motifs (PAMs) are shown in orange. [Figure 53] Frequency of DddAtox mutants, in which positively charged amino acid residues at the TYRO3 site (a), ROR1 site 1 (b), and HEK3 site (c) are substituted with alanine, being edited into the N-terminus of nCas9 (D10A). Protospacers are shown in blue, and protospacer-adjacent motifs (PAMs) are shown in orange. [Figure 54] ROR1 site 1(a), ROR1 site 2(b), ROR1 site 3(c), FANCF site(d), HBB site(e), HEK3 site(f), TRAC5 site 1(g), and EMX1 sites(h, I, j) are shown respectively. Protospacers are shown in blue, and protospacer adjacent motifs (PAMs) are shown in orange. C to T substitutions are shown in red. The target window for DddA is counted at the protospacer 5' upstream and is shown as a negative number. [Figure 55]Editing frequencies of AAAAA and E1347A in HeLa cells at TYRO3 site (a), ROR1 site 1 (b), ROR1 site 2 (c), ROR1 site 3 (d), FANCF site (e), HB site (f), HEK3 site (g), TRAC5 site 1 (h), TRAC5 site 2 (i), and EMX1 site 2 (j). Protospacers are shown in blue, and protospacer adjacent motifs (PAMs) are shown in orange. The target window for DddA is counted at the protospacer 5' upstream and is shown as a negative number. Target cytosine is shown in red. [Figure 56] This figure shows the time-dependent base editing and insertion / deletion ratios for AAAAA and E1347A at the TYRO3 site (a) and ROR1 site 1 (b). [Figure 57] Frequencies of edits, insertions, deletions, and alleles of GSVG fused to the N-terminus of nCas9(D10A), nCas9(H840A), and dCas9 at EMX1 site 2(a), FANCF site(b), TRAC5 site 1(c), TRAC5 site 2(d), ROR1 site 1(e), ROR1 site 2(f), ROR1 site 3(g), and HBB site(h). Protospacers are shown in blue, and protospacer adjacent motifs (PAMs) are shown in orange. C-to-T substitutions are shown in red. The GSVG target window is counted at the 5' upstream of the protospacer and is shown as a negative number. [Figure 58] Frequencies of edits, insertions, deletions, and alleles of GSVG fused to the C-terminus of nCas9(D10A), nCas9(H840A), and dCas9 at EMX1 site 2(a), EMX1 site 4(b), ROR1 site 2(c), and HBB site(d). Protospacers are shown in blue, and protospacer adjacent motifs (PAMs) are shown in orange. G-to-A substitutions are shown in red. The GSVG target window is shown counted 3' downstream from the initial position of the protospacer. [Figure 59]Frequency of time-dependent edits and insertions / deletions of E1347A, GSVG, SSVG, GSAG, and GSVS fused to the C-terminus of nCas9(H840A) in TYRO3(a) and EMX1 site 2(b). [Figure 60] Mitochondrial base editing of mDdCBE in HEK293T cells. Editing efficiency of ND4(a) and ND6(b). Target cytosine and TALE binding sites are shown in red and gray, respectively. Editing efficiency of ND4(c, d) and ND6(e, f) when only half of DddAtox fused to the TALE array and the other half did not have TALE. The TALE sequences on the left and right are indicated by L and R, respectively. ND6 TALE array mismatches with the reference genome are indicated by purple underlines. [Figure 61] (a) Schematic diagram of zinc finger cytosine (ZFD) using conventional ZFP DNA-binding protein. (b) Insertion site of adenine deaminase into ZFD (red arrow indicates insertion site). (c) Base editing efficiency (C-to-T) of the constructed ZFDdABE at the nuclear DNA Trac site. (d) Base editing efficiency (A-to-G) of the constructed ZFDdABE at the nuclear DNA Trac site. (WT-ZFD is a C-to-T deaminase that has only fragmented DddAtox without adenine deaminase.) (e) Efficiency (C-to-T) of ZF-DdABE targeting mitochondrial DNA at the ND1 site. (f) Efficiency (A-to-G) of ZF-DdABE targeting mitochondrial DNA at the ND1 site. [Figure 62](a) Schematic diagram of DdABE using TALE and split DddAtox. (Components include split DddAtox, adenine deaminase, and TALE array) (b) Base editing efficiency when only adenine deaminase is attached to TALE targeting the mitochondrial ND4 site (c) Base editing efficiency when adenine deaminase is attached to TALE-split DddAtox targeting the mitochondrial ND1 site (d) Base editing efficiency viewed at the single nucleotide level when adenine deaminase is attached to TALE-split DddAtox on the left and TALE-split DddAtox on the right (green boxes indicate the sites where TALE is attached) (e) Base editing efficiency viewed at the single nucleotide level when adenine deaminase is attached to TALE-split DddAtox on the left and DdCBE on the right (green boxes indicate the sites where TALE is attached) [Figure 63] (a) C-to-T and A-to-G base editing efficiency of DdABE with mitochondrial ND1 site target when UGI is absent and when UGI is present (red boxes indicate adenine deaminase) (b) C-to-T and A-to-G base editing efficiency of DdABE with mitochondrial ND4 site target when UGI is absent and when UGI is present (red boxes indicate adenine deaminase) (c) Single nucleotide unit of the DdABE with the highest efficiency among the components targeting mitochondrial ND1 site (green boxes indicate the site where TALE attaches) (d) Single nucleotide unit of the DdABE with the highest efficiency among the components targeting mitochondrial ND4 site (green boxes indicate the site where TALE attaches) [Figure 64](a) Top: Schematic diagram of a single TALE module where all constructs are in one TALE module (components include full-length DddAtox, adenine deaminase, and a TALE array). Bottom: Dual TALE module using two TALE modules (components include full-length DddAtox and a TALE array on one side, and adenine deaminase and a TALE array on the other). (b) Base editing efficiency of single-module and dual-module DdABE targeting mitochondrial ND1 sites. (c) Base editing efficiency of single-module and dual-module DdABE targeting mitochondrial ND4 sites. [Figure 65] This figure shows the results of verifying the base editing efficiency of a single module targeting the ND1 site (components being a TALE array, adenine deaminase (AD), and full-length DddAtox (GSVG, AAAA, and E1347A are variants)). [Figure 66] This figure shows the results of verifying the base editing efficiency of a dual module targeting the ND1 site (components being a TALE array, adenine deaminase (AD), and full-length DddAtox (GSVG, AAAA, and E1347A are variants)). [Figure 67] (a) This figure shows the base editing efficiency when only TadA(AD) adenine deaminease is attached to a TALE-binding protein that targets the ND1 site. (b) This figure shows the base editing efficiency when only TadA(AD) adenine deaminease is attached to a TALE-binding protein that targets the ND4 site. [Figure 68]The figures (from bottom to top) show the adenine and cytosine base editing efficiencies for dual-module, single-module, and split-DddA-AD in TALEs targeting the ND1 site. When UGIs are present on both sides, only cytosine base editing occurs. When AD is attached to one side, both cytosine and adenine base editing occur, and when there is no UGI, selective adenine base editing occurs. Similarly, in dual-module and single-module structures, selective adenine base editing occurs. [Best Mode for Carrying Out the Invention]
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by skilled professionals in the art to which this invention pertains. Generally, the nomenclature used herein is well known and commonly used in the art.
[0045] As used herein, “editing” is interchangeable with “editing” and refers to a method of altering a nucleic acid sequence by selective deletion of a specific genomic target. Such specific genomic targets include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.
[0046] As used herein, "single base" refers to one nucleotide in a nucleic acid sequence. When used in the context of single-base editing, this means that a base at a specific position in the nucleic acid sequence is replaced by a different base. Such substitutions can occur through many mechanisms, including substitution or modification without limitation.
[0047] As used herein, “target” or “target site” means a pre-defined nuclear sequence of any composition and / or length. Such target sites include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.
[0048] As used herein, “on target” means a subsequence of a specific genomic target that may be fully complementary to a programmable DNA-binding region and / or a single guide RNA sequence.
[0049] As used herein, “off-target” means a subsequence of a specific genomic target that may be partially complementary to a programmable DNA-binding region and / or a single guide RNA sequence.
[0050] 1. Split cytosine deaminonase The fusion protein according to the present invention is a fusion protein comprising cytosine deaminose or a variant thereof, wherein the cytosine deaminose or its variant comprises a first split and a second split derived from the cytosine deaminose or its variant, and the first and second splits are each bound to the DNA-binding protein.
[0051] The cytosine deaminosease is an amino group removal enzyme that can convert cytosine (C) to uridine (U).
[0052] The cytosine deaminosease may be the aforementioned cytosine deaminosease. Examples of the cytosine deaminosease include APOBEC1 (apolipoprotein B editing complex 1) and AID (activation-induced deaminase). However, most DNA deaminoses act only on single-stranded DNA and may not be suitable for ligating to DNA-binding proteins to perform base editing. Specifically, the cytosine deaminosease may be derived from a deaminosease (DddA) that acts on double-stranded DNA or its orthologue. More specifically, the cytosine deaminosease may be a double-stranded DNA-specific bacterial cytosine deaminosease.
[0053] The cytosine deaminosease is in a divided form, comprising an isolated first divided portion and a second divided portion, neither of which possesses deaminosease activity.
[0054] The full-length cytosine deaminose may include the sequence of Sequence ID No. 1, which corresponds to the tox split. The cytosine deaminose includes a first split and a second split, and neither the first nor the second split possesses deaminose activity.
[0055] [Sequence ID 1] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0056] In one embodiment, the first or second split of the cytosine deaminose may contain, at the N-terminus, one or more sequences selected from the group consisting of G33, G44, A54, N68, G82, N98, and G108 in the sequence of SEQ ID NO: 1. The first or second split of the cytosine deaminose may contain, at the C-terminus, one or more sequences selected from the group consisting of G34, P45, G55, N69, T83, A99, and A109 in the sequence of SEQ ID NO: 1.
[0057] Specifically, the cytosine deaminose may include the first split of SEQ ID NO: 23 (G1333-N) and the second split of SEQ ID NO: 24 (G1333-C), or the first split of SEQ ID NO: 25 (G1397-N) and the second split of SEQ ID NO: 26 (G1397-C), or the first split of SEQ ID NO: 23 (G1333-N) and the second split of SEQ ID NO: 26 (G1397-C), or the first split of SEQ ID NO: 25 (G1397-N) and the second split of SEQ ID NO: 24 (G1333-C).
[0058] (Sequence ID 23) Wild type DddAtox G1333-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (Sequence ID 27) GGCTCTGGTTCCTACGCCCTGGGTCCATATCAGATTAGTGCTCCCAACTCCCCGCCTACAACGGTCAGACAGTGGGGACCTTTTACTATGTCAACGACGCCGGGGGATTGGAATCCAAGGTTTTCTCTAGCGGTGGG
[0059] (Sequence ID 24) Wild type G1333-C PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (Sequence ID 28) CCAACACCTTATCCTAACTACGCTAACGCCGGGCACGTCGAGGGGCAGTCAGCTCTTTTTATGAGAGATAACGGCATTAGCGAAGGGCTTGTGTTCCATAATAATCCTGAGGGCACCTGTGGCTTCTGTGTAAATATGACC GAAACACTTCTGCTGAGAACGCTAAAATGACTGTCGTACCACCCGAAGGCGCAATCCCAGTTAAACGGGGCGCAACCGGCGAAACCAAAGTATTCACCGGAAACAGCAATAGTCCAAAGTCCCCCACCAAGGGAGGTTGC
[0060] (Sequence ID 25) Wild type DddAtox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0061] (Sequence ID 29) GGTAGCTACGCACTTGGTCCTTACCAGATTAGCGCACCCCAACTCCCCGCCTATAATGGTCAAACCGTCGGGACCTTTTACTACGTAAACGATGCTGGTGGGCTGGAATCCAAAGTATTCTCCTCAGGGGGCCCTACACCCTACCCCAACTACGCCAATGCT GGTCATGTAGAAGGGCAGTCAGCACTGTTTATGCGCGATAATGGTATAAGCGAGGGGTTGGTCTTCCATAACAACCCAGAGGGTACTTGTGGCTTCTGTGTGAATATGACTGAAACCCTTCTGCCCGAAAATGCCAAGATGACTGTCGTCCCACCTGAAGGC
[0062] (Sequence ID 26) Wild type DddAtox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0063] (Sequence ID 30) GCCATACCTGTGAAGCGGGGAGCAACAGGGGAGACAAAGGTGTTCACAGGCAACTCTAACAGTCCAAAGAGCCCCACCAAAGGCGGGTGT G1333N, G1333C, G1397N, and G1397C can be combined and used as split deaminationases. Specifically, they can be used in the form of Left-G1333-N+Right-G133-C, Left-G1397-N+Right-G1397-C, Left-G1397-N+Right-G1333-C, or Left-G1333-N+Right-G1397-C.
[0064] 2. Mutants The inventors of this application attempted to suppress undesirable base editing by using a DdCBE mutation in which an amino acid residue is substituted. We present a highly accurate DddA-derived cytosine base editor that can reduce the untargeted effects of DdCBE. The effect of such untargeted base editing is independent of the interaction between TALE and DNA, and DddA tox This phenomenon is caused by the spontaneous binding of deamination enzyme derivatives.
[0065] In this case, the amino acid residue to which the mutation is introduced is a contact site located on the surface where the dimers of DddAtox bind to each other. DddA tox HF-DdCBE was constructed by substituting alanine for amino acid residues located on the surface between the cleavage cells. HF-DdCBE was designed to fail to function properly if the two deaminationase pairs linked to TALE could not bind to DNA. Whole mitochondrial genome analysis confirmed that HF-DdCBE is highly efficient and precise, unlike conventional DdCBE which causes numerous unwanted, untargeted C-to-T conversions in human mitochondrial DNA.
[0066] In the case of DddAtox, theoretically, base editing cannot occur unless all of the split dimers are recruited to the target site on the DNA. However, actual experiments have shown that targeted base editing occurs even when one DdCBE is bound to DNA while the other is not. To solve this, attempts are made to prevent the binding of DdCBE pairs at undesirable positions by substituting the residues on the protein surface to which the split dimers bind to each other.
[0067] Based on this, the present invention relates to a novel variant of the cytosine deaminationase DddAtox in which amino acid residues are substituted to reduce non-selective base editing.
[0068] The cytosine deaminose enzyme comprises either the first split of SEQ ID NO: 23 (G1333-N) and the second split of SEQ ID NO: 24 (G1333-C), or the first split of SEQ ID NO: 25 (G1397-N) and the second split of SEQ ID NO: 26 (G1397-C).
[0069] The mutant of the cytosine deaminose may have one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30, and 31 of the first split of SEQ ID NO: 23 substituted with other amino acids. The mutant of the second split of SEQ ID NO: 24 has one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 substituted with other amino acids.
[0070] The mutant according to the present invention is obtained in which one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 10.0, 101, 102, and 103 of the first split of Sequence ID No. 25, or one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of the second split of Sequence ID No. 26, are substituted with other amino acids.
[0071] The aforementioned "other amino acids" refers to alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, glutamic acid, algin, histidine, lysine, and all known variants of the aforementioned amino acids, excluding the amino acid that the wild-type protein has in its original mutant position. For example, the aforementioned "other amino acids" may be alanine.
[0072] Specifically, the first split product of Sequence ID No. 23 may include at least one amino acid substitution selected from the group consisting of Y3A, L5A, I10A, S11A, V13A, G14A, T15A, F16A, Y17A, Y18A, V19A, K28A, F30A, and S31A (corresponding to Y1292A, L1294A, I1299A, S1300A, V1312A, G1313A, T1314A, F1315A, Y1316A, Y1317A, V1318A, K1327A, F1329A, and S1330A, respectively).
[0073] The second split of Sequence ID No. 24 may contain one or more amino acid substitutions selected from the group consisting of V13A, Q16A, S17A, F20A, M21A, E28A, G29A, L30A, V31A, F32A, H33A, K56A, M57A, T58A, and V60A (corresponding to V1346A, Q1349A, S1350A, F1353A, M1354A, E1361A, G1362A, L1363A, V1364A, F1365A, H1366A, K1389A, M1390A, T1391A, and V1393A, respectively).
[0074] Specifically, the first split product of Sequence ID No. 25 may contain one or more amino acid substitutions selected from the group consisting of C87A, V88A, T91A, E92A, L95A, K100A, M101A, T102A, and V103A (corresponding to C1376A, V1377A, T1380A, E1381A, L1384A, K1389A, M1390A, T1391A, and V1392A, respectively).
[0075] The second split product of Sequence ID No. 26 may contain one or more amino acid substitutions selected from the group consisting of K13A, V14A, F15A, and T16A (corresponding to K1410A, V1411A, F1412A, and T1413A, respectively).
[0076] TIFF2026062789000002.tif230170TIFF2026062789000003.tif86170
[0077] TIFF2026062789000004.tif234170TIFF2026062789000005.tif251170TIFF2026062789000006.tif96170
[0078] TIFF2026062789000007.tif234170TIFF2026062789000008.tif166170
[0079] TIFF2026062789000009.tif78170
[0080] The cytosine deaminase variant according to the present invention may include one or more sequences selected from the group consisting of the amino acid sequences described as above. The cytosine deaminase variant according to the present invention suggests the possibility of reducing the effect of causing unwanted editing of a large number of bases within a non-specific target site.
[0081] 3. Full-length cytosine deaminase The inventors of the present application have developed a new programmable cytosine deaminase using full-length DddA, which was created by modifying the positively charged amino acids of wild-type cytosine deaminase (Cytosine deaminase) DddA that causes toxicity in cells and is used in a split form. tox The present invention relates to a fusion protein comprising (i) a DNA-binding protein; and (ii) a cytosine deaminase or a variant thereof, wherein the cytosine deaminase or the variant thereof is a non-toxic full-length cytosine deaminase.
[0082] At the C-terminus of DddA tox positively charged amino acids (KRKKK) specifically gather. Since DNA has a negative charge, it binds to the positively charged amino acids of the protein. Substituting the positively charged amino acids with amino acids without electrodes, DddA tox This weakens the ability to bind to DNA, thereby reducing or eliminating intracellular toxicity. In other words, by substituting positively charged amino acids with non-polar amino acids, non-toxic combinations can be cloned using E. coli, and full-length DddA can be secured.
[0084] Wild-type DddA tox This method uses a separated form divided into two due to intracellular toxicity, but this imposes many limitations on experiments. In particular, when using Cas9, an orthogonal Cas9 mutant that recognizes other PAMs is used. In this case, because PAMs are restrictively present, it is often difficult to precisely deaminate cytosine to thymine at the desired location. Also, the target window in which the highest activity is expressed is 40 bp between the binding of the two Cas9 mutants, and undesirable cytosine at this site is also deaminated. However, since the full-length DddA is not separated, it is not subject to the constraints of PAMs. Furthermore, since the highest activity is observed in the TC motif within 10 bp at the target site, it also offers high accuracy.
[0085] We confirmed that Cas9 deaminates cytosine to thymine in the ACA, GC, and CC motifs of the R-loop formed by binding to the target site. This activity is not observed in the isolated form.
[0086] Full-length DddA can replace cytosine with thymine at a desired location using not only Cas9 but also the TALE module and zinc finger protein. Conventional DddAtox is in a separated form and therefore had to be transmitted as a pair, but full-length DddA can use only one module of the TALE module and zinc finger protein, allowing for unrestricted selection of the target site. Furthermore, it can target DNA not only at genomic sites but also at mitochondria, plant chloroplasts, or plastids, converting cytosine to thymine in specific DNA.
[0087] Furthermore, its small size allows the entire structure to be incorporated into AAV, a viral vector used in gene therapy. Conventional CBEs (cytosine base editors) replace cytosine with thymine in the R-loop formed by Cas9 binding to the target site, but the newly invented full-length DddA deaminates cytosine outside the R-loop. Therefore, it can convert cytosine to thymine in locations where editing with conventional CBEs is restricted.
[0088] Based on this, the non-toxic full-length cytosine deaminosease may have one or more, two or more, three or more, four or more, or five or more amino acids substituted for other amino acids in the wild-type deaminosease of Sequence ID No. 1.
[0089] The aforementioned "other amino acids" refers to alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, glutamic acid, algin, histidine, lysine, and all known variants of the aforementioned amino acids, excluding the amino acid that the wild-type protein originally has in its mutant position. For example, the aforementioned "other amino acids" may be alanine.
[0090] The aforementioned non-toxic full-length DddA may include sequences selected from the group consisting of sequence numbers 12 to 18, depending on the type.
[0091] Wild type (SEQ ID No: 1) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0092] A1341D KRKKA (SEQ ID No: 12) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTAGGC
[0093] AAAAA (SEQ ID No: 13) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0094] AAAAK (SEQ ID No: 14) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC
[0095] AAKAA (SEQ ID No: 15) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC
[0096] AAKAK (SEQ ID No: 16) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC
[0097] KAAAA (SEQ ID No: 17) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC
[0098] E1347A (SEQ ID No: 18) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0099] The full-length deaminosease mutant may contain one or more substitutions selected from the group consisting of the following in the amino acid sequence of SEQ ID NO: 1: The 37th place S is replaced with G. The 59th ranked G is replaced with S. The 109th ranked A is replaced with V, and The 129th ranked S is replaced with G.
[0100] In one embodiment, the full-length deaminosease mutant may include the sequence of Sequence ID No. 19, which includes the substitution of S at position 37 with G, G at position 59 with S, A at position 109 with V, and S at position 129 with G in the amino acid sequence of Sequence ID No. 1.
[0101] GSVG (SEQ ID No: 19) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0102] The full-length DddA GSVG can be cloned using common E. coli. It was confirmed at the human cell genomic context that full-length DddA GSVG deaminates cytosine to thymine in TC Motif at the target site. Full-length DddA GSVG can be cloned to both the N-terminus and C-terminus of Cas9. DddA GSVG ligated to the N-terminus of Cas9 at the same target site can replace cytosine with thymine. It was confirmed at the human cell genomic level that DddA GSVG ligated to the C-terminus of Cas9 replaces cytosine to thymine in the complementary sequence of TC Motif at the desired guanine, replaces cytosine to thymine in the complementary sequence of TC Motif at the desired guanine, and replaces the desired guanine with adenine.
[0103] In one embodiment, the full-length deaminosease mutant may include a sequence selected from the group consisting of SEQ ID NOs: 20 to 22.
[0104] SSVG (SEQ ID No: 20) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0105] GSAG (SEQ ID No: 21) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0106] GSVS (SEQ ID No: 22) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0107] 4. DNA-binding proteins The DNA-binding protein may be, for example, a zinc finger protein, a TALE (Transcription Activator-like Effector) protein, a CRISPR-related nuclease, or a combination of two or more of these.
[0108] The zinc finger motif of the aforementioned zinc finger protein has a DNA-binding domain, and the C-terminal portion of the finger specifically recognizes DNA sequences. A DNA-binding protein containing 3 to 6 zinc finger motifs recognizes DNA sequences.
[0109] In one embodiment, the first and second cytosine deaminose compounds can each be bound to the N-terminus or C-terminus of a zinc finger protein.
[0110] The C-terminus of the zinc finger protein (ZF-Left) is bound to the N-terminus of the first cytosine deaminose split, and the C-terminus of the zinc finger protein (ZF-Right) is bound to the N-terminus of the second cytosine deaminose split (CC configuration). The N-terminus of the zinc finger protein (ZF-Left) is bound to the C-terminus of the first cytosine deaminose split, and the C-terminus of the zinc finger protein (ZF-Right) is bound to the N-terminus of the second cytosine deaminose split (NC configuration). The C-terminus of the zinc finger protein (ZF-Left) is bound to the N-terminus of the first cytosine deaminose split, and the N-terminus of the zinc finger protein (ZF-Right) is bound to the C-terminus of the second cytosine deaminose split (CN configuration), or, The N-terminus of a zinc finger protein (ZF-Left) can bind to the C-terminus of the first cytosine deaminose split, and the N-terminus of a zinc finger protein (ZF-Right) can bind to the C-terminus of the second cytosine deaminose split (NN configuration).
[0111] The aforementioned ZF-Left may include the sequence of sequence number 2 below.
[0112] [Sequence ID 2] GIHGVPAAMAERPFQCRICMRNFSDRSNLSRHIRTHTGEKPFACDICGRKFAISSNLNSHTKIHTGSQKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFATSGNLTRHTKIHLR.
[0113] The aforementioned ZF-Right may include the sequence of sequence number 3 below.
[0114] [Sequence ID 3] GIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR.
[0115] ZF sequences may vary depending on the DNA target. ZFs can be customized according to the DNA target sequence. Since ZFs recognize 3 bp of DNA, typically 3 to 6 ZFs can be ligated together to create ZF combinations that recognize 9 to 18 bp of DNA. For example, these can be created using a library containing modules such as GNN, TNN, CNN, or ANN.
[0116] In some cases, the zinc finger protein may be linked to a deaminose enzyme via a linker. The linker may be a peptide linker containing 2 to 40 amino acid residues. The linker may be, but is not limited to, linkers of length 2aa, 5aa, 10aa, 16aa, 24aa, or 32aa.
[0117] In one embodiment, the linker may include the following: 2a.a Linker: GS, 5a.a Linker: TGEKQ (Sequence ID 8), 10a.a Linker: SGAQGSTLDF (Sequence ID 9), 16a.a Linker: SGSETPGTSESATPES (Sequence ID 10), 24a.a Linker:SGTPHEVGVYTLSGTPHEVGVYTL (Sequence ID 115), or, 32a.a Linker: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS (Sequence ID 11)
[0118] In a specific embodiment of the present invention, the split deaminose enzyme and the zinc finger protein can be linked via a linker, with the zinc finger protein binding to the N-terminus of half the deaminose enzyme containing the first split, and the zinc finger protein binding to the N-terminus of half the deaminose enzyme containing the second split. At this time, a C-to-T base transposition can occur at a spacer between the attachment sites of the left and right ZFPs. It has been confirmed that high editing efficiency is observed when all of the left and right ZFPs are linked to each of the two splits of the deaminose enzyme via a 24a.a linker.
[0119] The aforementioned TAL effector (TALE) consists of a repeating structure of 33 to 34 amino acid sequences, with approximately 9 repeating domains (RVD, Repeated variant domain). Each domain can recognize one nucleotide and can bind to a specific DNA sequence according to the 12th to 13th amino acid sequence (HD → Cytosine, NI → Adenine, NG → Thymine, NN → Guanine). The TAL effector (TALE) recognizes a single DNA strand within a target site. The distance between target sites can be 12 to 14 nucleotides.
[0120] The aforementioned TALE domain refers to a protein domain that binds to nucleotides in a sequence-specific manner through a combination of one or more TALE-Repeats. It contains, but is not limited to, at least one TALE-Repeat, specifically 1 to 30 TALE-Repeats. A TALE-Repeat is a site within a TALE domain that recognizes a specific nucleotide sequence.
[0121] The TALE domain comprises an N-terminal enclosing region and a C-terminal enclosing region of the TALE as a skeletal structure. The first TALE, which encloses the N-terminal region of the TALE, can be coded by sequence number 4 or 5. The second TALE, which encloses the C-terminal region of the TALE, can be coded by sequence number 6 or 7.
[0122] [Table 1]
[0123] Based on the cleavage site, a single TALE array or a first TALE array and a second TALE array can be joined, depending on the position where the TALE domains join.
[0124] The first TALE can bind to the first split of the cytosine deaminose (left TALE), and the second TALE can bind to the second split of the cytosine deaminose (right TALE). These have the structures N'-TALE-first split-C' and N'-TALE-second split-C', respectively.
[0125] If the cytosine deaminose is full-length, a single module TALE can be bound to the N-terminus of the cytosine deaminose. The single TALE module and cytosine deaminose are included in the NC direction. A dual module TALE may be included, in which case a first TALE may be bound to the N-terminus of the full-length cytosine deaminose, and a second TALE may be included separately. The first TALE module and cytosine deaminose are included in the NC direction, and have the structures N'-TALE-cytosine deaminose-C' and N'-TALE-C'.
[0126] TALE arrays can be customized to match the target DNA sequence. A TALE array is composed of repeating modules consisting of 33 to 35 amino acid residues, derived from the plant pathogen Xanthmonas. Each module recognizes and binds to DNA, specifically one A, C, G, and T base. The base specificity of each module is determined by the 12th and 13th amino acid residues, which are called repeat variable di-residues (RVDs). For example, a module with RVD NN recognizes G, NI recognizes A, HD recognizes C, and NG recognizes T. A TALE array can be composed of at least 14 to 18 or more modules and designed to recognize target DNA sequences of 15 to 20 bp.
[0127] Regarding the aforementioned CRISPR-related nuclease, the CRISPR array contains two RNAs: one is crRNA (CRISPR RNA) and the other is tracRNA (trans-activating CRISPR RNA). crRNA is transcribed in the protospace region and binds to tracRNA to form a tertiary structure. These two RNAs help recognize and cleave external DNA.
[0128] The aforementioned Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Csy1, Csy2, Csy3, Cse1, and Cs The endonucleases may be, but are not limited to, e2, Csc1, Csc2, Csa5, Csn2, CsMT2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, or Csf4.
[0129] The aforementioned Cas protein is found in Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus (Streptococcus pyogenes), Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus (Staphylococcus aureus), The Cas protein may be derived from a microbial genus containing an orolog selected from the group consisting of Nitratifractor, Corynebacterium, and Campylobacter, and may be simply isolated or recombinant from these genus.
[0130] The Cas protein may also contain mutated forms. This could mean that the protein has been mutated to lose its endonuclease activity, which breaks DNA double strands. For example, it could be one or more mutated forms such as a target-specific nuclease that has lost endonuclease activity but possesses nickase activity, or a form that has lost both endonuclease activity and nickase activity.
[0131] If the molecule possesses nickase activity, a nick may be introduced on the chain in which the base conversion occurred or its reverse chain (for example, the reverse chain of the chain in which the base conversion occurred) simultaneously with or sequentially, regardless of the order, the base conversion by the cytosine deaminose (for example, the nick may be introduced on the reverse chain of the chain in which PAM is located, at a position corresponding to the 3rd and 4th nucleotides in the 5' end direction of the PAM sequence). Such mutations (for example, amino acid substitutions) can occur in the catalytic domain (for example, the RuvC catalytic domain in the case of Cas9). In the case of Cas9 derived from Streptococcus pyogenes, the mutation may include a mutation in which one or more other amino acids selected from the group consisting of catalytic aspartate residues (such as aspartate at position 10 (D10)), glutamic acid at position 762 (E762), histidine at position 840 (H840), asparagine at position 854 (N854), asparagine at position 863 (N863), and aspartate at position 986 (D986). In this case, any other amino acid to be substituted may be alanine, but is not limited to alanine.
[0132] In some cases, one or more of the aspartic acid at position 1135 (D1135), arginine at position 1335 (R1335), and threonine at position 1337 (T1337) of the Streptococcus pyogenes Cas9 protein, for example, all three of them, may be mutated to recognize an NGA (where N is any base selected from A, T, G, and C) that differs from the PAM sequence (NGG) of wild-type Cas9.
[0133] For example, among the amino acid sequences of the Cas9 protein derived from Streptococcus pyogenes, (1) D10, H840, or D10+H840, (2) D1135, R1335, T1337, or D1135+R1335+T1337, or (3) Amino acid substitutions can occur in both residues (1) and (2).
[0134] The aforementioned "other amino acids" means alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, glutamic acid, arginine, histidine, lysine, and all known variants of the aforementioned amino acids, excluding the amino acid that the wild-type protein has at its original mutant position. For example, the aforementioned "other amino acids" may be alanine, valine, glutamine, or arginine.
[0135] In some cases, the guide RNA may be further included. The guide RNA may be one or more selected from the group consisting of, for example, CRISPR RNA (crRNA), trans-activated crRNA (tracrRNA), and single guide RNA (single guide RNA; sgRNA), and specifically, it may be a double-stranded crRNA:tracrRNA complex in which crRNA and tracrRNA are bound to each other, or a single-stranded guide RNA (sgRNA) in which crRNA or a part of tracrRNA and a part of tracrRNA are linked by an oligonucleotide linker.
[0136] 5. Addition of adenine deaminose The present invention relates to a fusion protein comprising (i) a DNA-binding protein, (ii) cytosine deaminoase or a variant thereof, and (iii) adenine deaminoase, wherein the cytosine deaminoase or variant thereof comprises (a) non-toxic full-length cytosine deaminoase or (b) a first and second split derived from cytosine deaminoase or a variant thereof, and the first and second splits are each in a form that binds to the DNA-binding protein.
[0137] The inventors of this application have created a base editor capable of editing A bases by linking a DddAtox cytosine deaminosease and an adenine deaminosease capable of A-to-G conversion to a TALE or ZFP protein that can bind to DNA.
[0138] Conventional DddAtox-based deaminases (DdCBEs) are cytosine deaminoses that use TALE repeats as the DNA binding module. Unlike DdCBEs, which only induce C-to-T mutations, DdABEs can induce A-to-G mutations, thus creating different mutation patterns.
[0139] DdABEs recognize double-stranded DNA and induce deamination on their own, thus lacking additional components such as RNA. In the case of mitochondria and chloroplasts, the mechanism for delivering RNA is unknown, making it impossible to apply the CRISPR system. However, DdABEs, lacking such components, can not only target genomic DNA within cells but also target DNA in organelles like mitochondria and chloroplasts to induce A-to-G transmutation of specific DNA.
[0140] Currently, DdCBE is the only gene editing technology that targets mitochondria and organelles. Therefore, while all conventional technologies combined could only introduce C-to-T mutations, DdABE can induce A-to-G mutations, resulting in a much more diverse spectrum of mutations that can be introduced. This makes it possible to create or treat mitochondrial disease models that were previously impossible.
[0141] Conventional DdCBEs require two TALE modules (one attached to the left and one to the right), making them unsuitable for use in gene therapy with AAV, a viral vector with low gene capacity. However, DdABE can be used as a "single module," requiring only one TALE module, allowing it to be used with AAV viruses and giving it a significant advantage in gene therapy.
[0142] DdABE offers high compatibility because it can be used with split DddAtox as needed, or with the full-length DddAtox variant.
[0143] The adenine deaminose can be selected from the group consisting of, for example, APOBEC1 (apolipoprotein B editing complex 1), AID (activation-induced deaminase), and tadA (tRNA-specific adenosine deaminase), and more specifically, tadA (tRNA-specific adenosine deaminase). The adenine deaminose may also be, for example, deoxy-adenine deaminase as a mutant of Escherichia coli TadA.
[0144] The cytosine deaminose is contained in a divided form, and the DNA-binding protein is a zinc finger protein, with the N-terminus of the zinc finger protein (ZF-Left) bound to the C-terminus of the first divided cytosine deaminose, and the C-terminus of the zinc finger protein (ZF-Right) bound to the N-terminus of the second divided cytosine deaminose (NC configuration). The adenine deaminose can bind to the C-terminus of the zinc finger protein (ZF-Left), the N-terminus or C-terminus of the first divided cytosine deaminose, the N-terminus of the zinc finger protein (ZF-Right), or the N-terminus or C-terminus of the second divided cytosine deaminose.
[0145] Adenine deaminose can be configured in the following ways: CC configuration (where the C-terminus of zinc finger protein (ZF-Left) is bound to the N-terminus of the first cytosine deaminose molecule, and the C-terminus of zinc finger protein (ZF-Right) is bound to the N-terminus of the second cytosine deaminose molecule); CN configuration (where the C-terminus of zinc finger protein (ZF-Left) is bound to the N-terminus of the first cytosine deaminose molecule, and the N-terminus of zinc finger protein (ZF-Right) is bound to the C-terminus of the second cytosine deaminose molecule); or NN configuration (where the N-terminus of zinc finger protein (ZF-Left) is bound to the C-terminus of the first cytosine deaminose molecule, and the N-terminus of zinc finger protein (ZF-Right) is bound to the C-terminus of the second cytosine deaminose molecule). In this configuration, it can bind to the C-terminus of the zinc finger protein (ZF-Left), the N-terminus or C-terminus of the first cleavage of cytosine deaminose, the N-terminus of the zinc finger protein (ZF-Right), and the N-terminus or C-terminus of the second cleavage of cytosine deaminose.
[0146] If the cytosine deaminosease is contained in a split form and the DNA-binding protein is TALE, the first TALE binds to the first split of the cytosine deaminosease, and the second TALE binds to the second split of the cytosine deaminosease, resulting in the structures N'-TALE-first split DDDA-C' and N'-TALE-second split DDDA-C', respectively. Adenine deaminosease can bind to the N-terminus or C-terminus of the first split of the cytosine deaminosease or to the N-terminus or C-terminus of the second split of the cytosine deaminosease.
[0147] If the cytosine deaminosease is contained in full form and the DNA-binding protein is TALE, it may contain a single TALE module, and the single TALE module and cytosine deaminosease may be contained in the NC direction, in which case adenine deaminosease may be bound to the C-terminal direction of the single TALE module, or to the N-terminus or C-terminus of the cytosine deaminosease.
[0148] If the cytosine deaminose is contained in full form and the DNA-binding protein is TALE, a dual TALE module may be included, further comprising a second split containing a first TALE module and cytosine deaminose in the NC direction, adenine deaminose and a second TALE, and the adenine deaminose can be bound to the N-terminus or C-terminus of TALE in the structures N'-TALE-cytosine deaminose-C' and N'-TALE-adenine deaminose-C'.
[0149] In some cases, the formula may further contain a UGI (uracil DNA glycosylase inhibitor) that can enhance base editing efficiency. UGIs can enhance base editing efficiency by inhibiting the activity of UDG (uracil DNA glycosylase), an enzyme that repairs mutated DNA and catalyzes the removal of u from DNA.
[0150] The present invention relates to a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminosease of the fusion protein or a variant thereof is derived from bacteria and is specific to double-stranded DNA.
[0151] The present invention relates to a composition for A-to-G base editing in prokaryotic or eukaryotic cells comprising the fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease. The present invention relates to a composition in which the cytosine deaminoase or a variant of the fusion protein is derived from bacteria, is specific to double-stranded DNA, has a DNA-binding protein bound to the N-terminus of the cytosine deaminoase or its variant, and has a DNA-binding protein bound to the C-terminus of the adenine deaminoase of the fusion protein.
[0152] The present invention relates to a composition for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or its variant is derived from bacteria and is specific to double-stranded DNA.
[0153] Specifically, the present invention relates to a composition for A-to-G base editing in prokaryotic and eukaryotic cells (without UGI) comprising 1) a DNA-binding protein, 2) a full-length double-stranded DNA-specific bacterial cytosine deaminoase or a variant thereof, and 3) a deoxyadenine deaminoase derived from E. coli TadA, wherein the DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalyst-deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the full-length double-stranded DNA-specific bacterial cytosine deaminoase is DddAtox derived from Burkholderia cenocepacia.
[0154] The present invention relates to a composition for A-to-G base editing in prokaryotic and eukaryotic cells (without UGI), comprising: 1) a left-side DNA-binding protein functionally linked to a full-length double-stranded DNA-specific bacterial cytosine deaminoase or a variant thereof; and 2) a right-side DNA-binding protein functionally linked to a deoxyadenine deaminoase derived from E. coli TadA, wherein the left-side or right-side DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalyst-deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and is a full-length double-stranded DNA-specific bacterial cytosine deaminoase derived from Burkholderia cenocepacia (DddAtox).
[0155] The present invention also relates to a composition for A-to-G and C-to-T base editing in prokaryotic and eukaryotic cells comprising 1) a DNA-binding protein, 2) a full-length double-stranded DNA-specific bacterial cytosine deaminoase or a variant thereof, 3) a deoxyadenine deaminoase derived from E. coli TadA, and 4) a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalyst-deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and is a full-length double-stranded DNA-specific bacterial cytosine deaminoase derived from Burkholderia cenocepacia (DddAtox).
[0156] The present invention further comprises a composition (without UGI) for A-to-G base editing in prokaryotic and eukaryotic cells, comprising: 1) a DNA-binding protein; 2) a split double-stranded DNA-specific bacterial cytosine deaminoase or a variant thereof; and 3) a deoxyadenine deaminoase derived from E. coli TadA, wherein the DNA-binding protein is a zinc finger protein or a transcription activator-like effector (TALE) array or a catalyst-deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the split double-stranded DNA-specific bacterial cytosine deaminoase is DddAtox derived from Burkholderia cenocepacia.
[0157] The present invention further comprises a composition for A-to-G and C-to-T base editing in prokaryotic and eukaryotic cells, comprising: 1) a DNA-binding protein; 2) a split double-stranded DNA-specific bacterial cytosine deaminoase or a variant thereof; 3) a deoxyadenine deaminoase derived from E. coli TadA; and 4) a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalyst-deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and is a split double-stranded DNA-specific bacterial cytosine deaminoase derived from Burkholderia cenocepacia (DddAtox).
[0158] The present invention relates to a method for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or its variant is of bacterial origin and is specific to double-stranded DNA.
[0159] The present invention relates to a method for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it, wherein the DNA-binding protein is zinc pinker protein, TALE protein, or CRISPR-related nuclease, the cytosine deaminosease or a variant of the fusion protein is of bacterial origin, is specific to double-stranded DNA, has a DNA-binding protein bound to the N-terminus of the cytosine deaminosease or its variant, and has a DNA-binding protein bound to the C-terminus of the adenine deaminose of the fusion protein.
[0160] The present invention relates to a method for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or its variant is of bacterial origin and is specific to double-stranded DNA.
[0161] The specific arrangement of components included in the composition or method according to the present invention is as follows: ND1-ZFP-Right-1397C-AD (Figure 61f-g: SEQ ID No: 410) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0162] ND1-ZFP-Left-1397C-UGI (Figure 61f-g: SEQ ID No: 411) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0163] ND1-ZFP-Right-1397N-UGI (Figure 61f-g: SEQ ID No: 412) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTLYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0164] Trac-ZFP-LEFT-1397N-UGI (Figure 61d: SEQ ID No: 413) MAPKKKRKVGIHGVPAAMGGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTLFQCRICMRKFATSGSLTRHTKIHTGEKPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFATSSNRTKHTKIHTHPRAPIPKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAWHSSLRVHTKIHLRSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0165] Trac-ZFP-right-1397C-AD (Figure 61d: SEQ ID No: 414) MAPKKKRKVGIHGVPAAMAERPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFADRSHLARHTKIHTGSQKPFQCRICMRKFALKQHLNEHTKIHTGEKPFQCRICMRNFSQSGNLARHIRTHTGEKPFACDICGRKFAHNSSLKDHTKIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0166] ND1-Left-TALE-1397N-UGI (Figure 62: SEQ ID No: 415) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0167] ND1-Left-TALE-1397C-UGI (Figure 62: SEQ ID No: 416) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0168] ND1-Right-TALE-1397N-UGI (Figure 62: SEQ ID No: 417)
[0169] ND1-Right-TALE-1397C-UGI (Figure 62: SEQ ID No: 418) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0170] ND4-Left-TALE-AD (Figure 62b: SEQ ID No: 419) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0171] ND4-Right-TALE-AD (Figure 62b: SEQ ID No: 420) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0172] ND1-Left-TALE-1333N-AD (Figure 63: SEQ ID No: 421) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0173] ND1-Left-TALE-1333C-AD (Figure 63: SEQ ID No: 422) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0174] ND1-Right-TALE-1333N-AD (Figure 63: SEQ ID No: 423)
[0175] ND1-Right-TALE-1333C-AD (Figure 63: SEQ ID No: 424)
[0176] ND1-Left-TALE-1397C-AD (Figures 62-63: SEQ ID No: 425) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0177] ND1-Right-TALE-1397C-AD (Figures 62-63: SEQ ID No: 426)
[0178] ND1-Left-TALE-1333N (Figure 63: SEQ ID No: 427) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG
[0179] ND1-Left-TALE-1333C (Figure 63: SEQ ID No: 428) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0180] ND1-Left-TALE-1397N (Figure 62-63: SEQ ID No: 429) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0181] ND1-Right-TALE-1333N (Figure 63: SEQ ID No: 430) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG
[0182] ND1-Right-TALE-1333C (Figure 63: SEQ ID No: 431) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0183] ND1-Right-TALE-1397N (Figures 62 - 63: SEQ ID No: 432) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0184] ND4-LEFT-TALE-1333N-AD (Figure 63: SEQ ID No: 433)
[0185] ND4-LEFT-TALE-1333C-AD (Figure 63: SEQ ID No: 434)
[0186] ND4-LEFT-TALE-1397C-AD (Figure 63: SEQ ID No: 435)
[0187] ND4-Right-TALE-1333N-AD (Figure 63: SEQ ID No: 436)
[0188] ND4-Right-TALE-1333C-AD (Figure 63: SEQ ID No: 437)
[0189] ND4-Right-TALE-1397C-AD (Figure 63: SEQ ID No: 438)
[0190] ND1-Left-TALE-AD-GSVG (Figure 64: SEQ ID No: 439)
[0191] ND1-Left-TALE-AD-E1347A (Figure 64: SEQ ID No: 440)
[0192] ND1-Left-TALE-AD-AAAAA (Figure 64: SEQ ID No: 441)
[0193] ND1-Right-TALE-AD (Figure 64: SEQ ID No: 442) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0194] ND1-Left-TALE-GSVG (Figure 64: SEQ ID No: 443) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0195] ND1-Left-TALE-E1347A (Figure 64: SEQ ID No: 444) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0196] ND1-Left-TALE-AAAAA (Figure 64 SEQ ID No: 445) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0197] ND1-TALE-left (SEQ ID No: 446) DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAAL
[0198] ND1-TALE-Right (SEQ ID No: 447) DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAAL
[0199] ND4-TALE-left (SEQ ID No: 448) MDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG
[0200] ND4-TALE-Right (SEQ ID No: 449) MDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHD GGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPE QVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALE TVQRLLPVLCQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPEQVVAIAS NIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG
[0201] TRAC-ZFP-LEFT (SEQ ID No: 450) FQCRICMRKFATSGSLTRHTKIHTGEKPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFATSSNRTKHTKIHTHPRAPIPKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAWHSSLRVHTKIHLR
[0202] TRAC-ZFP-Right (SEQ ID No: 451) FQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFADRSHLARHTKIHTGSQKPFQCRICMRKFALKQHLNEHTKIHTGEKPFQCRICMRNFSQSGNLARHIRTHTGEKPFACDICGRKFAHNSSLKDHTKIHLR
[0203] ND1-ZFP-Left (SEQ ID No: 452) FQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR
[0204] ND1-ZFP-Right (SEQ ID No: 453) YKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR
[0205] G1333N-DddAtox (SEQ ID No: 454) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG
[0206] G1333C-DddAtox (SEQ ID No: 455) GSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0207] G1397N-DddAtox (SEQ ID No: 456) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0208] G1397C-DddAtox (SEQ ID No: 457) GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0209] Adenine deaminase (AD: ABE 8e) (SEQ ID No: 458) SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0210] Full length DddAtox variant GSVG (SEQ ID No: 459) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0211] Full length DddAtox variant E1347A (SEQ ID No: 460) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0212] Full length DddAtox variant AAAAA (SEQ ID No: 461) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0213] UGI (SEQ ID No: 462) TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0214] Single module (TALE)-linker-AD-GSVG (SEQ ID No: 463) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0215] (TALE)-linker-AD-E1347A (SEQ ID No: 464) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0216] (TALE)-linker-AD-AAAAA (SEQ ID No: 465) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGVAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVTFGNSNSASPTAGGC
[0217] Dual module (TALE) -linker-AD (SEQ ID No: 466) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0218] (TALE)-GSVG (SEQ ID No: 467) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0219] (TALE)-E1347A (SEQ ID No: 468) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0220] (TALE)-AAAAA (SEQ ID No: 469) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0221] (TALE)-1397C-AD (SEQ ID No: 470) GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0222] (TALE)-1333N-AD (SEQ ID No: 471) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0223] (TALE)-1333C-AD (SEQ ID No: 472) GSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREV PVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0224] Link(AA means amino acid) 8AA: SGGGLGST (SEQ ID No: 473) 16AA: SGSETPGTSESATPES (SEQ ID No: 474) 32AA: SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID No: 475)
[0225] SOD2 MTS-3xHA: MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA (SEQ ID No: 476) COX8A MTS-3xFLAG: MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK (SEQ ID No: 477)
[0226] 5. Communication The fusion proteins according to the present invention can be delivered to cells by various methods in the art, such as, but are not limited to, microinjection, electroporation, DEAE-dextran treatment, lipofection, nanoparticle-mediated transfusion, protein delivery domain-mediated transfusion, and PEG-mediated transfusion.
[0227] In another aspect, the present invention relates to nucleic acids encoding the fusion protein. With respect to the nucleic acids, the terms "polynucleotide," "nucleotide," "nucleotide sequence," and "nucleotide" are used interchangeably. The polymer forms of nucleotides of any length may include deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. Polynucleotides may contain one or more nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure are possible before or after polymer assembly.
[0228] The polynucleotide may be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination sequence).
[0229] As a means of expressing the aforementioned fusion protein, known expression vectors such as plasmid vectors, cosmid vectors, and bacteriophage vectors can be used, and the vectors can be easily manufactured by those skilled in the art according to any known method using DNA recombination technology.
[0230] The vector may be a plasmid vector or a viral vector, and the viral vector may specifically be, but is not limited to, an adenovirus, adeno-associated virus, lentivirus, or retrovirus vector.
[0231] Recombinant expression vectors can contain nucleic acids in a form suitable for nucleic acid expression in host cells, meaning they contain one or more regulatory elements that can be selected based on the host cell so that the recombinant expression vector is used for expression, that is, functionally linked to the nucleic acid sequence to be expressed.
[0232] Within a recombinant expression vector, "functionally linked" means that the nucleotide sequence of interest is linked to a regulatory element in a manner that allows for the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system, or in the host cell if the vector is introduced into a host cell).
[0233] Recombinant expression vectors can contain a form suitable for messenger RNA synthesis, including the T7 promoter, which means they contain one or more regulatory elements that enable mRNA synthesis in vitro, i.e., messenger RNA synthesis by T7 polymerase.
[0234] "Regulatory elements" may include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression regulatory elements (e.g., transcription termination signals, e.g., polyadenylation signals and poly-U sequences). Regulatory elements include elements that direct the induction or constant expression of nucleotide sequences in many types of host cells, and elements that direct the expression of nucleotide sequences only in specific host cells (e.g., tissue-specific regulatory elements). Tissue-specific promoters can primarily direct expression in desired tissues of interest, e.g., muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). Regulatory elements can also direct expression in transient-dependent manner, such as cell-cycle-dependent or developmental-stage-dependent manner, which may or may not be tissue- or cell-type specific.
[0235] In some cases, a vector may contain one or more pol III promoters, one or more pol II promoters, one or more pol I promoters, or a combination thereof. Examples of pol III promoters include, non-limitingly, the U6 and H1 promoters. Examples of pol II promoters include, non-limitingly, the retroviral Roussarcoma virus (RSV) LTR promoter (optionally containing an RSV enhancer), the cytomegalovirus (CMV) promoter (selectively containing a CMV enhancer) (e.g., Boshart et al (1985) Cell 41:521-530), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerate kinase (PGK) promoter, and the EF1α promoter.
[0236] "Regulatory elements" may include enhancers, such as WPRE; CMV enhancer; R-U5' segment of the LTR of HTLV-I; SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin. It will be recognized by those skilled in the art that the design of expression vectors may depend on factors such as the selection of host cells to be transformed and the desired expression level. Vectors can be introduced into host cells and produce transcripts, proteins, or peptides containing fusion proteins or peptides encoded by the nucleic acids described herein (e.g., clustered regularly spaced short palindromic repeat (CRISPR) transcripts, proteins, enzymes, their mutants, their fusion proteins, etc.). Useful vectors include lentiviruses and adenoviruses, and types of such vectors may also be selected to target specific types of cells.
[0237] The vectors can be transmitted in vivo or intracellularly by local injection (e.g., direct injection into the lesion or target site), electroporation, lipofection, viral vectors, nanoparticles, as well as by methods such as PTD (Protein translocation domain) fusion protein therapy.
[0238] The nucleic acid can be injected in the form of ribonucleic acid, for example messenger ribonucleic acid mRNA, to enable gene base editing of cells, for example, in animal cells or plant cells without limitation.
[0239] The nucleic acid according to the present invention may be in mRNA form. When transmitted in mRNA form, the transcription process to mRNA is unnecessary compared to transmission using a DNA vector, resulting in faster initiation of gene editing. Transient protein expression is also highly likely.
[0240] The inventors of this application have confirmed that when a cytosine base editor is injected into plant cells in the form of ribonucleic acid, such as messenger ribonucleic acid, for plant organelle gene editing, the non-targeting effect is reduced compared to transmission via plasmid. They were the first to demonstrate that transforming plant cells with a cytosine base editor in mRNA form offers advantages in terms of non-targeting effect compared to plasmids in plant organelle gene editing.
[0241] The mRNA can be delivered directly or transmitted via a carrier. In some cases, the mRNA of nucleic acid cleavage enzymes and / or cleavage factors can be chemically modified or directly transmitted in the form of synthetic self-replicative RNA.
[0242] Methods for delivering mRNA to cells, or for delivering mRNA to cells of organisms such as humans and animals in vivo, are conceivable, and methods for delivering mRNA molecules to cells in vitro or in vivo are possible. For example, mRNA molecules can be delivered to cells by including lipids (e.g., liposomes, mycels, etc.), nanoparticles or nanotubes, or cationic compounds (e.g., polyethyleneimine or PEI). In some cases, bolistic methods such as gene guns or bioistic particle delivery systems can be used to deliver mRNA to cells.
[0243] The carrier may include, but is not limited to, cell-permeable peptides (CPPs), nanoparticles, or polymers.
[0244] The aforementioned CPP is a short peptide that facilitates the cellular absorption of various molecular cargoes (from nano-sized particles to small chemical molecules and large fragments of DNA).
[0245] With respect to the aforementioned nanoparticles, the compositions according to the present invention can be transmitted via polymer nanoparticles, metal nanoparticles, metal / inorganic nanoparticles, or lipid nanoparticles. The polymer nanoparticles may be, for example, DNA nanocrews synthesized by rolling circle amplification, or filamentous DNA nanoparticles. mRNA can be loaded onto the DNA nanocrews or filamentous DNA nanoparticles, and they can be coated with PEI to improve their endosomal escape ability. These complexes can be made to bind to the cell membrane, internalize, and then translocate to the nucleus via endosomal escape for transmission.
[0246] With respect to the metal nanoparticles, gold particles can be linked together and delivered to cells by forming a complex with a cationic endosomal disruptive polymer. The cationic endosomal disruptive polymer may be, for example, polyethyleneimine, poly(aladdin), poly(lysine), poly(histidine), poly-[2-{(2-aminoethyl)amino}-ethyl-aspartamide] (pAsp(DET)), a block copolymer of poly(ethylene glycol) (PEG) and poly(arginine), a block copolymer of PEG and poly(lysine), or a block copolymer of PEG and poly{N-[N-(2-aminoethyl)-2-aminoethyl]aspartamide} (PEG-pAsp DET)).
[0247] Regarding the aforementioned metal / inorganic nanoparticles, for example, mRNA can be encapsulated via ZIF-8 (zeolitic imidazolate framework-8).
[0248] In some cases, the mRNA can bind to anion-containing and cation-containing substances to form nanoparticles, which can then penetrate cells through receptor-mediated endocytosis or pagocytosis.
[0249] Cationic polymers include polyallylamine (PAH), polyethyleneimine (PEI), poly(L-lysine) (PLL), poly(L-arginine) (PLA), polyvinylamine monopolymers or copolymers, poly(vinylbenzyl-tri-C1-C4-alkylammonium salts), aliphatic or aromatic aliphatic dihalides and polymers of aliphatic N,N,N',N'-tetra-C1-C4-alkyl-alkylenediamines, poly(vinylpyridine) or poly(vinylpyridinium salts), poly(N,N-diallyl-N,N-di-C1-C4-alkyl-ammonium halide), quaternary di-C1-C4-alkyl-aminoethyl acrylate or methacrylate monopolymers or copolymers, and POLYQUAD. TM This may include polyaminoamides, etc.
[0250] The material may include a cationic liposome formulation as the cationic lipid. The lipid bilayer of the liposome protects the encapsulated nucleic acid from degradation and prevents specific neutralization by antibodies that can bind to the nucleic acid. During endosomal maturation, the endosomal membrain fuses with the liposome, enabling efficient endosomal escape of cationic lipid-nucleic acid cleavage enzymes. Cationic lipids may include polyethyleneimine, polyamidoamine (PAMAM)-derived dendritic structures, lipofectin (a combination of DOTMA and DOPE), lipofectase, lipofectamine (registered trademark) (e.g., lipofectamine (registered trademark) 2000, lipofectamine (registered trademark) 3000, lipofectamine (registered trademark) RNAiMAX, lipofectamine (registered trademark) LTX), Saint-Red (Synvolux Therapeutics, located in Groningen, Netherlands), DOPE, cytopectin (Gilead Sciences, located in Poster City, California), and upectin (JBL, located in San Luis Obispo, California). Typical cationic liposomes can be produced from N-[1-(2,3-dioreooxy)-propyl]-N,N,N-trimethylammonium chloride (DOTMA), N-[1-(2,3-dioreooxy)-propyl]-N,N,N-trimethylammonium methyl sulfate (DOTAP), 3β-[N-(N',N'-dimethylaminoethane)carbamoyl]cholesterol (DC-hol), 2,3-dioleyloxy-N-[2(sperminecarboxamide)ethyl]-N,N-dimethyl-1-propanamide trifluoroacetate (DOSPA), 1,2-dimyristyloxypropyl-3-dimethyl-hydroxyethylammonium bromide, or dimethyldioctadecylammonium bromide (DDAB).
[0251] Lipid nanoparticles can be transported to a carrier using liposomes. Liposomes are spherical vesicle structures consisting of a single or multiple lamellar lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposome dosage forms may mainly contain natural phospholipids and lipids, such as 1,2-distearolyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, phosphatidylcholine, or monosialogangliosides. In some cases, cholesterol or 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) may be added to the lipid membrane to eliminate instability in plasma. The addition of cholesterol reduces the rapid release of encapsulated viable compounds into plasma, while 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) increases stability.
[0252] 7. Base editing In other aspects, the present invention relates to a base editing composition comprising the fusion protein or the nucleic acid. In other words, the present invention relates to a base editing method comprising the step of treating cells with the composition. After a DNA-binding protein, such as TALE or ZFP, binds to target DNA, the cytosine deaminosease in the fusion protein hydrolyzes the amino group of cytosine, transforming it into uracil. Since uracil can base-pair with adenine, the cytosine-guanine base pair can be edited through the cellular DNA replication process to a uracil-adenine base pair and finally to a thymine-adenine base pair. Similarly, the adenine deaminosease in the fusion protein hydrolyzes the amino group of adenine, converting it into hypoxanthine. Since hypoxanthine can base-pair with cytosine, the adenine-thymine base pair can similarly be edited through the cellular DNA replication process to a hypoxanthine-cytosine base pair and finally to a guanine-cytosine base pair.
[0253] The cells may be, but are not limited to, eukaryotic cells (e.g., fungi such as yeast, cells derived from eukaryotes and / or eukaryotic plants (e.g., embryonic cells, stem cells, somatic cells, germ cells, etc.)), eukaryotes (e.g., humans, primates such as monkeys, dogs, pigs, cattle, sheep, goats, mice, rats, etc.), or eukaryotic plants (e.g., algae such as green algae, corn, soybeans, wheat, rice, etc.).
[0254] (1) Base editing of plant cell DNA The present invention relates to a composition or method for base editing of plant cell DNA. The present invention relates to a composition for base editing of plant cells comprising the fusion protein or nucleic acid encoding the same; and an NLS peptide, a chloroplast transit peptide, a mitochondrial targeting signal (MTS), a nuclear export signal protein or nucleic acid encoding the same.
[0255] The present invention also provides a composition for base editing of plant cells comprising the fusion protein or nucleic acid; and an NLS peptide or nucleic acid encoding it.
[0256] The present invention also provides a composition for base editing of plant cells comprising the fusion protein or nucleic acid; and a chloroplast signaling peptide or nucleic acid encoding it.
[0257] The present invention also provides a composition for base editing of plant cells comprising the fusion protein or nucleic acid; and a mitochondrial signaling protein or nucleic acid encoding it.
[0258] In some cases, the present invention provides a composition for base editing of plant cells, further comprising a nuclear export signaling protein or a nucleic acid encoding the same.
[0259] Specifically, this relates to compositions or methods for editing plant cell nuclear DNA, mitochondrial DNA, or chloroplast DNA.
[0260] Specifically, the fusion protein can be transmitted to plant cells via the following means. Injection using a Gene Gun (Bombardment); PEG-mediated protoplast transfection; Protoplast transfection (electroporation); or Protoplast injection (microinjection) using microinjectors.
[0261] The polynucleotide sequence encoding the fusion protein according to the present invention may be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination sequence).
[0262] The polynucleotide encoding the aforementioned fusion protein can be transmitted to plant cells via the following means.
[0263] Transformation using Agrobacterium, such as Agrobacterium tumefaciens. -Binary vector, - Viral vector: Geminivirus Examples include tobacco rattle virus (TRV), tomato mosaic virus (ToMV), foxtail mosaic virus (FoMV), barley yellow striate mosaic virus (BYSMV), and Sonchus yellow net rhabdovirus (SYNV); Virus-based transfection; Injection (Bombardment, Gene gun); PEG-mediated protoplast transfection; Protoplast transfection (electroporation); or Protoplast injection (microinjection) using microinjectors.
[0264] Examples of the aforementioned viruses include viral vectors such as Geminivirus, tobacco rattle virus (TRV), tomato mosaic virus (ToMV), foxtail mosaic virus (FoMV), barley yellow striate mosaic virus (BYSMV), and Sonchus yellow net rhabdovirus (SYNV).
[0265] The aforementioned vectors can be delivered into cells via local injection (e.g., direct injection into the lesion or target site), electroporation, lipofection, viral vectors, nanoparticles, as well as PTD (Protein translocation domain) fusion protein methods.
[0266] With respect to the protein or nucleic acid encoding it that is transported to the plant organ, the plant organ may be a mitochondria, chloroplast, or plastid (white or chromoid).
[0267] The proteins transported to the plant organs may be, for example, chloroplast signaling peptides or mitochondrial signaling proteins.
[0268] For example, chloroplast signaling proteins (CTPs) or mitochondrial signaling proteins (MTSs) bind to and transmit signals to chloroplasts and mitochondria within plant cells. When transmitted to chloroplasts and mitochondria, the remaining portion, excluding the N-terminal CTP or MTS protein, is transmitted into the interior of the chloroplasts and mitochondria in a preprotein form. During the process of entering the interior of the chloroplasts and mitochondria, the signaling protein portion detaches, allowing for targeted base editing of specific regions within the chloroplasts and mitochondria.
[0269] In addition to the aforementioned fusion protein or the nucleic acid encoding it, plant cells can be treated or transmitted with chloroplast signaling proteins (CTPs) or nucleic acids encoding them, or mitochondrial signaling proteins (MTSs) or nucleic acids encoding them, to edit the bases of plant mitochondria, chloroplasts, chromoplasm, or leucoplast DNA.
[0270] When a nuclear export sequence is attached to a base-editing protein during mitochondrial gene editing, base editing can be performed with greater efficiency. The nuclear export signal protein may, for example, be derived from MVM (Mirute virus of mice), but is not limited thereto. The nuclear export signal protein may, for example, contain the amino acid sequence of SEQ ID NO: 31, but is not limited thereto.
[0271] VDEMTKKFGTLTIHDTEK (Sequence ID 31) The present invention further includes a TAL (Transcription Activator-Like) effector (TALE) domain that cleaves wild-type DNA base sequences but not edited base sequences, and a FokI nuclease or nucleic acid encoding it, or a ZFN or nucleic acid encoding it, i.e., a mitochondrial TALE nuclease, or nucleic acid encoding it, or a ZFN or nucleic acid encoding it. This allows for more efficient mitochondrial base editing even when mitochondrial sequence cleavage proteins are used simultaneously.
[0272] (2) Base editing of animal cell DNA The present invention relates to a composition or method for base editing of animal cell DNA. The present invention relates to a composition for base editing of animal cells comprising the fusion protein or nucleic acid encoding the same; and an NLS peptide, a mitochondrial signaling protein (MTS), a nuclear export signaling protein or nucleic acid encoding the same.
[0273] The present invention also provides a composition for base editing of animal cells, comprising the fusion protein or nucleic acid; and an NLS peptide or nucleic acid encoding it.
[0274] The present invention also provides a composition for base editing of animal cells, comprising the fusion protein or nucleic acid; and a mitochondrial signaling protein (MTS) or nucleic acid encoding it. The present invention further provides a composition for base editing of animal cells, which optionally further comprises a protein or nucleic acid encoding a nuclear export signal.
[0275] The animal cells are non-human animal cells, and the bases of the mitochondrial DNA of the non-human animal cells can be edited by processing them with nuclear export signaling proteins or nucleic acids encoding them; and / or mitochondrial signaling proteins (MTS) or nucleic acids encoding them.
[0276] For example, a mitochondrial signaling protein (MTS) binds to the fusion protein or the nucleic acid encoding it, and the protein is then transmitted to the mitochondria. When transmitted to the mitochondria, the remaining portion, excluding the N-terminal MTS protein, is transmitted into the mitochondria as a whole protein. During the process of entering the mitochondria, the signaling protein portion detaches, allowing for targeted base editing of specific regions within the mitochondria.
[0277] Mitochondrial signaling pathways, TAL effectors, and cytosine deaminases (DddA) tox The present invention relates to a composition or editing method for base editing of mitochondrial DNA in non-human animal cells, to which a TALE-DdCBE (TALE DddA-derived cytosine base editor) or a nuclear export signal (NES) to which the nucleic acid encoding it is bound. The nuclear export signal can reduce nuclear DNA base editing against mitochondrial-nuclear analogous sequences.
[0278] According to the present invention, a nuclear export signal protein or the nucleic acid encoding it can be used to achieve more efficient animal mitochondrial DNA base editing. Furthermore, the nuclear export signal can reduce nuclear DNA base editing of mitochondrial-nuclear mitotic sequences, so that only mitochondrial DNA is edited.
[0279] The nuclear export signaling protein may, for example, be derived from MVM (Mirute virus of mice), but is not limited thereto. The nuclear export signaling protein may, for example, include the amino acid sequence VDEMTKKFGTLTIHDTEK (SEQ ID NO: 31), but is not limited thereto.
[0280] The following can be injected into eukaryotic cells: (1) a nuclear export signaling protein or nucleic acid encoding it, and (2) a TAL effector (TALE) domain that cleaves wild-type DNA base sequences but not edited base sequences, simultaneously with or sequentially before editing a DNA-binding protein, deaminose, or variant thereof, or nucleic acid encoding it, and FokI nucleohydrolase or nucleic acid encoding it, or ZFN or nucleic acid encoding it.
[0281] In particular, with respect to base editing of eukaryotic mitochondrial genes, the invention may include nuclear export signaling proteins or nucleic acids encoding them and / or mitochondrial signaling proteins (MTS) or nucleic acids encoding them.
[0282] According to the present invention, when a nuclear export sequence is attached to a base-editing protein during animal mitochondrial gene editing, base editing is performed with higher efficiency, and in animal embryos, nonspecific base editing of nuclear similar sequences is suppressed in particular.
[0283] The present invention further includes mitoTALEN (Mitochondrial TALE Nuclease), a mitochondrial nucleic acid hydrolase, or the nucleic acid encoding it, which allows for more efficient mitochondrial base editing even when a mitochondrial sequence cleavage protein is used simultaneously. MitoTALEN, a mitochondrial DNA nucleic acid hydrolase, can be used to cleave mitochondrial DNA, allowing for highly efficient cleavage of wild-type mitochondrial genomes and the acquisition of base-edited genomes in animals.
[0284] The present invention further includes mitoTALEN, a mitochondrial nucleic acid hydrolase, or a nucleic acid encoding it, which allows for more efficient mitochondrial base editing even when a mitochondrial sequence cleavage protein is used simultaneously. Specifically, it may include a TAL effector domain to which a mitochondrial signaling signal is bound, and a fusion protein (mitoTALEN) having a FokI nuclease or ZFN or a nucleic acid encoding it, or a nucleic acid encoding it.
[0285] Using mitoTALEN, a mitochondrial DNA nucleic acid hydrolase, mitochondrial DNA can be cleaved, allowing for highly efficient acquisition of base-edited genomes from animals by cleaving wild-type mitochondrial genomes.
[0286] In some cases, the formula may also contain a UGI (uracil DNA glycosylase inhibitor) to enhance base editing efficiency. UGIs inhibit the activity of UDG (uracil DNA glycosylase), an enzyme that repairs mutated DNA and catalyzes the removal of u from DNA, thereby increasing base editing efficiency.
[0287] Specifically, DddA-derived cytosine base editors (DdCBEs), consisting of the fragmented bacterial toxin DddAtox, a TALE designed to bind to DNA, and a uracil glycosylation inhibitor (UGI), enabled targeted cytosine-thymine base alterations in mitochondrial DNA. Examples demonstrate that highly efficient mitochondrial DNA editing is possible in mouse embryos. The target was MT-ND5 (ND5), which encodes a subunit of NADH dehydrogenase that catalyzes NADH dehydration and electron transfer to ubiquinone in mitochondrial genes. This includes mutations associated with human mitochondrial diseases, such as m.G12918A, and mutations that produce early stop codons, such as m.C12336T. This allows for the creation of mitochondrial disease models in mice, suggesting the potential for treating mitochondrial diseases.
[0288] (1) A nuclear export signaling protein or the nucleic acid encoding it can be linked to (2) a DNA-binding protein, a deaminoase or a variant thereof, or a nucleic acid encoding it, and (3) a nucleohydrolase, mitoTALEN, or a nucleic acid encoding it can be linked to (2). To transmit (1) to (3), one or more delivery means can be used in the same or different configurations.
[0289] The above (1) may be included in the first transmission means, (2) in the second transmission means, and (3) in the third transmission means. Each delivery system may simultaneously be a virus transmission means, one may be a virus transmission means and the other a non-virus transmission means, or both may be a non-virus delivery means.
[0290] The nuclear export signaling proteins (1) to (3) described above, DdCBE, and mitoTALEN can be mixed and transmitted.
[0291] One or more of the above (1) to (3) can be transmitted to the nuclear export signaling protein, DdCBE, and mitoTALEN, and some can be transmitted in such a way that the DNA sequences encoding (1) to (3) are located on the vector.
[0292] The DNA sequences encoding (1) to (3) above can be located on the same vector and transmitted simultaneously via a single vector, or they can be located on separate, different vectors and transmitted simultaneously.
[0293] The animals according to the present invention may include humans or non-human animals. The non-human transformed animals may be insects, annelids, mollusks, brachiopods, nematodes, coelenterates, sponges, chordates, or vertebrates, and the vertebrates may be fish, amphibians, reptiles, birds, or mammals, the insects may be fruit flies (Drosophila), the nematodes may be Caenorhabditis elegans (C. elegans), the fish may be zebrafish, and the mammals may be primates, carnivores, insectivores, rodents, artiodactyls, odd-toed ungulates, or proboscideans, and the rodents may be rats or mice.
[0294] The composition according to the present invention can be introduced into human or non-human animal embryos, and the embryos can be implanted into a surrogate mother to induce pregnancy and produce a base-edited animal. The composition according to the present invention can be introduced into animal fertilized eggs and cultured.
[0295] The fertilized egg obtained can be implanted in a surrogate mother and delivered. The non-human transformed animal may further include a step to confirm whether transformation is possible after delivery. The non-human transformed animal can be mated to produce offspring transformed animals.
[0296] The term "offspring" refers to all offspring that can survive by mating with the non-human transformed animal, and more specifically, it refers to, but is not limited to, the F1 generation produced by mating transformed animals with each other or with normal animals using the transformed animal as a parent, the F2 generation produced by mating F1 generation animals with normal animals, and subsequent generations.
[0297] The aforementioned mating may be characterized by mating with the transformed animal or a normal animal. The present invention may also include cells, tissues, and by-products isolated from the transformed animal or offspring transformed animal. The by-products mean all substances derived from the transformed rabbit, but are preferably selected from the group consisting of blood, serum, urine, feces, saliva, organs, and skin, but are not limited thereto. Modes for carrying out the invention
[0298] The present invention will be described in more detail below with reference to examples. These examples are merely illustrative and it will be obvious to those with ordinary skill in the art that the scope of the present invention is not limited by these examples.
[0299] Example 1. Zinc finger deaminase (ZFD) Base editing of nuclear or mitochondrial DNA has broad applications in medical and life science research, medicine, and biotechnology. The ZFD platform includes a DNA-binding protein, the interbacterial toxin deaminase DddAtox, and a uracil glycosylase inhibitor. UGI catalyzes targeted C-to-T base conversion in human cells without inducing unwanted small insertions and deletions. Plasmids encoding ZFD were constructed using publicly available zinc finger resources, achieving base editing at frequencies of up to 60% in nuclear DNA and 30% in mitochondrial DNA. Unlike CRISPR-based base editing, ZFD does not cleave DNA to create single-strand or double-strand breaks, thus avoiding unwanted insertions and deletions due to error-prone non-homologous end joining at the target site. Furthermore, recombinant ZFD proteins purified in E. coli spontaneously pass through human cells and convert target bases to other bases (base conversion). This provides proof-of-principle evidence for gene therapy that does not involve genes. Technologies for genome editing in eukaryotic cells and organisms include, but are not limited to, zinc finger nucleases (ZFNs), transcription activator-like effector (TALE) nucleases (TALENs), TALE-linked split interbacterial deaminase toxin DddA-derived cytosine base editors (aka DdCBEs), CRISPR-Cas9, and Cas9-linked deaminases (aka base editors) that have lost their cleaving activity. These tools, in principle, consist of two functional units: a DNA-binding portion and a catalytic portion. Thus, a zinc finger array or TALE array functions as the DNA-binding portion, while a nuclease (FokI in the case of ZFNs and TALENs) or deaminosease (split DddA in DdCBEs) functions as the DNA-binding portion. toxAPOBEC1 in CBE functions as a catalytic unit. Crispr-Cas9 possesses both nuclease function and RNA-guided DNA binding protein function. Customized, programmable nucleases such as ZFNs, TALENs, and Cas9 cleave DNA to induce double-strand breaks. While repairing these, they cause targeted gene knockout and knock-in. However, double-strand breaks induced by programmable nucleases can lead to undesirable large gene deletions at the target site, p53 activation, and chromosomal rearrangement during simultaneous DSB repair at both the target and non-target sites. In contrast, programmable base editors, including cytosine and adenine base editors (CBEs and ABEs), do not induce DSBs, thus avoiding such undesirable phenomena in cells, and can efficiently catalyze single nucleotide conversations without a repair template or donor DNA. However, because CBE and ABE contain Cas9 nickase variants, they still cause nicks or single-strand breaks by cleaving one strand of DNA, resulting in unwanted insertions and deletions at the target gene site.
[0300] This can catalyze C-to-T base substitutions in nuclear and mitochondrial DNA within cells. We demonstrated mitochondrial DNA editing in mice and chloroplast DNA editing in plants using a custom-designed DdCBE. Zinc finger deaminoses (ZFDs) that perform precise base editing without deletions in eukaryotic cells, distinct from human cells, were constructed by ligating fragmented DddAtox to a customized zinc finger protein. Zinc finger arrays (2x0.3~0.6k base pairs) are smaller in size compared to TALE arrays (2x1.7~2k base pairs) and S. pyogenes Cas9 (4.1 k base pair), allowing ZFD-encoded genes to be easily loaded into viral vectors with limited cargo space, such as AAV, for in vivo research and gene therapy applications. Unlike TALE arrays, zinc finger arrays are engineering-friendly because they do not have large volume domains at the C-terminus or N-terminus. The divided halves of DddAtox can fuse to the C-terminal or N-terminal portion of a zinc finger protein. Furthermore, zinc finger proteins possess unique cell-penetrating capabilities, enabling nucleic acid-free gene editing in human cells. These properties make zinc finger proteins an ideal platform for editing bases in the nucleus or other organelles, acting as DNA-binding modules.
[0301] 1-1. Materials and Methods Plasmid production For expression in mammals, the p3s-ZFD plasmid was prepared by cleaving the p3s-ABE7.10 plasmid (addgene, #113128) with HindIII and XhoI (NEB) enzymes, followed by deformation. The enzymatically cleaved p3s plasmid and the synthesized insert DNA were assembled using the HiFi DNA Assembly Kit (NEB). All insert DNA encoding MTS-, ZFP (Toolgen, Sangamo and Barbas modules-, split DddA-, or UGI) was synthesized by IDT. The pTarget plasmid was designed to determine the optimal length of the spacer sequence for ZFD activity. Each pTarget plasmid containing spacers of varying lengths and two ZFP binding sites was constructed by inserting the ZFP binding sequence and spacer sequence into a pRGS-CCR5-NHEJ reporter plasmid that had been degraded with two enzymes (EcoRI and BamHI, NEB). The pET-ZFD plasmid for protein production in E. coli was the pET-Hisx6-rAPOBEC1-XTEN-nCas9-UGI-NLS plasmid (addgene, #89508) was generated by cleaving with NcoI and XhoI (NEB) and then deforming it. The ZFD sequence was amplified from the p3s-ZFD plasmid using PCR, and the Hisx6 tag and GST tag sequences were synthesized using oligonucleotides (Macrogen). All plasmids for protein purification were prepared by inserting the sequences encoding the ZFD and tags into enzymatically cleaved pET plasmids using the HiFi DNA Assembly Kit (NEB). Plasmid transformation was performed on chemically acceptable DH5a E. coli cells, and the plasmids were purified using the AccuPrep Plasmid Mini Extraction Kit (Bioneer) according to the manufacturer's protocol. After confirming the entire sequence by Sanger sequencing, the desired plasmids were selected.
[0302] HEK293T cell culture and transfection HEK293T cells (ATCC CRL-11268) were cultured in Dulbecco's Modified Eagle Medium (Welgene) supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antifungal solution (Welgene). HEK293T cells (7.5 × 10⁶) 4 The cells were seeded in a 48-well plate. After 18-24 hours, when the cells had grown to about 70-80%, the left and right ZFD encoding plasmids (500 ng each) were transfected using Lipofectamine 2000 (1.5 μL, Invitrogen) or transfected together with the pTarget plasmid (10 ng). After harvesting the cells for 96 hours following transfection, cell lysis buffer (50 mM Tris-HCl, pH 8.0 (Sigma-Aldrich), 1 mM EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)), 5 μL of proteinase K (Qiagen) were added, and the cells were incubated at 55°C for 1 hour and at 95°C for 10 minutes. To perform whole-mtDNA sequencing, HEK293T cells were transfected with plasmids or mRNA encoding ND1 or ND2 target mitoZFD pairs at sequentially diluted concentrations. The amount of structure (ng) was 7.5 × 10⁻⁶. 4 The transfection was successful. mtDNA was extracted from cells 96 hours after transfection.
[0303] K562 cell culture and transfection K562 cells were cultured in RPMI 1640 culture medium supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-grade antifungal solution (Welgene). For ZFD transfer to K562 cells by electroporation, an Amaxa 4D-Nucleofector with program FF-120 (Lonza) was used. TM The X Unit system was used. 16-well Nucleocuvette TMWhen using strips, the maximum volume of substrate solution added to each sample was 2 μL. For proteins, 220 pmol (at maximum dose) or 110 pmol (half maximum dose) of the left and right ZFDs were transfected into K562 cells (1 x 10⁵). For plasmids, 500 ng of plasmids encoding the left and right ZFDs were transfected into K562 cells (1 x 10⁵). After 96 hours post-treatment, cells were collected by centrifugation at 100 g for 5 minutes and incubated with 100 μL of cell lysis buffer (50 mM Tris-HCl, pH 8.0 (Sigma-Aldrich), 1 M EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) and 5 μL of proteinase K (Qiagen) for 1 hour at 55°C and 10 minutes at 95°C. Methods previously used for direct transfer of ZFNs were referenced to directly transfer ZFDs or ZFD-encoding plasmids into K562 cells. A mixture of left and right ZFD proteins (final concentration 50 μM) or a mixture of left and right ZFD-encoding plasmids (500 ng each) was diluted in serum-free medium at pH 7.4 containing 100 mM L-arginine and 90 μM ZnCl2 to a final volume of 20 μL. K562 cells (1 × 10⁵) were centrifuged at 100 g for 5 minutes, and the supernatant was removed. The cells were then released into the diluted ZFD solution and incubated at 37°C for 1 hour. After incubation, the cells were centrifuged at 100 g for 5 minutes, and fresh culture medium was added. The cells were maintained at 30°C (temporary hypothermia) or 37°C for 18 hours, and then grown at 37°C for 2 days. Some cells were treated twice according to the above procedure. The cells were analyzed 96 hours after treatment.
[0304] ZFD protein expression and purification Rosetta(DE3) competent cells were transformed with plasmids encoding each pair of ZFDs, each possessing a C-terminal GST tag, and then cultured on LB agar plates containing kanamycin. After overnight culture, single colonies were selected and cultured overnight (pre-culture) in liquid medium containing 50 μg / ml kanamycin and 10.0 μM ZnCl2 at 37°C. The following day, a portion of the pre-culture was transferred to a large volume of liquid medium and cultured at 37°C until the absorbance A at 600 nm was ~0.5~0.70. The culture was placed on ice for about 1 hour, then 0.5 mM isopropyl β-D-1-thiogalactopyranoside (IPTG; GoldBio) was added to induce ZFD protein expression, and the culture was incubated at 18°C for 14 hours.
[0305] For the protein purification process, cells were lysed in a lysis bath (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 1 mM MgCl2 (Sigma-Aldrich), 10 mM 1,4-dithiothreitol (DTT; GoldBio), 1% Triton X-10 Sigma-Aldrich), 10% glycerol, 1 mM phenylmethylsulfonyl fluoride (Sigma-Aldrich), 1 mg / ml lysozyme extracted from chicken egg white (Sigma-Aldrich), 100 μM ZnCl2 (Sigma-Aldrich), 100 mM Sigma-Aldrich), pH 8.0). Further cell lysis was performed using sonication (3 min total, 5 s on, 10 s off). After these steps, the solution was centrifuged (13,000 rpm) and only the supernatant was extracted. The supernatant was incubated for 1 hour with Glutathione Sepharose 4B (GE Healthcare). After incubation, the resin-lysate mixture was poured onto the column and washed three times with washing buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 10 mM DTT (GoldBio), 1 mM MgCl2 (Sigma-Aldrich), 100 μM ZnCl2 (Sigma-Aldrich), 10% glycerol, 100 mM arginine (Sigma-Aldrich), pH 8.0). Proteins adhering to the resin were separated from the resin using elution buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 1 mM MgCl2 (Sigma-Aldrich), 40 mM glutathione (Sigma-Aldrich), 10% glycerol, 1 mM DTT (GoldBio), 100 μM ZnCl2 (Sigma-Aldrich), 100 mM arginine (Sigma-Aldrich), pH 8.0).Finally, the extracted proteins were concentrated to a concentration of ~15 ng / μL (200-240 pmol / μL, depending on the size of the protein).
[0306] In vitro deamination of PCR amplicons using ZFD Amplicons containing the TRAC region were prepared using PCR. 8 μg of amplicon was incubated with 2 μg each of the ZFD proteins (Left-G1397N and Right-G1397C) in NEB3.1 buffer containing 10.0 μM ZnCl2 at 37°C for 1-2 hours. After the reaction, 4 μL of proteinase K solution (Qiagen) was added, and the mixture was incubated at 55°C for 30 minutes to remove the ZFDs. The amplicon was purified using a PCR purification kit (MGmed). 1 μg of the purified amplicon was incubated with two USER enzymes (NEB) at 37°C for 1 hour. The amplicon was then incubated with 4 μL of proteinase K solution and purified again using a PCR purification kit. The purified PCR agarose gels were electrophoresed and imaged.
[0307] Targeted deep sequencing For the analysis of base editing ratios at target and off-target sites, target sites were subjected to superimposed primary PCR, secondary PCR amplification, and tertiary PCR using TruSeq HT Dual index-containing primers to generate deep sequencing libraries using PrimeSTAR® GXL DNA polymerase (TAKARA). The libraries were then paired-end sequenced using Illumina MiniSeq.
[0308] mRNA preparation DNA templates containing a T7 RNA polymerase promoter upstream of the ZFD sequence were PCR-generated using forward and reverse primers (forward: 5'-CATCAATGGGCGTGGATAG-3' SEQ ID No: 116, reverse: 5'-CATCAATGGGCGTGGATAG-3' SEQ ID No: 117, reverse: 5'-GACACCTACTCAGACAATGC-3' SEQ ID No: 118). mRNA was obtained using mMESSAGE mMACHINE. TM The mRNA was synthesized in vitro using the T7 ULTRA Transcription Kit (Thermo Fisher). The in vitro transcribed mRNA was processed according to the manufacturer's protocol using MEGAclear. TM Purification was performed using the Transcription Clean-Up Kit (Thermo Fisher).
[0309] Whole mitochondrial genome sequencing For whole mitochondrial genome sequencing, three steps are required. 1. mtDNA extraction from isolated mitochondria: First, 3 × 10⁵ HEK293T cells were trypsin-treated and transfected with ND1 or ND2 targeted mitoZFD pairs, then collected by centrifugation (500g, 4 min, 4°C) after 96 hours. The cells were then washed with phosphate-buffered saline (Welgene), centrifuged, and collected again. The supernatant was removed, and mitochondria were isolated from cultured cells using the reagent-based method of the Mitochondria Isolation Kit for Cultured Cells (Thermo Fisher) according to the manufacturer's protocol. mtDNA was extracted from the isolated mitochondria using the DNeasy Blood & Tissue Kit (Qiagen). 2. NGS library generation: To generate an NGS library from the extracted mtDNA, Nextera TMThe Illumina DNA Prep kit, which includes DNA CD Indexes (Illumina), was used. 3. NGS: The library was pooled and loaded into a MiniSeq sequencer (Illumina). The average sequencing depth was greater than 50.
[0310] Analysis of whole-genome DNA editing in the mitochondrial genome To analyze the NGS data of the entire mitochondrial genome sequencing, Fastq files were aligned to the GRCh38.p13 (release v102) reference genome using BWA, and BAM files were generated with SAMtools (v.1.9) after correcting the read pairing information and flags. Next, the REDItoolDenovo.py script from REDItools (v.1.2.1) was used to identify the locations of all cytosine and guanine in the mitochondrial genome with a base edit rate of 1% or more. Locations with a base edit rate of 50% or more in the cell line were considered single nucleotide mutations and excluded from all samples. Target sites of each ZFD were excluded for off-target analysis. The remaining locations with an edit frequency of ≥1% were considered off-target, and the number of edited C / G nucleotides was calculated. To calculate the T / A base edit frequency from the mean off-target C / G of the mitochondrial genome, it was averaged over all C / G. The specificity ratio was calculated by dividing the on-target mean by the off-target mean. The graph of the entire mitochondrial genome was generated by displaying the base editing rates at both the target and non-target locations.
[0311] 1-2. Optimization of ZFD structure To develop ZFDs for base editing in human and other eukaryotic cells, we first attempted to optimize the amino acid linker and spacer length of the Zing Finger Protein (ZFP) that attaches to half of the DddAtox. A C-to-T base transposition occurs in the spacer between the attachment sites of the left and right ZFPs. We selected ZFN pairs that target the appropriately characterized human CCR5 gene. Using these, we constructed ZFDs with various amino acid linkers of 2, 5, 10, 16, 24, and 32, and created a series of target plasmids with diverse spacers ranging from 1 to 24 base pairs in length, containing ZFP binding sites on the left and right sides of the ZFD and repeating TC sequences (Figure 1a, b, and Table 1).
[0312] [Table 1]
[0313] DddAtox can be split at two locations (G1333 and G1397), and each half can fuse to either the left or right ZFP. The base editing efficiency of the 24 generated ZFD configurations (= 6 linkers x 2 splitting locations x ZFP fusion locations (left or right)) was measured for each target plasmid with 24 spacers. In Hek293T cells, measurements were taken via deep sequencing on day 4 after transfection.
[0314] [structure] ·Left-ZFD: SV40 NLS-ZFP(S162-left)-linker-DddAtox half-4aa linker-UGI ·Right-ZFD: SV40 NLS-ZFP(S162-right)-linker-DddAtox half-4aa linker-UGI - SV40 NLS: PKKKRKV (SEQ ID No: 478) - ZFP(S162-left) GIHGVPAAMAERPFQCRICMRNFSDRSNLSRHIRTHTGEKPFACDICGRKFAISSNLNSHTKIHTGSQKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFATSGNLTRHTKIHLR (SEQ ID NO: 2) - ZFP(S162-right) GIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR (SEQ ID NO: 3)
[0315] - Linker between zinc finger protein and DddAtox half: - 2aa: GS - 5aa: TGEKP (SEQ ID No: 479) - 10aa: SGAQGSTLDF (SEQ ID No: 9) - 16aa: SGSETPGTSESATPES (SEQ ID No: 10) - 24aa: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID No: 115) - 32aa: GSGGSGGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID No: 11) - Split-DddAtox G1333-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID No: 27) - Split-DddAtox G1333-C PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPV KRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID No: 272)
[0316] - Split-DddAtox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDN GISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID No: 273)
[0317] - Split-DddAtox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID No: 26) - 4aa linker SGGS (SEQ ID NO: 480) - UGI TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA LVIQDSNGENKIKML (SEQ ID NO: 481)
[0318] ZFDs with short linkers (2 and 5 amino acid (AA) linkers) showed low efficiency. On the other hand, ZFDs with linkers of 10AA or more showed a C-to-T base editing efficiency in the range of 1% to 24% in spacers of 4 or more base pairs (Figure 1c and Figures 2a, c). The ZFD pair with a 24AA linker showed the highest editing efficiency. To find the optimal linker combination, we fixed a 24AA linker to the left ZFD of a ZFD pair and combined it with ZFDs of various linker lengths on the right side and checked the base editing efficiency, and similarly tried the reverse (Figure 1d and Figure 3). It was found that having 24AA linkers on both sides was the most efficient. Furthermore, we confirmed that DddAtox showed even better efficiency when split at the G1397 region than when split at the G1333 region (Figure 1c and Figures 2a, b). These most efficient ZFD pairs edit cytosin with high efficiency of >6.7% in 7–21 bp spacers (Figure 1c and Figures 2a, c).
[0319] 1-3. Base editing in in vivo nuclear DNA targets Next, we investigated whether ZFDs with a 24AA linker in human cells could catalyze C-to-T base editing at in vivo chromosomal target sites. Twenty-two pairs of ZFDs were created targeting 11 sites (one site with two ZFD pairs) within a total of eight genes (Figure 4). Of these, 14 pairs of ZFDs were obtained from publicly available zinc finger resources. The other eight pairs of ZFDs were created by modifying previously characterized ZFNs (specifically identified as CCR5 and TRAC). This was done because, while ZFNs cleave target DNA with a spacer length of 5-7 bp, ZFDs function with a spacer of at least 7 bp. Therefore, we created ZFDs that could function by attaching or detaching one or two zinc fingers to such ZFN pairs. Since the FokI nuclease can fuse to either the N-terminal or C-terminal portion of a ZFP, creating ZFNs with four different configurations, we also created two pairs of ZFDs with other configurations (Trac-NC in Figure 4b) to test whether half of the divided DddAtox could fuse not only to the C-terminal portion of a conventional ZFP but also to the N-terminal portion of a ZFP (NC configurations are shown in Figures 4a and 5).
[0320] [structure] ·C type: SV40 NLS-Zinc finger protein-24aa linker-DddA tox half-4aa Linker-UGI • N type: SV40 NLS-DddA tox half-24aa linker-Zinc finger protein-4aa linker-UGI - SV40 NLS PKKKRKV (SEQ ID No: 478) - 24aa linker SGTPHEVGVYTLSGTPHEVGVYTL - Split-DddA tox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDN GISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG - Split-DddA tox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC - 4aa linker SGGS - UGI TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA LVIQDSNGENKIKML - ZFP CCR5-1 Left (C type) [S162 ZFN-Left] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR
[0321] CCR5-1 Right (C type) [S162 ZFN-Right] GIHGVPAAMAERPFQCRICMRNFS RSDNLSV HIRTHTGEKPFACDICGRKFA QKINLQV HTKIHTGEKPFQCRICMRNFS RSDVLSE HIRTHTGEKPFACDICGRKFA QRNHRTT HTKIHLR
[0322] CCR5-2 Left (C type) [S162 ZFN-Left] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR (SEQ ID NO: 2)
[0323] CCR5-2 Right (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIHGVPAAMAERPFQCRICMRNFS QSGDLRR HIRTHTGEKPFACDICGRKFA RSDNLSV HTKIHTGSQKPFQCRICMRNFS QKINLQV HIRTHTGEKPFACDICGRKFA RSDVLSE HTKIHLR (SEQ ID NO: 482)
[0324] TRAC-Left (C type) [Adapted from Paschon, D.E. et al., 2019] GIHGVPAAMAERPFQCRICMRNFS DQSNLRA HIRTHTGEKPFACDICGRKFA TSSNRK THTKIHTGSQKPFQCRICMRNFS LQQTLAD HIRTHTGEKPFACDICGRKFA QSGNLAR HTKIHLR (SEQ ID NO: 483)
[0325] TRAC-Left (N type) [Adapted from Paschon, D.E. et al., 2019] FQCRICMRKFA TSGSLTR HTKIHTGEKPFQCRICMRNFS RSDHLST HIRTHTGEKPFACDICGRKFA TSSNRTK HTKIHTHPRAPIPKPFQCRICMRNFS RSDNLSE HIRTHTGEKPFACDICGRKFA WHSSLRV HTKIHLR (SEQ ID NO: 484)
[0326] TRAC-Right (C type) [From Paschon, D.E. et al., 2019] GIHGVPAAMAERPFQCRICMRNFS RSDHLST HIRTHTGEKPFACDICGRKFA DRSHLAR HTKIHTGSQKPFQCRICMRKFA LKQHLNE HTKIHTGEKPFQCRICMRNFS QSGNLAR HIRTHTGEKPFACDICGRKFA HNSSLKD HTKIHLR (SEQ ID NO: 485)
[0327] MFAP1 Left (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYSCGICGKSFS DSSAKRR HCILHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYKCMECGKAFN RRSHLTR HQRIHTGEKPYECNYCGKTFS VSSTLIR HQRIHLR (SEQ ID NO: 486)
[0328] MFAP1 Right (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS TSG SLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA QSSNLVR HTKIHLR (SEQ ID NO: 487)
[0329] CCDC28B Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS DPGHLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 488)
[0330] CCDC28B Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHTGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHLR (SEQ ID NO: 489)
[0331] KDM4B Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS DCRDLAR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 490)
[0332] KDM4B Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYRCEECGKAFR WPSNLTR HKRIHTGEKPYRCEECGKAFR WPSNLTR HKRIHTGEKPYSCGICGKSFS DSSAKRR HCILHLR (SEQ ID NO: 491)
[0333] NUMBL Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS DCRDLAR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 492)
[0334] NUMBL Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYKCGQCGKFYS QVSHLTR HQKIHTGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHLR (SEQ ID NO: 493)
[0335] INPP5D-1 Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS RSDKLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 494)
[0336] INPP5D-1 Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYTCSDCGKAFR DKSCLNR HRRTHLR (SEQ ID NO: 495)
[0337] INPP5D-2 Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS RSDKLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 496)
[0338] INPP5D-2 Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKH (SEQ ID NO: 497)
[0339] DVL3 Left (C type) [De novo designed using Barbas zinc finger modules] GIHGVPAAMAERPFQCRICMRNFS TSGHLVR HIRTHTGEKPFACDICGRKFA TSGHLVR HTKIHTGEKPFQCRICMRNFS TSGELVR HIRTHTGEKPFACDICGRKFA QSSNLVR HTKIHLR (SEQ ID NO: 498)
[0340] DVL3 Right (C type) [S162 ZFN-left] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR (SEQ ID NO: 499)
[0341] In HEK293T cells, the cyto-to-t base editing efficiency of ZFDs containing NC structures ranged from 1.0% to 60%. On the other hand, insertion-deletion rates were <0.4%, indicating that they rarely occurred (Figures 4b and 6). As can be seen from the plasmid-based experiments, ZFDs targeting the CCR5 with a 5bp spacer exhibited very low efficiency. For targets with a spacer length of at least 7 bp, the other 20 ZFD pairs showed an average editing efficiency of 12.0 ± 3.4%, which is comparable to the average 8.3 ± 2.2% shown by Cas9-based base editing genes (Cas9-derived Base editor 2). Furthermore, cyto-to-t base editing was observed not only in the TC conxtext but also in the A C yaGC C But it happened (Figures 4c-4f). A C This is NUMBL's C6, GC C At C7 of INPP5D-2, the C base editing efficiencies were 4.58% and 1.85%, respectively.
[0342] JPEG2026062789000012.jpg249170
[0343] 1-4. Directly deliver purified ZFD protein to human cells. Instead of transmitting plasmid DNA encoding gene-editing proteins, transmitting the purified gene-editing proteins themselves to cells reduces off-target effects, avoids innate immune responses caused by external DNA, and prevents the insertion of external plasmid DNA into the in vivo genome. Other groups have shown that ZFPs can spontaneously enter mammalian cells both in vitro and in vivo. To demonstrate protein-mediated base editing by ZFDs, ZFD pairs targeting the highly efficient TRAC gene were selected, and recombinant ZFD proteins with one or four NLSs were purified from E. coli. First, the base-editing efficiency of ZFD proteins was experimentally tested in vitro using a PCR amplicon with a TRAC site, confirming very high efficiency. Efficiency was confirmed by gene cleavage through a uracil-specific cleavage reagent (USER), a mixture of uracil DNA glycosylase and DNA glycosylase-lyase EndoNuclease VIII (Figure 7). The TRAC-NC ZFD protein was transmitted to human leukemia cells K562, a cell type that is difficult to transfect, by two methods: electroporation and direct delivery without electroporation. The ZFD protein was highly efficient. Base editing from C to T showed 26.5% (electroporation) and 17% (direct delivery) (Figure 4g). In summary, these results indicate that plasmids encoding ZFD or purified recombinant ZFD proteins can be used to base edit nuclear DNA in human cells.
[0344] 1-5. Mitochondrial DNA base editing with ZFDs Unlike CRISPR-based systems, a major advantage of systems fused with DddAtox split into custom-designed DNA-binding proteins is that such programmable base-editing scissors can be used to edit organelle DNA, such as mitochondrial DNA. To deliver ZFDs to mitochondria, mitoZFDs were constructed by ligating MTS and NES to the N-terminuses of nine ZFDs designed to target mitochondrial genes (Figure 8). The ZFP portions of the ZFDs were obtained from publicly available Zingfinger resources. The ZFDs were constructed with spacer lengths of 7–15 bp, and the DNA-binding sites were 12 bp on both the left and right sides.
[0345] [structure] ·C type: MTS-FLAG tag-NES-Zinc finger protein-24aa linker-DddA tox half-4aa Linker-UGI • N type: MTS-HA tag-NES-DddA tox half-24aa linker-Zinc finger protein-4aa linker-UGI - MTS (Mitochondrial Targeting Sequence of human mitochondrial ATP synthase F1β subunit) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ (SEQ ID No:274) - FLAG tag (C type) DYKDDDDK (SEQ ID No:275) - HA tag (N type) YPYDVPDYA (SEQ ID No:276)
[0346] - NES (Nuclear export signal) VDEMTKKF (Minute virus if mice; MVM NES) - Split-DddA tox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDN GISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG - Split-DddA tox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC - 4aa linker SGGS - UGI TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA LVIQDSNGENKIKML
[0347] - ZFP * The zinc fingers are connected by TGEKP linkers (ZF1-linker-ZF2-linker-ZF3-linker-ZF4). * The ND2-target mitoZFD has a QQ variant. This variant has either an R or K amino acid replaced with Q. It is shown in red. TIFF2026062789000013.tif246170
[0348] JPEG2026062789000014.jpg209170
[0349] In HEK293T cells, the efficiency of mitochondrial DNA base editing by mitoZFD was 2.6% to 30% (average 14±3%) (Figure 9a). MitoZFD with an NC configuration (16±4.5%, n=6) showed higher efficiency as it increased with the presence of a CC configuration (8.3±2.7%, n=3). CMost cytosins with context were base-edited with varying efficiencies (Figure 9b-g). Furthermore, A in the ND2 region CC Cytosins with (C8 and C9) contexts also showed efficiencies of approximately 7.4% and 19.9%, respectively (Figure 9b). This indicates that ZFD-mediated C-to-T base editing is not limited to TC motifs.
[0350] Next, we isolated single-cell derived clonal populations from mitochondrial DNA (mtDNA) mutant cells to demonstrate that mitoZFD is non-toxic and that mutant mtDNA is maintained in the clonal population. Of 30 single-cell derived clonal populations isolated from HEK293T cells treated with ND1-specific mitoZFD, 5 showed base editing efficiencies of ND1 genes ranging from 35% to 98% (Figure 10a). Similarly, of 36 single-cell derived clonal populations isolated from HEK293T cells treated with ND2-specific mitoZFD, 7 showed base editing efficiencies of ND2 genes ranging from 26% to 76% (Figure 10b). All other clonal populations showed low efficiencies of 0.4% to 1.0%, most of which appear to be sequencing errors. Similar efficiencies were observed in cells not treated with ZFD (Figure 10c). These results indicate that mitoZFD does not uniformly induce heteroplasmic mutations in individual cells. In cells treated with ZFD, most were wild-type, while those with heteroplasmic mutations had mutation rates of up to 98%. This change persisted even after clonal expansion (Figures 11 and 12).
[0351] 1-6. mitoZFDs and TALE-based DdCBEs We found that the ND1-specific mitoZFD we created had a different mutation pattern from the TALE-based DdCBE that targets the same gene (Figure 9f-h). In the case of the two mitoZFDs, base editing from C to T occurs in the cytosin at position C5 or C8 (Figure 9f), whereas DdCBE modifies C8, C9, and C 11 The region is subjected to base editing (Figure 9g). As a result, the amino acid changes caused by mitoZFD are completely different from those caused by DdCBE (Figure 9h). The left and right portions to which mitoZFD is attached are divided by a spacer length of 8-bp, whereas in the case of DdCBE, the spacer length is divided by a length of 16-bp. This difference can result in different mutation patterns. These results indicate that mitoZFD and DdCBE can complementarily generate a variety of mutations in mitochondrial DNA.
[0352] To increase the number of mutation patterns, we tested whether ZFD singles and DdCBE singles could be mixed and used as hybrid pairs. Ten hybrid pairs targeting the ND1 gene showed good activity in HEK293T cells with an average base editing efficiency of 17 ± 3.4% (Figure 13). In fact, one hybrid pair (TALE-L / ZFD-R1) showed better efficiency than two DdCBE pairs and ten ZFD pairs targeting the same site, with a maximum base editing efficiency of 41% (Figure 13b). Furthermore, the hybrid pairs showed different mutation patterns from DdCBE and mitoZFD (Figure 13c). Some (e.g., ZFD-L1 / TALE-R and ZFD-L2 / TALE-R) caused C-to-T mutations at only one site without other associated mutations. In contrast, most mitoZFD pairs and DdCBE pairs caused C-to-T changes at various sites within the spacer. These results demonstrate that ZFD / DdCBE mixed pairs can create specific mutation patterns and generate certain mutations that cannot be produced by ZFD pairs or DdCBE pairs alone.
[0353] 1-7. Specificity of mitoZFD as a mitochondrial genome-wide target To confirm that mitoZFD induces off-target effects in mitochondrial DNA, genes were extracted from cells treated with two pairs of mitoZFD targeting the ND1 or ND2 gene, and whole mitochondrial genome sequencing was performed. HEK293T cells were transfected with mRNA or plasmids encoding mitoZFD pairs in varying amounts (5–500 ng). As expected, on-target efficiency was dose-dependent. High concentrations (100, 200, and 500 ng) of mRNA or plasmid showed >30% on-target efficiency, but also generated hundreds of off-target effects exceeding 1% (Figure 14–17). At low concentrations (5 and 10 ng), these off-target effects were almost eliminated, but on-target efficiency was also significantly lower. An intermediate concentration of 50 ng of mRNA proved most appropriate, maintaining high on-target efficiency without inducing hundreds of off-target effects. To eliminate the remaining off-target effects, we attempted to remove nonspecific DNA contact by introducing the R(-5)Q mutation into each zinc finger region. The resulting ZFD mutant (shown as QQ in Figure 18) maintained high on-target activity while exhibiting high specificity with virtually no off-target effects compared to mtDNA from cells not treated with ZFD (Figures 18a and b).
[0354] Base editing is a relatively new technique that allows for the editing of target bases without causing double-strand breaks in DNA and without the need for DNA repair templates. Base editing can be used in cells, animals, and plants to cause base editing from C to T or A to G, allowing for the study of the functional effects of single nucleotide variants (SNPs) and the repair of disease-causing point mutations in therapeutic applications. Two types of base editing techniques have been developed: CRISPR-based adenine and cytosine base editing scissors and DddA-based base editing techniques. CRISPR-based base editing scissors consist of catalytically-impaired Cas9 or Cas12a as the DNA binding unit, and the deaminase is single-strand DNA specific, derived from rat or E. coli. On the other hand, in cases such as DdCBE, a TALEDNA binding array and double-strand DNA specific DddAtox are used.
[0355] Compared to DdCBE, ZFD is smaller in size because the zinc finger protein in ZFD is smaller, while the TALE array in DdCBE is larger. As a result, genes encoding ZFD pairs can be easily loaded into AAV vectors, which have small loading space, whereas genes encoding DdCBE pairs cannot. Furthermore, the simple properties of ZFP make engineering easier. Half a DddAtox can be fused to the N-terminus or C-terminus of a ZFP, allowing it to be designed to function upstream or downstream of the ZFP binding site. In addition, recombinant ZFD proteins spontaneously penetrate human cells without electroporation or lipofection. This enables gene-free gene therapy. ZFD pairs or ZFD / DdCBE mixed pairs can create unique mutation patterns, a feature not obtainable with DdCBE alone. These properties make ZFD a powerful new platform for modeling and treating mitochondrial diseases.
[0356] Example 2. TALE-DdCBE plant chloroplast, mitochondrial gene editing Plant organelles, including mitochondria and chloroplasts, each possess their own genomes, encoding numerous genes essential for respiration and photosynthesis, respectively. Genetic editing of plant organelles, an unmet condition for plant genetics and biotechnology, has been limited by a lack of suitable tools for targeting the DNA of such organelles. To assemble the DddA-derived cytosine base-editing plasmid (DdCBE), a golden gate cloning system was developed consisting of 16 expression plasmids (8 for transmitting the resulting protein to mitochondria and another 8 for transmission to chloroplasts) and 424 TALE subarray plasmids, using the completed DdCBE plasmid to efficiently create enhanced point mutations in mitochondria and chloroplasts. DdCBE base editing induced efficiencies of up to 25% (mitochondrial) and 38% (chloroplast) in potassium of lettuce or rapeseed. To avoid non-targeted mutations that could cause the DdCBE-encoding plasmid, DdCBE messenger RNA was transmitted to lettuce protoplasts, demonstrating base editing in chloroplasts without DNA. Furthermore, by introducing point mutations into the chloroplast 16S rRNA gene, we created lettuce varieties and young plants with up to 99% editing efficiency that were resistant to streptomycin or spectinomycin.
[0357] DdCBEs consist of isolated non-toxic domains derived from the bacterial cytosine diaminase toxic agent DddAtox, along with a TALE array and UGI designed for specific locations. They function as heterodimers, initiating cytosine deamination reactions by substituting cytosine with thymine in the spacer spaces between TALE protein binding sites on target DNA. We used DdCBEs as a result to demonstrate rapid and convenient binding of DdCBE plasmids expressed in mitochondria and chloroplasts, thus proving highly efficient organelle base editing in plants.
[0358] 2-1. Method Production of expression plasmids in plant protoplasts The DdCBE Golden Gate Destination Vector was constructed using the Gibson assembly method. Sequences encoding the TAL N-terminal domain, HA tag, FLAG tag, TAL C-terminal domain, split DddAtox, and UGI were codon-optimized for dicotyledonous plant (Arabidopsis thaliana) expression and synthesized by Integrated DNA Technology. Sequences encoding the CTP from AtinfA and AtRbcS, and the MTS from the ATPase delta subunit and ATPase gamma subunit were amplified from Arabidopsis thaliana cDNA. For plant expression, the mammalian CMV promoter was replaced with the PcUbi promoter and pea3A terminator in the backbone plasmid. For vector construction for in-vitro DdCBE mRNA transcription, a T7 promoter cassette was introduced between the PcUbi promoter and the DdCBE encoding region in the DdCBE Golden Gate Destination Vector.
[0359] The TALE array gene was constructed using a single golden gate assembly. The DdCBE expression plasmid was constructed as a golden gate assembly for BsaI degradation and T4 ligation using 424 TALE array plasmids and a destination vector. Single golden gate cloning was performed as follows: 20 repeats at 37°C and 50°C for 5 minutes each, followed by a final reaction at 50°C for 15 minutes, and then at 80°C for 5 minutes. All plant protoplast transduction vectors were purified using the plasmid plasmamidi flap kit (qiazen). The DNA and amino acid sequences of the vectors are as follows.
[0360] [Table 2]
[0361] The specific amino acid sequences for the DdCBE structure and the TALE repeat are as follows: [Table 3] TIFF2026062789000017.tif247170TIFF2026062789000018.tif246170TIFF20260627890 00019.tif241170TIFF2026062789000020.tif244170TIFF2026062789000021.tif200170
[0362] mRNA in vitro transfer The DdCBE DNA template was prepared by PCR using fusion DNA polymerase (Thermo Fisher). DdCBE mRNA was synthesized and purified using an in vitro mRNA synthesis kit (Enzynomics).
[0363] Protoplast extraction and transduction Lettuce seeds are surface-sterilized in 70% ethanol for 30 seconds and in 0.4% Lux solution for 15 minutes, then rinsed three times with sterile distilled water. The lettuce seeds are sown at 25°C under 16 hours of light and 8 hours of darkness in a medium containing 0.5x MS medium and 2% sucrose. Rapeseed seeds are surface-sterilized in 70% ethanol for 3 minutes and in 1.0% Lux solution for 30 minutes, then rinsed three times with sterile distilled water. Rapeseed seeds are sown at 25°C under 16 hours of light and 8 hours of darkness in a medium containing 1x MS medium and 3% sucrose.
[0364] Protoplast extraction and transduction will be carried out according to previous research. Lettuce leaves at 7 days old and rapeseed leaves at 14 days old will be treated in an enzyme solution under dark conditions with shaking (40 rpm) for 3 hours. The protoplast-enzyme mixture will be washed with the same amount of W5 solution, and intact protoplasts will be obtained from the sucrose solution by centrifuging 80 g for 7 minutes. The protoplasts will be treated in W5 solution at 4°C for 1 hour, followed by transduction using polyethylene glycol.
[0365] Lettuce protoplasts and rapeseed protoplasts are resuspended in MMG solution, then transduced with plasmids or mRNA using PEG, and cultured at room temperature for 20 minutes. The PEG-protoplast mixture is rinsed three times with the same amount of W5 solution, slowly inverting each time, and then cultured for 10 minutes. The protoplasts are then centrifuged at 100g for 5 minutes to form pellets.
[0366] Protoplast culture Lettuce protoplasts were resuspended in lettuce protoplast culture medium (LPCM) after introduction of a plasmid encoding DdCBE. The protoplasts containing the medium were mixed 1:1 with medium containing 2.4% raw melting agarose and immediately placed in a 6-well dish. After the mixture solidified, 1 ml of liquid medium was placed on top of the embedded protoplasts and cultured at 25°C in the dark for 1 week. After the initial culture, the liquid medium on top was replaced with fresh medium weekly, and the embedded protoplasts were cultured for 1 week under 16 hours of weak light and 8 hours of darkness, followed by 2 weeks under 16 hours of light and 8 hours of darkness. Small calluses induced from the protoplasts were cultured in redifferentiation medium for 4 weeks at 25°C under 16 hours of light and 8 hours of darkness. In preparation for base editing efficiency analysis, the protoplasts were cultured in liquid medium without embedding at 25°C in the dark for 1 week. To test for antibiotic resistance, small calluses implanted for one month were cultured for four weeks at 25°C under 16 hours of light and 8 hours of darkness in redifferentiation medium containing 50 mg / L of streptomycin or 50 mg / L of spectinomycin. After four weeks, antibiotic-resistant green calluses or adventitious buds were transferred to fresh redifferentiation medium containing 200 mg / L of streptomycin or 50 mg / L of spectinomycin.
[0367] Rapeseed protoplasts into which the DdCBE-encoding plasmid was introduced were resuspended in rapeseed culture medium. The protoplast-medium mixture was transferred to a 6-well dish and incubated at 25°C in the dark for 2 weeks. After 2 weeks, the protoplasts were incubated for 3 weeks under 16 hours of weak light and 8 hours of darkness. The medium was replaced with very fresh medium.
[0368] DNA and RNA extraction Total DNA or RNA from cells cultured in liquid medium or transformed callus was extracted using a dienizi plant mini-kit or an al-ennizi plant mini-kit. Cultured cells and callus were harvested by centrifugation at 10,000 rpm for 1 minute. cDNA from total RNA was reverse transcribed using RNA to cDNA EcoDry Premix (TAKARA).
[0369] Deep sequencing The target region was amplified using a fusion enzyme and appropriate primers (see subtable 1). Three rounds of PCR (first, nested PCR, second, PCR, third, indexing PCR) were performed to prepare a DNA sequence analysis library. After collecting equal amounts of DNA, sequence analysis was performed using a MiniSeq (Illumina) instrument. Paired indic sequence files were analyzed using Cas-analyzer and the source code of a computer program.
[0370] 2-2.Results To produce chloroplast-targeted DdCBEs (cp-DdCBEs) and mitochondrial-targeted DdCBEs (mt-DdCBEs), a Golden Gate assembly system was developed (Figure 19). The expression plasmids, with their bound proteins, encode chloroplast signaling peptides or mitochondrial targeting sequences, the N- or C-terminal domains of TALE, split DddAtox (G1333N, G1333C, G1397N, and G1397C), and UGIs, which are optimized for expression in dicotyledonous plants and are regulated by a parsley ubiquitin (PcUbi) promoter and a pea3A terminator. The customized TALEDNA binding sequences and DdCBE plasmids can be mixed with the expression vector and six TALE subarray plasmids in an E-tube to form a single subcloning step. The modular TALE subarray plasmid, comprising a total of 424 sequences (6x64 recognizing 3 sequences + 2x16 recognizing 2 sequences + 2x41 recognizing sequences), can produce cp-DdCBE and mt-DdCBE sequences that recognize sequences of 16-20 nucleotide lengths, including a conserved T at the 5' end. Consequently, functionally, DdCBE heterodimers recognize DNA sequences of 32 to 40 bp.
[0371] To determine whether DdCBEs can promote chloroplast base editing, four pairs of cp-DdCBEs plasmids were constructed suitable for the chloroplast 16S rRNA gene encoding the RNA compound of the 30S ribosomal subunit. Each pair was transinjected into lettuce and rapeseed protoplasts, and after 7 days, base editing efficiency was measured by deep sequencing (Figure 20a, b). The most efficient cp-DdCBE pair induced C·G to T·A substitution in the 15 bp spacer region between the two TALE binding sites (Left-G1397-N + Right-G1397-C) with an efficiency of 30% in lettuce protoplasts and 15% in rapeseed protoplasts (Figure 20b). Similar to conventional results in mammalian cells and mice, cp-DdCBEs preferentially substituted cytosine (C9, C13), a 5'-TC motif, with thymine. Interestingly, the 5'-AC motif, cytosine (C7), was replaced with thymine by another cp-DdCBE (Left-G1333-N + Right-G1333-C) with an efficiency of 4.2% in lettuce protoplasts. Furthermore, we investigated the persistence of cp-DdCBE-mediated base editing in lettuce protoplasts over a 14-day culture period (Figure 24). Editing efficiency increased sustainably for up to 10 days and was maintained throughout the culture period.
[0372] In photosystem II, two additional experiments were performed on psbA and psbB, chloroplast genes that encode D1 and CP-47 photosynthetic proteins, respectively (Figures 20c, d, and 25). Of the cp-DdCBEs targeting the Ps gene, the most efficient was (Left-G1397-C + Right-G1397-N), which was able to induce up to 25% of C·G to T·A substitutions in lettuce protoplasts (Figure 20d). This base editing scissors efficiently substituted only the two cytosines (C11, C12) of 5'-TCC with thymine. This suggests that 5'-TCC may first be substituted into 5'-TTC, and then into 5'-TTT. In rapeseed protoplasts, the other combination (Left-G1333-N + Right-G1333-C) showed the highest efficiency at four cytosine positions (C3, C4, C11, C12), reaching a maximum efficiency of 3.5% (C3). While rapeseed genes have 5'-TCC sequences at C3 and C4, lettuce has a 5'-ACC sequence due to relatively small single nucleotide polymorphisms. Therefore, DdCBE is effective in editing two cytosines (C3, C4) in rapeseed genes, but not in lettuce genes. Furthermore, the cp-DdCBE combination targeting the psbB gene in rapeseed protoplasts replaced the two cytosines with TCC sequences with efficiencies ranging from 0.36% to 4.1% (Figure 25). In summary, these results demonstrate that the editing efficiency of DddAtox is determined by the cytosine position and sequence, including the split position (G1333 vs. G1397) and direction (Left-G1333-N vs. Left-G1333-C), and that cp-DdCBE enables efficient base editing of plant chloroplast genes.
[0373] Next, we attempted to achieve mitochondrial DNA base editing in plants using customized mt-DdCBE. To this end, we created plasmids (using the Golden Gate cloning system) encoding mt-DdCBE targeting the atp6 gene in lettuce and rapeseed, and the rps14 gene in rapeseed. After introducing the plasmids into lettuce and rapeseed protoplasts, we measured the base editing efficiency by deep sequencing 7 days after introduction (Figures 20e, f, and 26). The most efficient mt-DdCBE combination (Left-G1397-N + Right-G1397-C in lettuce and Left-G1397-C + Right-G1397-N in rapeseed) altered the atp6 gene target site by 23% in lettuce protoplasts and 23% in rapeseed protoplasts by substituting C·G with T·A (Figure 20). Furthermore, the mt-DdCBE combination induced 11% C·G to T·A substitution in rapeseed protoplasts at the rps14 target site. These results demonstrate that mt-DdCBE is effective in base editing plant mitochondrial DNA.
[0374] To investigate whether DdCBE-induced cpDNA and mtDNA editing is maintained during redifferentiation, lettuce and rapeseed callus were collected from DdCBE-treated protoplasts four weeks after injection (Figure 21a), and the base editing efficiency of each callus was measured using deep sequencing and Sanger sequencing (Figures 21b and 27). Chloroplast or mitochondrial genes to which base editing was induced by DdCBE were measured with efficiencies of up to 38% in 22 out of 26 lettuce callus and up to 25% in 7 out of 14 rapeseed callus (Figure 21c). Furthermore, base editing of the chloroplast psbA gene showed an efficiency of up to 3.9% in lettuce callus (Figure 27). In addition, mitochondrial base editing in rapeseed callus was measured with efficiencies of up to 25% for atp6 and 1.9% for rps14, respectively (Figure 27). These results indicate that plant protoplasts can tolerate DdCBE expression and that DdCBE induces organelle base editing during the redifferentiation process into protoplasts.
[0375] Next, we attempted to demonstrate DNA-free base editing in organelles using in vitro transcribed cp-DdCBE mRNA instead of plasmids. In vitro transcripts encoding cp-DdCBE targeting the 16S rRNA gene in lettuce were injected into protoplasts, and the base editing efficiency at the target site was analyzed (Figure 21a). C-to-T mutations in the protoplasts were measured up to 25% (Figures 21d, 28). As expected, the DdCBE mRNA and DNA sequences disappeared 7 days after introduction into the protoplasts (Figure 29). This method can be said to avoid the potential integration of a portion of plasmid DNA into the host genome.
[0376] With the help of stable organelle editing maintained in callus redifferentiated from protoplasts, we investigated antibiotic resistance to streptomycin and spectinomycin, which irreversibly bind to 16S rRNA and inhibit protein synthesis through editing of the 16S rRNA gene in chloroplast DNA. Several single nucleotide polymorphisms in the 16S rRNA gene are commonly observed to cause streptomycin resistance in prokaryotes and eukaryotes, and in particular, the 16S rRNA C860T mutation (at position C912 in E. coli) causes streptomycin resistance in tobacco. The C860T point mutation in tobacco is identical to that at position C9 in lettuce (Figures 20a, 21b, 22b, 22d). Lettuce callus obtained by redifferentiation from DdCBE-treated protoplasts was transferred to a medium supplemented with streptomycin and spectinomycin. Mock-treated groups turned white when exposed to antibiotics, indicating protoplast dysfunction of the callus. In contrast, DdCBE-treated callus remained green, demonstrating resistance to these antibiotics. The editing efficiency by DdCBE in resistant lettuce and young plants was analyzed. T substitutions at the C9 position, such as the C860T mutation, were observed in up to 98.6% of callus and shoots obtained from drug treatment (Figure 21e, f). Interestingly, T editing at the nearby C13 position showed up to 20% efficiency in the absence of specniomycin, but was not observed in the presence of antibiotics, demonstrating that this mutation was selected by drug treatment. In summary, these results suggest that DdCBE induction of organelle mutations in protoplasts can be maintained after cell division and plant development, and that chloroplast editing can achieve homoplasmy through drug selection.
[0377] Furthermore, off-target effects of TALE diaminase targeting the 16S rRNA site were analyzed in protoplasts, callus, and suits. Five potential off-target sites were selected based on sequence similarity in the target site (50 base pairs bilaterally) (Figure 31) or in chloroplast genes in single-cell-derived, antibiotic-resistant callus and suits (Figure 32). No off-target effects were observed. On the other hand, when a plasmid encoding DdCBE was introduced into protoplasts, off-target changes of TC to TT were induced at three of the five potential off-target sites with low efficiency ranging from 1.2% to 4.1%. When an in vitro transcript (mRNA) was used instead of the plasmid encoding TALE diaminase, the efficiency of off-target effects in protoplasts was significantly reduced (Figure 22). These results suggest that overexpression or long-term plasmid-derived expression of DdCBE may increase off-target mutations, while transient mRNA-derived expression using mRNA is preferable for avoiding off-target base editing.
[0378] In summary, we developed a Golden Gate cloning system using a 424TALE subarray plasmid and 16 expression plasmids to assemble plasmids encoding DdCBE for organelle DNA editing in plants. Customized DdCBEs targeting three chloroplast DNA genes and two mitochondrial DNA genes achieved highly efficient C-to-T substitution in lettuce and rapeseed protoplasts. In particular, organelle editing in plants was maintained during cell division and plant development. Furthermore, mutations in the chloroplast 16S rRNA gene resulted in nearly homoplasmic (99%) antibiotic-resistant lettuce callus and plants. Without antibiotic selection, editing efficiencies of 25% in mitochondria and 38% in chloroplasts were observed. We expect the Golden Gate cloning system to be a valuable resource for plant organelle DNA editing.
[0379] Example 3. TALE-DdCBE Animal DNA Editing DddA-derived cytosine base editors (DdCBEs), including the fragmented bacterial toxin DddAtox and TALE and UGI designed to enable binding to DNA, enabled targeted cytosine-thymine base alterations in mitochondrial DNA. This demonstrated highly efficient mitochondrial DNA editing in mouse embryos. The target mitochondrial gene was MT-ND5 (ND5), which encodes a subunit of NADH dehydrogenase that catalyzes NADH dehydration and electron transfer to ubiquinone. This included mutations associated with human mitochondrial diseases, such as m.G12918A, and mutations generating early stop codons, such as m.C12336T. Through this, a mitochondrial disease model could be created in mice, suggesting the potential for treating mitochondrial diseases.
[0380] 3-1. Method Plasmid assembly. The TALEN (transcription activator-like effector nucleases) system was borrowed to construct an expression plasmid containing the DddA half and the final TALE-DddAtox structure. In the TALEN system expression plasmid, monomers in the nuclear transport signal and FokI dimer were replaced with the mitochondrial signaling pathway (MTS), the DddA deaminationase half, and a uracil glycosylation inhibitor (UGI). Sequences encoding MTS, DddA, and UGI were synthesized using IDT. To construct the expression vector, the DNA fragments required for Gibson assembly were amplified using Q5 DNA polymerase (NEB) and then purified. The purified gene fragments were assembled using the HiFi DNA assembly kit (NEB), chemically transformed into E coli DH5a (Enzynomics), and confirmed by Sanger sequencing. Therefore, eight different expression plasmids were obtained, in which the BsaI restriction enzyme site for golden gate cloning was located between the sequences encoding the N-terminal and C-terminal domains. For the assembly of the DdCBE plasmid, the expression plasmid was combined with a module vector (each encoding a TALE sequence), BsaI-HFv2 (10 U), T4 DNA ligase (200 U), and reaction buffer in a single tube. Subsequently, restriction enzyme and ligase reactions were repeated 20 times in a thermocycler at 37°C for 5 minutes and 50°C for 5 minutes, followed by treatment at 50°C for 15 minutes and 80°C for 5 minutes. The conjugated plasmids were introduced into E coli DH5α via chemical transformation, and the final structure was confirmed by Sanger sequencing. For cell line introduction, the plasmids were midiprep.
[0381] Culture and transfection of mammalian cell lines. The NIH3T3 (CRL-1658, American Type Culture Collection (ATCC)) cell line was cultured at 37°C in a 5% CO2 environment. The cell line was grown without antibiotics in 10% (v / v) elegance serum-enriched DMEM (Gibco) medium, and mycoplasma testing was not performed. For lipofection, cells were placed in 12-well cell culture plates (SPL, Seoul, Korea) in 1.5 × 10⁶ wells. 4 Cells were introduced at a given cell density 18–24 hours before transfection. A total of 1,000 ng of plasmid DNA was introduced using Lipofectamine 3000 (Invitrogen), with 500 ng of each DdCBE cleavage. Cells were harvested 4 days after transfection.
[0382] mRNA preparation. The mRNA template was amplified by PCR using Q5 DNA polymerase (NEB), and the following primers were used (F: 5′-CATCAA TGGGCGTGGATAG-3′SEQ ID No. 268, R: 5′-GACACCTACTCAGACAATGC-3 SEQ ID No. 269). DdCBE mRNA was synthesized using an in vitro RNA transcription kit (mMESSAGE mMACHINE T7 Ultra kit, Ambion) and then purified using the MEGAclear kit (Ambion).
[0383] Animals. All experiments, including those involving rats, were conducted with the approval of the Animal Management and Use Committee of the Institute for Basic Science. Superovulating C57BL / 6J females were mated with C57BL / 6J males, and ICR strain females were used as surrogate mothers. Mice were housed in a facility free of specific pathogens under conditions of constant temperature and humidity, maintaining a 12-hour day-night cycle (20-26°C, 40-60%).
[0384] Microinjection was performed into mouse conjugates. The processes immediately preceding microinjection, including superovulation, embryo collection, and microinjection, proceeded as described in the previous paper. For microinjection, a mixture containing left DdCBE mRNA (300 ng / μl) and right DdCBE mRNA (300 ng / μl) was diluted in DEPC-treated injection buffer (0.25 mM EDTA, 10 mM Tris, pH 7.4) and injected into the conjugate cytoplasm using a Nikon ECLIPSE Ti micromanipulator and a FemtoJet 4i microinjector (Eppendorf). After microinjection, embryos were placed in KSOM + AA (Millipore) microdrops and cultured at 37°C and 5% CO2 for 4 days. Two-cell stage embryos were transplanted into the fallopian tubes of 0.5-dpc pseudo-pregnant surrogate mothers.
[0385] Genotype analysis. Blastocyst-stage embryos and tissues were placed in degradation buffer (25 mM NaOH, 0.2 mM EDTA, pH 10) and cultured at 95°C for 20 minutes. After incubation, the pH was adjusted to 7.4 using HEPES (free acids, no pH adjustment) to achieve a final concentration of 50 mM. For DNA isolation of mouse offspring, DNeasy Blood & Tissue Kits (Qiagen) were used, and analysis was performed by Sanger sequencing and targeted deep sequencing.
[0386] Mitochondrial DNA isolation for high-resolution sequence analysis. To isolate mitochondria from NIH3T3 cells cultured in a 12-well plate, the cell culture medium was removed, and 200 μl of Mitochondrial isolation buffer A (ScienCell) was added to the culture plate. After scraping the cells using a cell lifter, they were placed in a microcentrifuge tube and crushed using a disposable mortar and pestle. After 15 crushings, the well-crushed homogenate was centrifuged at 1,000 × g at 4°C for 5 minutes. The supernatant was placed in a new microcentrifuge tube and centrifuged at 10,000 × g at 4°C for 20 minutes. The precipitate was released into 20 μl of lysate (25 mM NaOH, 0.2 mM EDTA, pH 10) and boiled at 95°C for 20 minutes. To lower the pH, 2 μl of 1 M HEPES (free acids, without pH adjustment) was added to the mitochondrial lysate. 1 μl of the thus prepared solution was used as a PCR template for high-resolution sequence analysis.
[0387] High-processing sequence analysis. To prepare a deep sequencing library, nested primary and secondary PCR were performed using Q5 DNA polymerase, and the final index sequence was added. The library was used for fair-end read sequencing analysis using MiniSeq (Illumina). For whole-mitrogony genome analysis, isolated mitochondrial DNA was prepared using the tagmentation DNA prep kit (Illumina) according to the manufacturer's protocol. In all analyses, fair-end sequence results were combined using a single fastqjoin file and analyzed using CRISPR RGEN Tools (http: / / www.rgenome.net / ).
[0388] Data analysis and display. Microsoft Excel (2019) and PowerPoint (2019) were used to create pictures, graphs, and tables. Geneious (version 2021.0.1) and Snapgene 5.2.3 were used for genome sequence alignment, primer preparation, and cloning design, with NC_005089 used as the reference sequence.
[0389] 3-2.Results Assembly of the DdCBE plasmid. To facilitate the assembly of the DdCBE-customized TALE sequence, an expression plasmid encoding half of the split DddAtox was constructed, and a Golden Gate cloning system was used with a total of 424 sequences (6×64 sequences with 3 recognitions + 2×16 sequences with 2 recognitions + 2×4 sequences with 1 recognition) (Figure 33a). As shown in Table 4 below, six TALE plasmids and the expression plasmid were mixed in the same tube to create a ready-to-use DdCBE plasmid containing 15.5 to 18.5 repeat variable diresidue sequences (Figure 38).
[0390] [Table 4]
[0391] The DdCBE configuration is shown in Table 5 below. As a result, DdCBE recognizes 17 to 20 DNA sequences, including the conserved thymine sequence at the 5' end. Therefore, a functional DdCBE pair recognizes a total of 32 to 40 DNA base sequences.
[0392] [Table 5] JPEG2026062789000024.jpg235161JPEG2026062789000025.jpg80170
[0393] In vitro mitochondrial base editing. To attempt in vivo mitochondrial DNA editing using the Golden Gate cloning system, we selected the ND5 gene encoding the Mus musculus mitochondrial NADH-ubiquinone oxidoreductase chain 5 protein. The ND5 protein is a core subunit of NADH dehydrogenase (ubiquinone) and catalyzes electron transfer from NADH to the respiratory chain. In humans, mutations in the ND5 gene are known to be associated with MELAS (mitochondrial encephalomyopathy, lactic acidosis, and stroke-like episodes) and some symptoms of Leigh syndrome or LHON (Leber's hereditary optic neuropathy). To mimic human dysfunction, we attempted to create a mouse model with genetic mutations in mitochondrial genes.
[0394] First, several types of DdCBE plasmids were assembled, designed to induce two silent mutations, m.C12539T and m.G12542A. These plasmids were transfected into NIH3T3 mouse cell lines, and the base editing frequency was measured after 3 days. As expected, cytosine bases in the target range were edited to thymine with an efficiency of up to 19% (Figure 34a). While DddAtox has been previously reported to exclusively deaminate only cytosines in "TC" sequences, our experimental results also showed that only two cytosines in the TC context were edited. Deletions or other types of point mutations were not meaningfully generated within the editing target range.
[0395] In vivo mitochondrial base editing. The most effective DdCBE pair (left-G1397-N and right-G1397-C) was used in in vitro experiments. Four days after microinjection of an in vitro transcript encoding this DdCBE pair into one-cell stage C57BL6 / J embryos, 9 out of 32 embryos were successfully edited (28%, Table 6).
[0396] [Table 6]
[0397] The TALE-DddAtox deaminationase efficiently generated the C·G to T·A base transfer, with efficiencies ranging from 2.2% to 25% in m.C12539 and from 0.63% to 5.8% in m.G12542. Subsequently, embryos injected with DdCBE were transplanted into surrogate mothers, yielding offspring with m.C12539T and m.G12542T (Figure 39). Three of the four offspring (F0) underwent C·G to T·A editing, ranging from 1% to 27% (Figure 34c). Two offspring showed similar mutation values in the toes and tail, and maintained this efficiency even 14 days after birth. Furthermore, these mitochondrial DNA mutations were detected in various tissues of adult F0 mice 50 days old (Figure 34d). These results suggest that mitochondrial DNA heteroplasms generated by DdCBE within the single-cell complex are maintained during development and differentiation.
[0398] To confirm that the mutations induced by DdCBE are transmitted to the next generation, female F0 mice were crossed with wild-type C57BL6 / J males to produce F1 offspring. The m.C12539T and m.G12542T mutations were observed with efficiencies of 6–26% in two offspring. Furthermore, these mitochondrial edits were observed at similar rates in 11 different tissues (Figure 35b).
[0399] DdCBE-mediated MT-ND5 G12918A mutation. Next, we attempted to create the m.G12918A mutation, which can cause mitochondrial diseases in humans. Notably, this mutation causes several mitochondrial diseases, such as Leigh syndrome, MELAS syndrome, and LHON syndrome. Because the cytosine base at this position is adjacent to a thymine, base editing using DdCBE is possible (Figure 36a). Four pairs of DdCBE were assembled, and it was confirmed that editing up to 6.4% was possible in NIH3T3 (Figure 36b). Subsequently, the most efficient DdCBE combination was microinjected into mouse conjugates, and its efficiency was observed in blastocysts. Eleven out of 44 embryos (25%) had the m.G12918A mutation, with efficiencies ranging from 0.25 to 23% (Figure 36c). Next, embryos microinjected with DdCBE were transplanted into surrogate mothers to obtain offspring with the G12918A mutation (Figure 39b). Of the 11 newborn mice, 4 were found to have this mutation in approximately 3.9–31.6% of the offspring (Figure 36d). Although no phenotypes were observed immediately after birth, these results suggest that DdCBE can be used to create an animal model of mitochondrial disease, likely because the offspring were very young and normal and mutant DNA were present in a heterogeneous state.
[0400] MT-ND5 Nonsense Mutation. Finally, we investigated whether loss-of-function mutations in ND5 could be maintained in mice by creating nonsense mutations in the gene. Using m.C12336 as the target cytosine, we introduced an early stop codon at position 199 of the ND5 protein (Q199*; Figure 37a). First, we transfected NIH3T3 cell lines with four DdCBE combinations to check the base editing efficiency and confirmed that the most effective DdCBE pair induced nonsense mutations with an efficiency of approximately 5.7% (Figure 37b). This DdCBE caused cytosine-thymine editing and confirmed that the mutation (Q200Q) that produces the silent mutation of m.G12341A, albeit with slightly lower efficiency, was also edited within the target range. In 19 out of 37 mouse embryos (=51%), the m.C12336T and m.G12341A mutations were confirmed with efficiencies of 32% and 23%, respectively (Figure 37c).
[0401] Based on these results, mouse embryos were transplanted into surrogate mothers, and offspring with m.C12336T and m.G12341A mutations were obtained (Figure 39c). Of the 27 embryos in total, 9 F0 mice (23%) showed C·G to T·A editing, with an efficiency range of 0.22 to 57% (Figure 37d, e). This indicates that the nonsense mutation in ND5 does not cause embryonic selection.
[0402] Example 4. Animal mitochondrial DNA editing 4-1. Construction of an expression vector for animal mitochondrial base editing with a nuclear export signal bound to it. A vector (Figure 40) was constructed to enable the expression of a protein bound to a nuclear export signal via TALE-DdCBE in animal cells. The cytomegalovirus promoter (CMV promoter) was used for the vector. From the N-terminus, it contains the mitochondrial transport signal, protein purification / detection tag, TAL array N-terminal domain, repeat site, C-terminal domain, DddA cytosine deaminationase split half, uracil glycosylase inhibitor, and nuclear export signal (Figure 40a). The nuclear export signal can be derived from, for example, the NS2 protein of MVM (Mirute virus of mice), but other sequences are also possible. The expressed protein is released from the nucleus and transported to mitochondria for base editing. The target DNA used in this process was the mitochondrial ND5 gene - ND5-like gene on chromosome 4, mitochondrial TrnA - chromosome 5, and mitochondrial Rnr2 - chromosome 6 (Figure 40b).
[0403] 4-2. DdCBE-NES in animal cell lines NIH3T3 cell line (ATCC CRL-1658) was transfected one day before in the afternoon at 1.5 × 10⁶ 4 The samples were dispensed into 12-well plates containing 1 ml of cell growth medium (DMEM + 10% Bovine Calf Serum) per well. The following morning, cells were transfected with untreated Mock, DdCBE, and DdCBE-MVM NES plasmids according to the Lipofectamine 3000 manufacturer's protocol. After culturing in a carbon dioxide incubator (37°C, 5% CO2) for 3 days, the cells were taken up, and the whole DNA was purified using the Qiagen Blood & Tissue Kit. Subsequently, the DNA was amplified using mitochondrial gene-specific PCR primers, and next-generation sequencing was performed using Illumina Miniseq equipment. The results were then analyzed for base editing efficiency using Cas-analyzer (www.rgenome.net).
[0404] Figures 40c, d, and e show the results of transfecting the NIH3T3 mouse cell line with DdCBE-NES to induce mutations in the mouse mitochondrial genes ND5, TrnA, and Rnr2. It can be seen that the efficiency differs depending on the combination of DdCBE.
[0405] 4-3. mitoTALEN in animal cell lines A TALEN that recognizes the sequence shown in Figure 40f was constructed, and MTS was ligated to this structure and introduced into mitochondria together with DdCBE. The experimental method was the same as in Example 4-2. A mismatch was intentionally introduced at the TALE recognition site so that one mismatch occurred with wild-type mtDNA and two mismatches occurred with mutant mtDNA. This is because TALE may not recognize a single sequence mismatch. As a result, it was confirmed that the efficiency was slightly increased in the +2 mismatch experimental group when treated with TALEN compared to the group treated with DdCBE alone.
[0406] 4-4. DdCBE-NES in animal embryos Using the DdCBE expression vector and the DdCBE-NES expression vector as templates, DNA is obtained by PCR, amplified by the T7 promoter and the DdCBE or DdCBE-NES expression region. This DNA is then used as a template to synthesize mRNA using T7 polymerase.
[0407] A pair of DdCBE mRNA or a pair of DdCBE-NES mRNA is mixed in a microinjection solution and microinjected into mouse fertilized eggs. Fertilized eggs cultured for 4 days developed into blastocysts, which were then lysed. Using these blastocysts as templates, the target region of mitochondrial DNA, distinct from nuclear DNA, was amplified using PCR. Further amplification of the index and sequencing adapter was then performed using additional PCR. High-throughput sequencing was performed using Illumina Miniseq equipment, and the base editing efficiency was confirmed using Cas-analyzer (www.rgenome.net). Additionally, nuclear DNA with a similar sequence to the mitochondrial target region was amplified and sequenced using PCR.
[0408] As a result, DdCBE induced mutations not only in mitochondrial DNA but also in similar DNA sequences in the nucleus (mitochondrial: 13.1%, nucleus: 3.2%). In the case of DdCBE-NES, the mutation efficiency of mitochondrial targets was increased to 18.2%, while the mutation efficiency of nuclear DNA was reduced to 0.2% (Figure 41a). While nuclear DNA mutations did not occur with other targets such as TrnA or Rnr2, a statistically significant increase in base editing efficiency was observed in mitochondrial DNA (*p<0.05, **p<0.01, ns not statistically significant).
[0409] 4-5: DdCBE and mitoTALEN in animal embryos To increase the proportion of mitochondrial DNA within the cell after C-to-T conversion, a TALEN that cleaves unedited mitochondrial DNA sequences was injected along with DdCBE at the ND5 gene site. The microinjection method and sequencing confirmation method were the same as in Examples 4-2 and 4-4. The group microinjected with DdCBE alone showed an editing efficiency of 11%, which increased to 33.3% when treated simultaneously with mitoTALEN, demonstrating a statistically significant increase in editing efficiency. Similarly, the group microinjected with DdCBE-NES alone showed an editing efficiency of 20.5%, while an efficiency of 36.8% was observed when treated simultaneously with mitoTALEN. This was also statistically significant (Figure 41b).
[0410] When microinjected fertilized eggs were transplanted into surrogate mothers, DdCBE similarly showed an editing efficiency of 10.9% in the newborn mouse offspring. However, when treated with DdCBE-NES and mitoTALEN together, an efficiency of 23.4% was achieved (Figure 41c).
[0411] When nuclear export sequences are attached to base-editing proteins during editing of animal mitochondrial genes, base editing is performed with higher efficiency, and in animal embryos, nonspecific base editing of nuclear-like sequences is suppressed in particular. Furthermore, even when mitochondrial sequence cleavage proteins are used simultaneously, higher efficiency of mitochondrial base editing can be expected.
[0412] Example 5. Divided DddA tox Deaminationase mutants We present a highly accurate DddA-derived cytosine base editor that can reduce the untargeted effects of DdCBE. This untargeted base editing effect is a phenomenon caused by the spontaneous binding of DddAtox deaminationase splits, independent of the interaction between TALE and DNA. Therefore, we created HF-DdCBE by substituting the amino acid residues located on the surface between DddAtox splits with alanine, so that HF-DdCBE does not function properly if the two deaminationase pairs linked to TALE cannot bind to DNA. Whole mitochondrial genome analysis confirmed that HF-DdCBE is highly efficient and precise, unlike conventional DdCBE which causes numerous undesirable untargeted C-to-T conversions in human mitochondrial DNA.
[0413] 5-1. Method Plasmid construction. Point mutations were introduced into the DdCBE expression plasmid. The plasmid was amplified using Q5 Site-Directed Mutagenesis (NEB) mutation primers (Table 7), and the results were confirmed by Sanger sequencing.
[0414] [Table 7] TIFF2026062789000028.tif249170TIFF2026062789000029.tif249170TIFF2026062789000030.tif191170
[0415] For the assembly of bound surface mutations, a miniprep-prepped mutant expression plasmid was combined with a module vector (each encoding a TALE sequence), BsaI-HFv2 (10U), T4 DNA ligase (200U), and reaction buffer in a single tube. Restriction enzyme and ligase reactions were then performed 20 times in a thermocycle at 37°C for 5 minutes and 50°C for 5 minutes, followed by treatment at 50°C for 15 minutes and 80°C for 5 minutes. The bound plasmid was introduced into E coli DH5a by chemical transformation, and the final structure was confirmed by Sanger sequencing. For cell line introduction, the plasmid was midiprep-prepped.
[0416] Culture and transfection of mammalian cell lines. The HEK 293T / 17 (CRL-11268, American Type Culture Collection (ATCC)) cell line was cultured at 37°C in a 5% CO2 environment. The cell line was grown without antibiotics in DMEM supplemented with 10% (v / v) fetal bovine serum (Gibco) medium, and mycoplasma testing was not performed. For lipofection, cells were placed in 24-well cell culture plates (SPL, Seoul, Korea) at a rate of 1 × 10⁶ 5 Cells were introduced at the specified cell density 18–24 hours prior to transfection. A total of 1,000 ng of plasmid DNA was introduced using Lipofectamine 2000 (Invitrogen), with 500 ng of each DdCBE cell segment. Cells were harvested 4 days after transfection.
[0417] Genome and mitochondrial DNA isolation for high-resolution sequence analysis. To isolate genomic DNA, the cell culture medium was removed, and then lysis buffer supplemented with proteinase K from the DNeasy Blood & Tissue Kit (Qiagen) was added to the cell culture plate to separate the cells from the bottom. Genomic DNA was then isolated according to the manufacturer's protocol. For whole mitochondrial genome sequence analysis, 200 μl of Mitochondrial isolation buffer A (ScienCell) was added to the culture plate from which the cell culture medium had been removed. After scraping the cells using a cell lifter, they were placed in a microcentrifuge tube and crushed using a disposable mortar and pestle. After 20 crushings, the well-crushed homogenate was centrifuged at 1,000 × g at 4°C for 5 minutes. The supernatant was placed in a new microcentrifuge tube and centrifuged at 10,000 × g at 4°C for 20 minutes. The precipitate was released into 10 μl of lysis solution (25 mM NaOH, 0.2 mM EDTA, pH 10) and boiled at 95°C for 20 minutes. To lower the pH, 1 μl of 1 M HEPES (free acids, without pH adjustment) was added to the mitochondrial lysis solution. 1 μl of this prepared solution was used as a PCR template for high-resolution sequence analysis.
[0418] High-processing sequence analysis. To prepare a deep sequencing library, nested primary and secondary PCR were performed using Q5 DNA polymerase, and the final index sequence was added. The library was used for fair-end read sequencing analysis using MiniSeq (Illumina). For whole-mitrogony genome analysis, isolated mitochondrial DNA was prepared using the tagmentation DNA prep kit (Illumina) according to the manufacturer's protocol. In all analyses, the results of fair-end sequencing analysis were combined using a single fastqjoin file and analyzed using CRISPR RGEN Tools (http: / / www.rgenome.net / ).
[0419] 5-2.Results When attempting chloroplast editing in plants, non-targeted base mutations appear on the chloroplast gene, thus raising questions about the accuracy of DdCBE. Two reasons for non-targeted base editing by DdCBE were considered. First, the non-specific binding between the TALE protein and DNA; second, the naturally occurring and unintentional binding of DddAtox hemispheres (Figure 42a). In this study, we focused on DddAtox split hemispheres and improved the binding surface of the two split proteins to prevent the binding of unwanted DddAtox hemispheres.
[0420] First, we focused on the mitochondrial ND1 (mtND1) gene and investigated whether each subunit (Left-TALE or Right-TALE) could bind to DNA individually and interact with other halves of TALE-free DddAtox (which lacks a binding sequence) to induce cytosine-thymine base editing. In human kidney germ cell lines (HEK293T), a DdCBE pair targeting the human mitochondrial ND1 (mtND1) gene (Left-TALE:G1397N (a structure in which the N-terminal G1397 DddAtox half binds to the C-terminus of the TALE sequence and assumes responsibility for the left-side portion of the overall recognition sequence) + Right-TALE:G1397C (a structure in which the C-terminal G1397 DddAtox half binds to the C-terminus of the TALE sequence and assumes responsibility for the right-side portion of the overall recognition sequence)) effectively edited the target sequence C11 from cytosine to thymine with an efficiency of 60.7% (Figure 43a). Furthermore, when treated with TALE-free DddAtox halves that can pair with each subunit, base editing was observed, although less effective than with the original DdCBE pair. Therefore, Left-TALE and Right-TALE, respectively, can be seen to bind to the target ND1 sequence and edit the base with efficiency of 31% and 8.1% (Figure 43). In other words, DdCBE pairs fused with both TALE proteins, respectively, are 2.0 times (=60.7% / 31%) or 7.5 times (=60.7% / 8.1%) more efficient than when bound to only one TALE. Clearly, the N-terminal half of DddAtox bound to TALE can bind to the C-terminal half of DddAtox without a TALE sequence, enabling the deamination reaction, and vice versa.
[0421] Since DddAtox splits into two hemispheres (G1333 and G1397), we created DdCBE pairs targeting the mtND1 gene at the G1333 position (Left-TALE:G1333-N and Right-TALE:G1333-C). We then independently used the Left-TALE:G1333-N and Right-TALE:G133-C structures to bind to the paired TALE-free DddAtox hemisphere to confirm whether cytosine-thymine editing occurred. As expected, each TALE-fused DdCBE showed a base editing efficiency of 32.7% (left-side TALE conjugate) or 18.1% (right-side TALE conjugate) at the C8 position, which is comparable to the 56.1% shown by the original DdCBE pair. Therefore, it can be seen that DdCBE pairs bound to TALE act with 1.7 times (=56.1% / 32.7%) or 3.1 times (=56.1% / 18.1%) more efficiency than when they act separately. In summary, these results suggest that DdCBE can induce undesirable non-targeted base editing even when binding only one TALE sequence. Since TALE proteins bind to DNA even with a small number of mismatched sequences, DdCBE has the potential to induce non-targeted base editing in organelles or nuclear genes.
[0422] To reduce non-targeted base editing resulting from the binding of split DddAtox halves, we attempted to develop a high-precision DdCBE. We hypothesized that improving the binding surface sites of the split dimers could prevent or inhibit self-binding. Therefore, using PyMOL software, we used Python code (InterfaceResidues.py) to find the amino acid residues on the binding surface of two split DddAtox halves (split positions G1333 and G1397) within a range of 1 square angstrom. As a result, we were able to find 9 amino acid residues in G1397-N (the N-terminal DddAtox half split at position G1397), 4 residues in G1397-C (the C-terminal DddAtox half split at position G1397), 14 amino acids in G1333-N (the N-terminal DddAtox half split at position G1333), and 15 amino acid residues in G1333-C (the C-terminal DddAtox half split at position G1333) (Figures 42b and 42c). Subsequently, several mutant DddAtox hemispheres were created by substituting the identified amino acid residues with alanine. These binding surface mutant DdCBEs were then treated in the HEK293T cell line with wild-type DdCBE pairs or TALE-free DddAtox pairs, respectively, and their base editing efficiency was observed. In some binding surface mutants of the G1397 split, including C1376A, M1390A, and F1412A, cytosine-thymine editing efficiency could not be observed at the target site, not only with TALE-free DddAtox but also when coupled with the wild type. This suggests that such mutations cannot interact with other DddAtox hemispheres even when located close together. Some mutations, such as E1381A and V1377A, showed highly efficient base editing when coupled with TALE-free hemispheres, indicating that such mutations are changes unrelated to dimer interaction.
[0423] Importantly, some mutations, such as K1389A, K1410A, and T1413A, showed high activity when bound to wild-type DdCBE pairs, but low activity when used with TALE-free pairs. For example, the K1410A mutation showed an efficiency of 53.2%, similar to that of wild-type DdCBE pairs (60.7%), but when used with TALE-free pairs, it showed an efficiency of 0.9%, a difference of 59.1 times (= 53.2% / 0.9%). As mentioned above, the wild-type showed a difference of 7.5 times (= 60.7% / 8.1%). Furthermore, these mutations selectively edited bases more selectively than wild-type DdCBE. The mutants edited C8, C9, or C within the editing range. 13 C 11 It showed strong selectivity for this, which can be compared to the wild type in which all four cytosines were edited to be at least 6.7% (Figure 43b).
[0424] Furthermore, in G1333, we screened 29 mutations (14 G1333N and 15 G1333C) and were able to obtain several suitable binding surface mutations (Figure 44). Many of the mutations either decreased efficiency when coupled with the wild type (I1299A, Y1316A, Y1317A, and F1329A) or increased efficiency when coupled with TALE-free (S1300A and T1314A). However, it is noteworthy that mutations such as K1389A, T1391A, and V1393A showed good base editing when coupled with the wild type, but low efficiency with TALE-free. For example, K1389A showed a 38-fold difference (=45.4% / 1.2%), compared to only about a 3.1-fold difference (=56.1% / 18.1%) when paired with the wild-type DdCBE. Furthermore, in the case of K1389A, it was sometimes observed to be more efficient than the wild type. In addition, these mutations are C 11 or C 13It was observed that editing efficiency was concentrated at the C8 position, which can be compared to the increased editing efficiency at all cytosine positions in over 19% of wild-type DdCBEs. Also noteworthy is that K1410A in G1397-C preferentially edits C8, while K1389A in G1333-C selectively and more efficiently edits C11. On the other hand, wild-type DdCBE vs (G1333 or G1397 DddA tox The splits all exhibit low selectivity. These results suggest that the aforementioned binding surface mutations may reduce the effect of causing undesirable editing of multiple bases within the target site in DdCBE.
[0425] Example 6. Full-length deaminationase The DddA-derived cytosine base editing enzyme (DdCBE), containing the divided bacterial toxin DddAtox, a TALE array, and a uracil glycosylase inhibitor (UGI), enables the conversion of target cytosines to thymine in eukaryotic cell nuclei, mitochondrial DNA (mtDNA), and plant chloroplast DNA. DddAtox, which induces toxicity in bacteria, originates from Burkholderia cenocepacia and enzymatically deaminates cytosine within double-strand DNA. To avoid host cell toxicity, DddAtox is divided into inactive halves, each bound to a TALE DNA-binding protein to create DdCBE. The two inactive forms, bound to the TALE array, become functional when bound in close proximity to the target DNA by TALE. The cytosine-to-thymine base conversion is induced in a region of 14-18 bases between the two TALE binding sites. Unlike CRISPR-derived base editors that cannot edit organelle DNA, DdCBE enables targeted base editing in both nuclear and organelle DNA, but it has the drawback of requiring two TALE configurations instead of one to induce it. The first drawback is that using two TALE arrays limits the targetable sites because the TALE needs to bind to the target DNA site at both the 5' and 3' ends to thymine. The second drawback is that transmitting two TALE configurations instead of one is often inefficient and challenging. Dose-limited viral vectors, such as adeno-associated virus (AAV) vectors (dose, ~4.7kbps) widely used in gene therapy, cannot accommodate two DdCBE encoding sequences because the combination of dimeric DdCBEs is too large (2 × 4.1kbps, including promoter and polyA signaling). Cloning two TALE array DNAs into a single vector of a larger dose can be difficult due to the high similarity of the two TALE array sequences. Finally, using two TALE arrays instead of one can exacerbate off-target effects. To overcome the limitations of the dimerized DdCBE with fragmented DddAtox, we present mDdCBE (monomer DdCBE) in the non-toxic full-length DddAtox form, which induces cytosine-to-thymine conversion in target DNA of the nucleus and organelles.
[0426] 6-1. Method Plasmid construction. DddA mutants were PCR-amplified using synthesized full-length DddAtox (gBlock, IDT) as a template, with primers from Table 8 and Q5 DNA polymerase (NEB). These PCR products were cloned using Gibson assembly (NEB) at the p3s-BE3 site where Apobec1 was cleaved with BamHI and Sma I (NEB). TALE-DddAtox (Addgene #158093, #158095, #157842, #157841) were prepared by cleaving the plasmid with BamHI and Sma I, PCR-amplified the DddA mutant using primers from Table 8, and cloning using Gibson assembly. The secured plasmids were transformed into chemically produced E. coli DH5a by thermal shock, and the plasmid sequences of the surviving colonies were analyzed by Sanger sequencing. The final plasmids were midiprep (Macherey-Nagel) for cell transfection.
[0427] [Table 8]
[0428] Random mutagenesis was performed. Error-prone PCR was conducted using full-length DddAtox (gBlock, IDT) synthesized according to the manufacturer's protocol using the GeneMorph ii Random mutagenesis kit (Agilent) as a template. In summary, random mutations of 0–16 mutations / kb were introduced using 1 ng, 100 ng, and 700 ng of DddAtox DNA as templates. Full-length DddAtox gBlock was pre-amplified by PCR using the primers shown in Table 8. The combined PCR results were cloned into p3s-UGI-Cas9 (H840A) cleaved with Sma1 and Xho1 using Gibson assembly (NEB). E. coli DH5α cells, prepared chemically, were transformed with the plasmid by thermal shock, and the plasmid sequences of the surviving colonies were analyzed by Sanger sequencing. Of the plasmids analyzed, the p3s-UGI-nCas9(H840A)-DddAtox plasmid containing a coding frame was transfected into HEK293T cells along with sgRNA, and its editing activity was confirmed by targeted deep sequencing.
[0429] Mammalian cell culture and transfection. HEK293T (ATCC, CRL-11268) cells and HeLa (ATCC, CCL-2) cells were cultured at 37°C under 5% CO2. Cells were cultured in DMEM supplemented with 10% (v / v) fetal bovine serum (Welgene) and 1% penicillin / streptomycin (Welgene). Cells were seeded in 48-well plates (Corning) at densities of 3 x 10⁵ cells (HEK293T) and 4 x 10⁴ cells (HeLa) 24 hours before transfection and transfected with Lipofectamine 2000 (Invitrogen) and Cas9-fused DddA plasmid (750 ng) and sgRNA (250 ng). TALE-DddA was transfected into HEK293T cells using a 200 ng plasmid and Lipofectamine 2000. The sgRNA sequences are shown in Table 9.
[0430] [Table 9]
[0431] Preparation of genomic and mitochondrial DNA. Cells transfected with the Cas9-fusion DddA mutant were harvested 2 days after transfection, and cells transfected with TALE-DddA were harvested 3 days after transfection. Genomic and mitochondrial DNA were isolated using the DNeasy blood and tissue kit (Qiagen). For large-scale analysis, DNA was extracted using 100 μL of cell lysis buffer (50 mM Tris-HCl, pH 8.0 (Sigma-Aldrich), 1 mM EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) containing 5 μL of proteinase K (Qiagen). The lysates were reacted at 55°C for 1 hour, followed by reaction at 95°C for 10 minutes.
[0432] 6-2.Results The amino acid sequences of the wild type and the novel full-length DddA were compared, and the changed amino acids are shown in gray boxes in Figure 45.
[0433] As shown in Figure 46, DddA was linked to the N-terminus of Cas9 using a linker consisting of 16 amino acids, and UGI (uracil glycosylase inhibitor) and NLS (nuclear localization signal) were linked to the C-terminus using a linker consisting of 4 amino acids. Conversely, DddA was linked to the C-terminus of Cas9 using a linker consisting of 16 amino acids, and UGI and NLS were linked to the N-terminus using a linker consisting of 4 amino acids.
[0434] In this invention, experiments were conducted using DddA-Cas9 (D10A, D10A and H840A)-UGI. DddA is bound to the DNA-binding protein zinc finger protein, the TALE module, and cytosine can be replaced with thymine using only one module. Conventional splits require two modules, but full-length DddA can do so with only one module. These two DNA-binding proteins link NLS (nuclear localization signal), MTS (mitochondrial targeting sequence), and CTP (chloroplast transit peptide) to replace cytosine with thymine not only in genomic regions but also in mitochondria and plant chloroplasts, which is not possible with Cas9. As shown in Figure 47, the activity of replacing cytosine with thymine in the TC molecule was confirmed in the human cell genomic contexts ROR1 site (a), HEK3 site (b), and TYRO3 site (c). The activity of substituting cytosine with thymine at a target location 25 bp away from the target site was confirmed (a). A1341D KRKKA showed activity of substituting the second cytosine of CC-motif with thymine (a, b). The catalytic mutant E1347A also showed activity of substituting cytosine with thymine of TC-motif (a, b, c). The red underline indicates the site of Cas9 binding. Efficiency is expressed as a percentage of the cytosine substitution with thymine in deletion-free reads in the overall sequencing reads. It also shows the percentage of deletions in the overall sequencing reads.
[0435] As shown in Figure 48, the red square boxes indicate the portion where DddAtox activity was confirmed after division. The same portion was divided into three target sites using full-length DddA and the activity was measured. The divided product uses different orthogonal Cas9s with different PAMs to replace cytosine between two Cas9s with thymine, but in this case, it is difficult to precisely replace the desired cytosine with thymine. However, full-length DddA can target the portion where Cas9 binds to the same target site by dividing it into three parts, allowing for precise replacement of the desired cytosine with thymine. Efficiency is expressed as the percentage of cytosine replaced with thymine in deletion-free reads within the overall sequencing read. It also shows the percentage of deletions in the overall sequencing read.
[0436] As shown in Figure 49, the activity of full-length DddA was measured in the human cell genome contexts TRAC site 1 (a), TRAC site 2 (b), FANCF (c), and HBB (d). The red underlines indicate the sites where Cas9 binds. Efficiency is expressed as the percentage of cytosine substitution with thymine in deletion-free reads out of the total sequencing reads. It also expresses the percentage of deletions out of the total sequencing reads.
[0437] As shown in Figure 50, DddA activity was measured in human cell genome contexts TYRO3 (a), ROR1 (b), HEK3 (c), EMX1 site 2 (d), TRAC site 1 (e), and HBB (f) using DddA-dCas9(D10A, H840A)-UGI. Efficiency is expressed as the percentage of cytosine substitutions with thymine in the overall sequencing reads. No deletions were observed.
[0438] To obtain non-toxic, full-length DddAtox mutants useful for base editing, two approaches were employed: structurally-based mutations and random mutations. In the first approach, DddAtox mutants with reduced DNA binding or reduced catalytic activity were fused to inactive CRISPR-Cas9 (dCas9) or nickase (nCas9) mutants to develop a novel base-editing device in which target cytosine in human-derived cells was replaced with thymine. For this purpose, we attempted to subclone DddAtox into an expression vector by substituting the positively charged amino acid with alanine (Figure 51a). We hypothesized that such mutants would weaken the binding of negatively charged dsDNA, potentially avoiding toxicity. Most alanine-substituted mutants failed to form E. coli transformants (Figure 51b). Base analysis of plasmid DNA isolated from the obtained transformants revealed that various frame mutations were induced in the protein-forming region. These full-length DddAtox mutants, while under the control of a mammalian expression promoter, were weakly expressed in E. coli, leading to apoptosis. Fortunately, we were able to obtain several triple, quadruple, or quintuple ("AAAAA") alanine substitution mutants without frameshift mutations. Furthermore, the active site mutation E1347A was successfully cloned.
[0439] Next, we investigated whether AAAAA mutants fused to D10A nCas9 or dCas9 and UGI could induce base editing in human embryonic kidney 293T (HEK293T) cells (Figure 51c, d). Base editing device 2 (or 3), consisting of rat APBEC1 deaminose, uracil glycosylase inhibitor (UGI), and dCas9 (or D10A nCas9), was active in a narrow region within the protospacer region, but the AAAAA mutant induced up to 43% cytosine-to-thymine substitution just 5' above the protospacer region. Unexpectedly, the E1347A mutation was found at the same frequency of 37% (nCas9 fusion) or 16% (dCas9 fusion) C -3Base editing was induced at the 51 position (Figure 51c), confirming that the E1347A mutation does not induce complete deamination inactivation of DddAtox and possesses sufficiently high deamination activity to achieve base editing in human-derived cells. However, the E1347A mutant coupled with the quintuple AAAAA mutation failed to induce base editing. Furthermore, E1347A without migration mutations fused to dCas9 or nCas9 and UGI, AAAAA, and other mutants substituted with alanine (Figure 53) confirmed that editing occurred at positions up to 25 bases above the protospacer region, and at many other sites, editing efficiencies of up to 26% were observed (Figure 54). Moreover, it was highly efficient in HeLa cells with editing frequencies of up to 60% (Figure 55). Base editing induced by such fusion proteins was maintained in cells for up to 21 days, suggesting that such base editing is non-cytotoxic (Figure 56).
[0440] Due to a change in the editing window of the cytosine base editor, the alanine-substituted mutant attempted to fuse with the C-terminus of H840A nCas9. Unexpectedly, we failed to obtain an intact component without frameshift mutations. Therefore, we performed error-prone PCR to introduce random mutations into the DddAtox coding sequence and were able to obtain a non-toxic full-length DddAtox mutant with four-point mutations S1326G, G1348S, A1398V, and S1418G (referred to as "GSVG") (including the sequence of SEQ ID NO: 276, in which S1326G, G1348S, A1398V, and S1418G each involve substitution of S at position 37 with G; G at position 59 with S; A at position 109 with V; and S at position 129 with G in the amino acid sequence of SEQ ID NO: 269, respectively, as shown in Figure 52a). Furthermore, this mutant fused to the C-terminus of dCas9, D10A nCas9, and Cas9, and to the N-terminus of dCas9, nCas9, and Cas9. In human-derived cells, these fusion proteins, excluding wild-type Cas9, induced cytosine-to-thymine conversion at various sites with an efficiency of up to 38% (Figures 52b, 57, 58). Interestingly, fusion proteins containing the GSVG mutant fused to the C-terminus of dCas9, D10A nCas9, and H840A nCas9 showed cytosine base editing 3' downstream of the protospacer adjacent motif (PAM), while fusion proteins containing the same mutant fused to the N-terminus, namely dCas9 and nCas9, induced base editing 5' upstream of the protospacer (Figure 52c). As expected, fusion proteins containing Cas9 induced insertions and deletions rather than base substitutions.
[0441] To find out which mutations are important for GSVG variants, S SVG, G G VG, GS A G, and GSV SWe attempted to generate four revertants by site-directed mutagenesis. SSVG, GSAG, and GSVS revertants were obtained, but a GGVG mutant fused to the C-terminus of nCas9 was not. G1348 is located immediately adjacent to E1347, the core site of the catalytic site. The G1348S mutation reduced catalytic activity and avoided cytotoxicity in E. coli. We measured the editing frequencies of the three revertants and the GSVG mutant at two target sites in transfected cells for up to 21 days. The frequency of cytosine-to-thymine editing induced by GSAG and GSVS gradually decreased by ~2 times from 3 to 21 days post-transfection, and these two revertants were slightly cytotoxic, while GSVG and SSVG remained (Figure 59). These results suggest that G1348S is essential in the GSVG mutant, S1326G is neutral, while A1398V and S1418G reduce cytotoxicity.
[0442] In summary, the results describe the invention of a base editing device with a newly modified editing window, formed by fusing a non-toxic, full-length DddAtox mutant with reduced affinity to dsDNA (AAAAA), weakened deamination activity (E1347A and possibly GSVG), or reduced cytotoxicity (GSVG) to dCas9 or nCas9. Such base editing groups can be named dCas9-mDdBE (a base editing device derived from DddA, consisting of a full-length monomer DddAtox mutant fused to the C-terminus of dCas9), nCas9-mDdBE, mDdCE-dCas9, and mDdCE-nCas9, and are BE2 or BE3 or later base editing devices used for base editing at upstream or downstream positions of protospacers that are unreachable in BE2 or BE3.
[0443] Furthermore, we investigated whether non-toxic full-length DddAtox mutants could be used for mitochondrial DNA editing. Among numerous mutants, only two, GSVG and E1347A, successfully fused to the C-terminus of a TALE array designed to bind to mitochondrial genes, ND4 and ND6. Monomer DdCBE (mDdCBE) containing the GSVG mutant achieved base editing at target nucleotide positions at up to 31% (ND4) (Figure 60a) and 27% (ND6) (Figure 60b), comparable to the originally split DdCBE pair. mDdCBE containing E1347A also converted target cytosine to thymine, albeit with reduced efficiency at up to 7.2% (ND4) and 8.9% (ND6) editing ratios. Interestingly, the original DdCBE pair specialized for the ND4 gene (G1333 split) had an editing efficiency of 0.8% at the C4 position, while the two mDdCBEs containing GSVG showed higher editing efficiencies of 26% and 31%. These results suggest that the split dimer DdCBE and mDdCBE may have different mutation patterns, and that mDdCBE may be complementary to the dimer DdCBE, thus enabling the induction of diverse mutations at a given target site.
[0444] One potential advantage of mDdCBE compared to the split-dimer DdCBE is that the off-target effects due to nonspecific TALE-DNA interactions are halved compared to the dimer DdCBE. Dimer DdCBE with split DddAtox can function at half sites that can only bind one subunit, which leads to unwanted off-target mutations. The inactive DddAtox half of a DdCBE pair can recruit the other inactive half to form a functional deaminosease. To confirm this hypothesis, we co-transfected HEK293T cells with plasmids encoding one subunit of dimer DdCBE and plasmids encoding the DddAtox half without TALE, and measured the editing frequency at both mitochondrial target sites. As expected, cytosine-to-thymine editing was observed at the target site at frequencies of 0.7–3.6% (Figure 60c–f). These results suggest that the DddAtox halves split in a DdCBE pair may interact with each other at half of the site, potentially causing undesirable off-target mutations, and that mDdCBE can avoid half of the off-target mutations induced by dimeric DdCBE.
[0445] Example 7. Highly efficient A-to-G base editing in human cells using DdABE Mitochondrial DNA base editing using DddA-derived cytosine base editing (DdCBE) has enabled the creation of various cell lines and animal disease models, opening new avenues for treating mitochondrial genetic diseases. However, DdCBE is limited to inducing almost exclusively TC-to-TT base editing, covering only about 1 / 8 of all possible cases. Therefore, we developed TALED, a transcription activator-like effector (TALE) ligated with two types of deaminoses. Here, TALE contains a DddAtox cytosine deaminosement mutant that has lost its catalytic ability and has been customized to bind to the desired DNA region, as well as the TadA protein, a DNA adenine deaminosement derived from E. coli. Unlike conventional base editing technologies that were only capable of cytosine base editing in the conventional TC context in human mitochondria, TALED enables base editing for all A-to-G sequences. In fact, the customized TALED was able to induce adenine base editing with high efficiency (up to about 50%) in multiple targets in human cells.
[0446] To develop a novel base editing technology that did not exist before, we selected the ABE8e TadA variant (TadA*) from among many TadA variants. This is because it not only enables highly efficient adenine editing but has also been improved to be compatible with a wide variety of DNA-binding proteins, thus improving compatibility with actual TALE or ZFP.
[0447] We experimented to see if we could actually induce base editing in mitochondrial DNA by fusing TadA* and MTS into a TALE customized to target ND1 or ND4 for the first time. Targeted deep sequencing results showed that the fusion protein does have adenosine base editing efficiency, although it is very low. It induced adenine base editing with a maximum efficiency of 1.2% at the ND1 site (Figure 67a) and 0.6% at the ND4 site (Figure 67b). Although the efficiency is very low, this shows that when fused to a TALE, it can induce base editing even on double-stranded target DNA, compared to TadA* which is known to function specifically only on single-stranded DNA.
[0448] With the result that adenine base editing can occur in mitochondrial DNA, we considered fusing it with the already known DddAtox protein to further increase its efficiency. The DddAtox protein is an interbacterial toxin derived from Burkholderia cenocepacia that performs cytosine deaminement. This protein functions on the double helix of DNA, so using it allows the TadA* adenine deaminose enzyme to access the target DNA more effectively. In conventional DdCBE using DddAtox, the DddAtox protein is split in half, with each half forming a TALE (Left-TALE (L-TALE)) that attaches to the left DNA portion and a Right-TALE (R-TALE) that attaches to the right DNA portion. These split proteins are then fused together, and a uracil glycosylase inhibitor (UGI) is added to increase cytosine base efficiency (TALE-Split DddAtox-UGI). The reason for using DddAtox in this split form is that using the full-length protein results in cytotoxicity. First, we experimented by attaching TadA* instead of UGI to one side of a DdCBE targeting the ND1 site, creating L-TALE-Split-DddAtox-TadA* and R-TALE-1397C-TadA*. Surprisingly, when human cells were transmitted with TadA* on one side and 1397N and UGI on the other, we were able to confirm that both A-to-G and C-to-T base editing occurred (Figure 62c). In the case of conventional DdCBE, approximately 20% cytosine base editing occurred and no adenine base editing occurred at all. However, when TadA* was present on one side, cytosine base editing was reduced to half the level, and about 10% adenine base editing occurred (Figure 62c). The efficiency of adenine base editing and cytosine base editing was similar (Figure 62c).
[0449] Of course, the simultaneous occurrence of cytosine and adenine base editing may be useful for inducing random mutations, but in treating diseases, especially mitochondrial genetic diseases such as LOHN and MEALS that occur with C-to-T mutations, it is desirable to induce only adenine base editing. Therefore, UGI was removed to eliminate the simultaneous occurrence of cytosine base editing. In DdCBE, UGI is fused as an inhibitor to prevent the repair of U by uracil glycosylase, a repair protein within cells, during DNA repair, when the cytosine deaminationase DddAtox deaminates C to U. Therefore, it was thought that removing such UGI would maintain adenine base editing efficiency while suppressing cytosine base editing. Surprisingly, we confirmed that ND1-targeted TALE deaminosease pairs without UGI caused almost no cytosine base editing (<0.5%) and highly efficient adenine base editing (approximately 50%) (Figure 63a, c). This is far more efficient than those with UGI. Furthermore, we confirmed this with ND4-targeted TALE deaminosease pairs, and confirmed that they edited only adenine bases with high efficiency (approximately 35%), similar to targeting ND1 (Figure 63b, d). Thus, we developed TALED, a novel adenine deaminase that functions on DNA double strands by fusing the DddAtox system with TadA*, and were able to perform adenine base editing in human mitochondria for the first time in the world. In addition, we were able to edit adenine bases approximately 50 times more efficiently than with TALE attached only with TadA*.
[0450] Next, we attempted to induce adenine base editing using either the full-length E1347A DddAtox mutant with further reduced catalytic activity, or mutants (AAAAA and GSVG) that maintained the developed catalytic activity while eliminating cytotoxicity. Since the objective was not to induce cytosine base editing, but to induce adenine base editing in the DNA double helix, we could use the full-length E1347A DddAtox mutant with reduced cytosine base editing activity, and since we confirmed that the efficiency of cytosine base editing would be lost without UGI, we could also use mutants that simultaneously eliminated only cytotoxicity. We created two types of TALEDs containing the full-length mutants (Figure 64a). The first type contained both TadA*(AD) and the full-length DddAtox mutant in one TALE (mTALED), and the second type involved separately fusing TadA*(AD) and the full-length DddAtox mutant into each TALE (dTALED). These two types were then tested (Figure 64a). Surprisingly, both types of TALED targeting ND1 were confirmed to induce adenine base editing with high efficiency (Figure 64b). mTALED showed an efficiency of up to approximately 45%, and dTALED also showed an adenine base editing efficiency of approximately 50% (Figure 64b). Similar experiments were conducted at both the ND1 and ND4 sites, and adenine base editing was induced with similarly high efficiency (Figure 64c). Furthermore, even when using the full-length E1347A DddAtox mutant, which lacked cytosine base editing activity, it was confirmed to induce adenine base editing with high efficiency (Figures 64b, c). This indicates that, despite the loss of cytosine deamine ability, the role of helping to bring TadA* close enough to the DNA double helix is maintained. When the above results are examined in detail at the single-base level (Figures 65, 66), it can be seen that adenine base editing occurs very close to where TALE binds to DNA. Furthermore, when using two TALEs, base editing occurs only in the spacer between them, but surprisingly, the mTALED used by a single TALE also had a similar target length (Figures 65, 66).
[0451] Furthermore, we were curious whether the system would function with zinc finger protein (ZFP) systems and with nuclear DNA. Therefore, we created an NC-type ZFP targeting nuclear DNA and fused it with the fragmented DddAtox and TadA* (Figure 61a). At this time, we fused TadA* to the ZFP at various positions (Figure 61b). Among the various forms, we were able to create individuals that edited up to approximately 10% of adenine bases in nuclear DNA (Figure 61d). Because UGI is present on one side, high efficiency in cytosine base editing was also observed (Figure 61c). Having confirmed that the ZFP-DddAtox-TadA* system functions in human cell nuclear DNA, we further experimented to see if it would function in mitochondria. At this time, we referenced the structure that functioned most efficiently in nuclear DNA. Instead of the Nuclear Localization Signal (NLS) transmitted to nuclear DNA, a Mitochondrial Targeting Sequence (MTS) directed to mitochondria was attached. Experiments were conducted using ZFPs (NC type) that target the ND1 site, with 1397N fused to the left ZFP and 1397C and TadA* fused to the right ZFP. As a result, adenine base editing was found to occur with an efficiency of approximately 3% (Figure 61.g). Although the adenine base editing efficiency was lower than that of TALED, by optimizing various conditions such as the linker that connects proteins, it should be possible to achieve adenine base editing with good efficiency using this ZFP system.
[0452] To date, gene editing technology has made remarkable progress. CRISPR-based gene scissors (CRISPR-Cas9, base editor, prime editor, etc.) have improved off-target capabilities, increased efficiency, and diversified. However, despite these many advancements, there have been limitations in treating mitochondrial genetic diseases. This is because CRISPR-based technologies consist of a catalyzing protein and a gRNA that acts as a guide to the target. However, unlike proteins, there is no way to deliver gRNA to mitochondria, and until now, technologies for handling mitochondrial genes have been limited to cutting the DNA and removing the mitochondrial DNA. However, the David R. Liu group in the United States demonstrated DdCBE, the first to induce base editing in mitochondria. Because DdCBE contains the cytosine deaminonase DddAtox, which functions on DNA double strands, it can fuse with the DNA-binding protein TALE to deliver only the protein and induce base editing. However, such DdCBE restrictively induces base editing only on TC-type cytosines, which has many limitations in actually creating disease models or treating genetic diseases. Therefore, we created TALED, the world's first mitochondrial adenine base editing system, which boasts a high efficiency of up to approximately 50% and performs base editing with various types of adenine in the surrounding target region. Furthermore, when UGI is present, it can simultaneously edit cytosine and adenine bases, making it useful for inducing random mutations. In the absence of UGI, however, cytosine base editing does not occur, and only adenine base editing takes place, making it usable as a specific adenine base editing technology. We also demonstrated that it can be applied to ZFP systems, enabling adenine base editing even in nuclear DNA. The development of TALED offers solutions to many mitochondrial genetic diseases, makes it possible to create corresponding disease models, and will be a useful tool for conducting much untapped research related to mitochondrial genes.
[0453] Having described in detail certain aspects of the present invention, it will be clear to those with ordinary skill in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the invention. Therefore, the substantial scope of the invention is defined by the appended claims and their equivalents. [Industrial applicability]
[0454] According to the present invention, by substituting specific amino acid residues on the surface of the cytosine deaminose split bond during gene editing, the nonselectivity of undesired cytosine deaminose can be reduced.
[0455] In relation to full-length cytosine deaminosease, it is possible to edit parts that are difficult to edit using conventional cytosine base editors. Among conventional cytosine base editing methods, Apobec1, which is used as a deaminose, is known as an oncogene, so its use for therapeutic purposes is limited. However, the full-length deaminosease developed in this study is judged to not have this problem.
[0456] Including the DNA-binding protein, it is small, approximately 2.5 kb, making it suitable for gene therapy using AAV vectors. It facilitates mRNA and RNP transfer and can be used for the production of useful substances using prokaryotes.
Claims
1. (i) DNA-binding proteins; and (ii) A fusion protein comprising a first and a second split product derived from cytosine deaminose or a variant thereof, The first and second splits are fusion proteins, each in a form that binds to the DNA-binding protein.
2. (i) DNA-binding proteins; and (ii) A fusion protein comprising a non-toxic full-length cytosine deaminoase derived from cytosine deaminoase or a variant thereof.
3. The fusion protein according to claim 1, characterized in that each of the first and second splits lacks cytosine deaminose activity.
4. The fusion protein according to claim 1, wherein the first split includes one or more sequences selected from the group consisting of G33, G44, A54, N68, G82, N98, and G108 from the N-terminus of the sequence of Sequence ID No.
1.
5. The fusion protein according to claim 1 or 2, characterized in that the cytosine deaminoase is derived from a deaminoase (DddA) that acts on double-stranded DNA or its orthologue.
6. The fusion protein according to claim 1, wherein the second split includes one or more sequences selected from the group consisting of G34, P45, G55, N69, T83, A99, and A109 from the sequence of Sequence ID No. 1 to the C-terminus.
7. The fusion protein according to claim 1, wherein the mutant of the cytosine deaminose has one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30, and 31 of a first split of the sequence of Sequence ID No. 1, which includes the sequence from the N-terminus to G44, replaced with other amino acids.
8. The fusion protein according to claim 1, wherein the mutant of the cytosine deaminose has one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 of a second split of the sequence of Sequence ID No. 1, which includes the sequence from P45 to the C-terminus, replaced with other amino acids.
9. The fusion protein according to claim 1, wherein the mutant of the cytosine deaminose has one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of the sequence of Sequence ID No. 1 that includes the sequence from the N-terminus to G108, replaced with other amino acids.
10. The fusion protein according to claim 1, wherein the mutant of the cytosine deaminose has one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of a second split of the sequence of Sequence ID No. 1, which includes the sequence from A109 to the C-terminus, replaced with other amino acids.
11. The fusion protein according to claim 2, wherein the non-toxic full-length cytosine deaminoase of SEQ ID NO: 1 has one or more amino acids selected from the group consisting of 37, 59, 109, and 129 substituted with other amino acids.
12. The fusion protein according to any one of claims 7 to 11, wherein the other amino acid is alanine.
13. The fusion protein according to claim 2, wherein the non-toxic full-length cytosine deaminoase is at least one selected from the group consisting of SEQ ID NOs: 12 to 22.
14. The fusion protein according to claim 1 or 2, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease.
15. The fusion protein according to claim 1 or 2, characterized in that the DNA-binding protein is linked to cytosine deaminoase or a variant thereof via a peptide linker containing 2 to 40 amino acid residues.
16. The linker comprises the fusion protein according to claim 15: 2a.a Linker: GS and, 5a.a Linker: TGEKQ (Sequence ID 8), 10a.a Linker: SGAQGSTLDF (Sequence ID 9), 16a.a Linker: SGSETPGTSESATPES (Sequence ID 10), 24a.a Linker: SGTPHEVGVYTLSGTPHEVGVYTL (Sequence ID 115) or 32a.a Linker: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS (Sequence ID 11)
17. The fusion protein according to claim 1, wherein each of the first and second splits is bound to the N-terminus or C-terminus of a zinc finger protein.
18. The fusion protein according to claim 1 or 2, wherein a single TALE array or a first TALE array and a second TALE array are each bound to the cytosine deaminose enzyme.
19. (iii) The fusion protein according to claim 1 or 2, further comprising adenine deaminose enzyme.
20. The fusion protein according to claim 19, wherein the adenine deaminoase is a mutant of Escherichia coli TadA and is deoxy-adenine deaminoase.
21. The fusion protein according to claim 19, wherein the adenine deaminose is bound to the N-terminus or C-terminus of a DNA-binding protein or cytosine deaminose or a variant thereof.
22. A nucleic acid encoding the fusion protein according to any one of claims 1 to 21.
23. The nucleic acid according to claim 22, which is ribonucleic acid or DNA.
24. A base editing composition comprising a fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22.
25. The composition according to claim 24, further comprising UGI (uracil glycosylase inhibitor).
26. A composition for base editing of eukaryotic cells, comprising a fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22.
27. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A composition for base editing in plant cells, comprising an NLS (nuclear localization signal) peptide or a nucleic acid encoding it.
28. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A composition for base editing of plant cells, comprising a chloroplast transit peptide or a nucleic acid encoding it.
29. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A composition for base editing of plant cells, comprising a mitochondrial targeting signal (MTS) or nucleic acid encoding it.
30. The composition according to claim 29, further comprising a nuclear export signal (NES) or a nucleic acid encoding it.
31. The composition according to any one of claims 27 to 30, characterized in that the fusion protein is transmitted to plant cells by the following: Injection using a Gene gun (Bombardment); PEG-mediated protoplast transfection; Protoplast transfection (electroporation); or Protoplast injection (microinjection) using microinjectors.
32. The composition according to claim 31, characterized in that the nucleic acid is transmitted to plant cells through the following: Transformation using Agrobacterium, such as Agrobacterium tumefaciens or Agrobacterium rhizogene; Viral phenotypic infection; Injection using a Gene gun (Bombardment); PEG-mediated protoplast transfection; Protoplast transfection (electroporation); or Protoplast injection (microinjection) using microinjectors.
33. A base editing composition according to any one of claims 27 to 30, relating to mitochondria, chloroplasts, or plastids (white or chromoid bodies) of the aforementioned plant.
34. A composition according to any one of claims 27 to 30, further comprising a TAL (Transcription Activator-Like) effector (TALE) that cleaves wild-type DNA base sequences but not edited base sequences; and a FokI nuclease or nucleic acid encoding it, or a ZFN (zinc finger nuclease) or nucleic acid encoding it.
35. A method for editing the bases of nuclear, mitochondrial, or plastid DNA of a eukaryotic cell, comprising the step of processing the composition according to any one of claims 27 to 30.
36. The method according to claim 35, further comprising a TALEN or ZFN or nucleic acid encoding a TALEN or ZFN that cleaves wild-type DNA base sequences but does not cleave edited base sequences, thereby increasing the efficiency of base editing.
37. A method for base editing plant cells, comprising the step of treating plant cells with a composition according to any one of claims 27 to 30.
38. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A method for base editing plant cells, comprising the step of treating plant cells with an NLS (nuclear localization signal) peptide or a nucleic acid encoding it.
39. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A method for base editing plant cells, comprising the step of treating plant cells with a chloroplast transit peptide or a nucleic acid encoding it.
40. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A method for base editing plant cells, comprising the step of treating plant cells with a mitochondrial targeting signal (MTS) or a nucleic acid encoding it.
41. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A composition for base editing of animal cells, comprising an NLS (nuclear localization signal) peptide or a nucleic acid encoding it.
42. A fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and A composition for base editing of animal cells, comprising a mitochondrial targeting signal (MTS) or nucleic acid encoding it.
43. The composition according to claim 42, further comprising a nuclear export signal or a nucleic acid encoding the same.
44. The composition according to claim 42, further comprising a TAL (Transcription Activator-Like) effector (TALE) that cleaves wild-type DNA base sequences but not edited base sequences; and a FokI nuclease or nucleic acid encoding it, or a ZFN or nucleic acid encoding it.
45. A method for base editing animal cells, comprising the step of treating animal cells with the composition described in claim 41 or 42.
46. A method for base editing animal cells, comprising the steps of treating animal cells with a fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and an NLS (nuclear localization signal) peptide or a nucleic acid encoding the same.
47. A method for base editing animal cells, comprising the steps of treating animal cells with a fusion protein according to any one of claims 1 to 21 or a nucleic acid according to claim 22; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding it.
48. The method according to claim 46 or 47, further comprising a TALEN or ZFN or nucleic acid encoding a TALEN or ZFN that cleaves wild-type DNA base sequences but does not cleave edited base sequences, thereby increasing the efficiency of base editing.
49. A composition for A-to-G base editing in prokaryotic or eukaryotic cells comprising the fusion protein described in claim 19 or a nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or a variant thereof is derived from bacteria and is specific to double-stranded DNA.
50. A composition for A-to-G base editing in prokaryotic or eukaryotic cells comprising the fusion protein described in claim 19 or the nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease. The cytosine deaminoase or its variant of the aforementioned fusion protein is derived from bacteria and is specific to double-stranded DNA. A composition in which a DNA-binding protein is bound to the N-terminus of the cytosine deaminose or its variant, and a DNA-binding protein is bound to the C-terminus of the adenine deaminose of the fusion protein.
51. A composition for C-to-T base editing in prokaryotic or eukaryotic cells comprising the fusion protein described in claim 19 or a nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or a variant thereof is derived from bacteria and is specific to double-stranded DNA.
52. A method for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or nucleic acid encoding it according to claim 19, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or a variant thereof is of bacterial origin and is specific to double-stranded DNA.
53. A method for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the step of treating the prokaryotic or eukaryotic cells with the fusion protein or nucleic acid encoding it according to claim 19, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease. The cytosine deaminoase or its variant of the aforementioned fusion protein is derived from bacteria and is specific to double-stranded DNA. A method wherein a DNA-binding protein is bound to the N-terminus of the cytosine deaminose or its variant, and a DNA-binding protein is bound to the C-terminus of the adenine deaminose of the fusion protein.
54. A method for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the step of treating prokaryotic or eukaryotic cells with the fusion protein or nucleic acid encoding it and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-related nuclease, and the cytosine deaminose of the fusion protein or a variant thereof is of bacterial origin and is specific to double-stranded DNA.